Innovating from the core: validating an AI-adaptive music product for anxiety relief inside a music-distribution company.
The full dissertation document, in business-school order: front matter, ten chapters, references and appendices. Chapters 1–2 set the corporate-innovation question and the literature; chapter 3 analyses the company and the industry with VRIO, PESTEL, Porter's five forces, a competitor and positioning analysis, TAM/SAM/SOM and SWOT; chapter 4 states the method; chapters 5–8 report what the evidence says, what business model it implies, what has been built, and where the pre-registered gates currently stand; chapters 9–10 give the roadmap, the risk register and the conclusion.
Every figure in this document comes from a repository source. Framework cells the sources do not support are written as to verify in the cell rather than guessed. Per-number provenance: site/pages/dissertation.sources.md. Reference fields the repository does not supply: site/pages/dissertation.references-todo.md. Fact-checked companion: site/pages/journey.html with journey.sources.md and journey.audit.md.
0Front matter
Abstract
This dissertation asks how an established, profitable AI music-distribution company can use its own assets to enter an adjacent health-adjacent market, and whether the evidence justifies doing so. The setting is TDMusic, a Beijing-headquartered distributor with a catalogue, rights infrastructure, a published recommendation algorithm and worldwide distribution, and the candidate venture is Lilt: a heart-rate-adaptive calming-audio product positioned as general wellness. The study runs as insider action research organised through design thinking, with a convergent mixed-methods evidence base — a six-stream literature and market review, six interviews, a 200-row netnography, two focus groups with a blind stimulus test, a pre-registered seven-block survey, and a within-subject three-condition efficacy pilot — integrated through a Bayesian layer in which every source becomes a likelihood ratio shrunk by its quality, and every decision is taken against thresholds fixed before the data. Business-school frameworks (VRIO, PESTEL, five forces, competitor and positioning analysis, TAM/SAM/SOM, SWOT, Value Proposition and Business Model Canvases, Ansoff, Three Horizons, stage-gate) carry the strategic analysis. The central empirical finding is a content reversal: users in acute anxiety reject the melodic, artist-linked music that is the company's core asset and ask for featureless adaptive sound, while preferring real music for evening wind-down. The posterior for real music in acute states fell to 8 per cent; in wind-down it stands at 80 per cent. The verdict at the time of writing is NOT YET: engagement and willingness-to-pay gates are cleared, the efficacy pilot has not run and the rights check is unfinished.
Word count 248 (limit 250). Sources: research/METHODOLOGY_DESIGN.md §0–1; research/DECISION_MODEL.md §3, §5; company/COMPANY_BACKGROUND.md §1–2; interviews/user-interviews.md; site/assets/decision.js.
Keywords
Corporate entrepreneurship · corporate innovation · design thinking · insider action research · mixed methods · Bayesian process tracing · stage-gate · music and anxiety · heart-rate variability · adaptive audio · digital wellness · willingness to pay · VRIO · PESTEL · Business Model Canvas.
Executive summary
What this section establishes. The whole argument on one screen: the problem, the method, the four findings that matter, the recommendation, and the decision status. Everything below is the evidence for these five paragraphs.
Problem. TDMusic earns $10.5 M revenue and $2.2 M net income from music distribution and is a top-three distributor in China, a Tier-1 YouTube partner and a top supplier to Tencent Music (TDMusic Inc., 2022–2026). Its growth question is not whether it can distribute more music but whether its assets — catalogue, rights, an AI production pipeline, a published recommender and 220+ digital service providers — can be pointed at a new customer need. Music for anxiety relief is the candidate: anxiety affects an estimated 359 million people worldwide, of whom 27.6 per cent receive treatment (World Health Organization [WHO], 2023), and music listening lowers state anxiety across independent meta-analyses, albeit on trial evidence graded low quality (Bradt et al., 2013; de Witte et al., 2020).
Method. The founder-researcher runs the project as insider action research (Coghlan & Brannick, 2019), organised through design thinking and evaluated with decision analytics. Six evidence streams feed eight explicit beliefs; each belief has a prior with a written rationale, evidence items carrying likelihood ratios shrunk for quality, and a posterior computed in odds form (Fairfield & Charman, 2017; Humphreys & Jacobs, 2015). Proceed, pivot and kill thresholds were written down before the survey and before the pilot.
Findings. Four results carry the dissertation. First, a content reversal: three of three acute-anxiety interviewees independently rejected melody, lyrics and beat and asked for featureless sound; the focus-group blind test reproduced it (10 of 13 chose the neutral clip for the acute moment, 9 of 13 chose the real song for wind-down); the survey quantified it (24 per cent chose a real song in the acute scenario against 46 per cent neutral; 47 per cent chose the real song for wind-down; McNemar χ² = 14.29, p < .001). Second, engagement holds: 57 per cent of wearable owners would connect a device and 54 per cent value a measured result, against an "action gap" in which 38 per cent do nothing after a high stress reading. Third, price holds: the Van Westendorp range for US/EU contains the $6.99 test price and discounted purchase intent is 18 per cent. Fourth, efficacy is still borrowed: it rests on meta-analyses, not on the firm's own pilot, which has not run.
Recommendation. Do not fund an MVP yet. Split the product by arousal state rather than betting on one content type — an engineered neutral bed for the acute moment (which needs no catalogue rights) and real, artist-linked music for wind-down (which does). Run the legal check first because it is the cheapest decisive test, then the efficacy pilot, then re-apply the gates. Keep every claim inside the general-wellness line in all three regulatory regimes.
Decision status. NOT YET, as set out immediately below.
Decision status
Pre-registered PROCEED requires all five conditions below and that the efficacy pilot has run. Three of the five posterior conditions are met by the literature and survey alone; the two efficacy conditions were deliberately set above what the meta-analytic prior can reach, so they cannot be cleared without the pilot. No kill or pivot rule is triggered on the current numbers except the content pivot, which has already been acted on.
What is pending. (1) The E1 efficacy pilot — 20–30 participants × 3 sessions, scheduled 1–25 September; it carries pre-registered likelihood ratios of 4.0 or 0.30 on both efficacy nodes and is a necessary condition for PROCEED. (2) The legal check on adaptation and derivative rights for at least 50 catalogue tracks, in progress to 12 September, the only pending row graded decisive (LR 8.0 if positive, 0.15 if negative). (3) App Review — Lilt Adaptive Sound 1.0 (build 4) was submitted on 3 September, received a Guideline 2.1 information request, was answered and resubmitted on 4 September, and is WAITING_FOR_REVIEW. The landing-page A/B is listed as running to 20 September; no conversion figure exists yet.
Posteriors and gate arithmetic are computed exactly as site/assets/decision.js computes them, with the survey rows applied and the pilot and legal check pending. Gate rules: research/DECISION_MODEL.md §5. Dates: site/index.html roadmap; TODO_2026-08-19.md §11, §14–15.
1Introduction
What this chapter establishes. That this is a corporate-innovation problem inside a going concern rather than a start-up pitch, and that it reduces to two testable questions — is the need real and reachable, and can this firm serve it measurably — inside a fixed scope of three months and $10–20k.
1.1Context: TDMusic and the corporate-innovation question
TDMusic Inc. (太簇音乐, Tempered Digital Music) is a technology-driven digital-content distribution and monetisation company headquartered in Beijing and operating across the United States, Europe, Japan and China. Company documents report revenue of $10.5 M, net income of $2.2 M, a $70 M valuation and 52 employees, an angel round led by Challenger Capital, and a market position described as top-three music distributor in China, Tier-1 YouTube partner and top supplier to Tencent Music. Its 2022 catalogue figures record roughly 106,000 songs in reserve and 75,000 released, approximately 85 per cent owned or directly licensed, 1.05 billion streams between January and July 2022, and distribution to more than 200 countries through 220+ digital service providers.
The tutor's subject is innovation in the corporate. That framing matters because it changes what counts as an answer. A start-up would be asked whether the market is attractive; an established firm must additionally be asked whether it holds assets that make this market attractive to it, and whether it can run an exploratory venture without the exploitation business smothering it. The firm has already validated one instance of the general pattern: a Kindle Direct Publishing pilot that mined demand signals, produced content with AI assistance, published it and measured the result, reporting click-through rates two to three times benchmark, 4.2–4.5 star ratings and titles profitable within two to four months. The project reported here applies the same engine — demand signal, AI production, measured library — to the firm's core asset, music, in a vertical with a genuine clinical literature behind it.
1.2Problem statement
Three facts make the problem non-trivial. First, the need is large and under-served: an estimated 359 million people had an anxiety disorder in 2021 and roughly 27.6 per cent of those needing treatment receive it (WHO, 2023). Second, the incumbent solution set is commercially stressed: Calm's 2025 revenue is estimated at ~$210 M, down 24 per cent year on year with ~3.5 M subscribers, down ~500,000; Headspace's estimated revenue fell from ~$348 M in 2024 to ~$140 M ARR in 2025 alongside a 13 per cent workforce cut, and industry-wide 30-day retention for meditation apps is about 4.7 per cent with an average of one to four lifetime sessions (Creswell & Goldberg, 2025). Third, the evidence base the whole category leans on is thin: only Calm and Headspace have any randomised-trial coverage at all, half of Headspace's own trials disclose a conflict of interest, and no equivalent independent review exists for Endel, Brain.fm or any Chinese competitor (O'Daffer et al., 2022).
So the problem is not "can we build a calming-audio app" — plainly we can, and one is already in TestFlight. The problem is whether a firm whose advantage is real music and the rights to it can win in a market whose users, in the state where they most need help, may not want real music at all. That tension is the dissertation's spine.
1.3Research questions and objectives
| Question | Claim to establish | Evidence type | Where it is estimated |
|---|---|---|---|
| RQ1 · NeedIs there a real, reachable, under-served need? | N1 a sizeable segment has frequent "need to calm down" moments; N2 current solutions under-serve specific outcomes; N3 the need is state-dependent (acute versus wind-down) and so is the preferred content | Survey (quantitative); interviews and focus groups (qualitative); netnography (qualitative) | Survey blocks B–C; focus-group themes T1–T2; interview synthesis |
| RQ2 · AbilityCan TDMusic serve it with its assets, measurably? | A1 content preference by state; A2 one session moves state anxiety and heart-rate measures; A3 users connect a wearable and value the result; A4 someone pays; A5 rights are feasible | Pilot experiment (quantitative); survey (quantitative); landing page (behavioural); legal check | Pilot mixed model; survey blocks D–F; landing conversion; legal memo |
| DecisionProceed, pivot or kill? | Posterior beliefs cross pre-registered gates | Decision analytics | Belief register and gate rules (chapter 8) |
Transcribed from research/METHODOLOGY_DESIGN.md §0.
The objectives follow directly: (i) establish whether the need exists and how it segments; (ii) establish which of the firm's assets actually transfer; (iii) measure whether one session changes anything, on self-report and on a device the user already owns; (iv) establish whether anyone pays, and at what price; (v) establish whether the rights permit the asset-leveraging version of the product; and (vi) end the phase with a documented decision rather than a pitch.
1.4Scope and boundaries
Four boundaries are worth stating explicitly because they constrain what this dissertation can conclude. The wellness boundary means no efficacy claim in this document is a medical claim, and no marketing language tested here uses the word "treat". The overseas-first boundary means the China figures in chapter 3 are context, not a market plan. The budget boundary means the efficacy pilot is n = 20–30, which the methodology declares in advance to be a feasibility signal for self-report and under-powered for physiology. The phase boundary means retention, the single most load-bearing commercial variable in this category, cannot be tested at all within the phase; it is carried as a watch item rather than a gate.
1.5Structure of the dissertation
Chapter 2 reviews the literature in five strands and names the gap. Chapter 3 analyses the company and the industry with the standard strategy toolkit. Chapter 4 sets out the method, including the Bayesian layer and the pilot design. Chapter 5 reports the findings by design-thinking mode and closes with the belief register. Chapter 6 turns the findings into a concept and a business model. Chapter 7 describes what has actually been built and how the algorithm works. Chapter 8 states the validation plan and the decision rule. Chapter 9 gives the roadmap, the risk register and the resource boundary. Chapter 10 answers the research questions as far as the evidence currently allows, and states what it cannot yet answer.
2Literature review
What this chapter establishes. That five separate literatures — corporate entrepreneurship, design thinking, music and autonomic arousal, digital wellness products, and decision-making under uncertainty — each supply one piece of this project, and that none of them answers the question the project actually faces.
2.1Corporate entrepreneurship and ambidexterity
The project's own framing document names four lenses for treating this as corporate rather than start-up innovation. The Ansoff matrix places music for anxiety as "diversification-lite": a new customer need (health) served by an adapted product (music) built on existing assets. McKinsey's Three Horizons assigns AI music creation and distribution to Horizon 1 as the cash engine, the music-for-wellbeing product to Horizon 2, and a closed-loop "music as digital therapeutic" to Horizon 3. O'Reilly and Tushman's ambidexterity justifies a small autonomous team exploring while the distribution business exploits. And core-asset leverage lists the unfair advantages a pure start-up does not have: an AI music generation pipeline, owned and controlled catalogue and rights, artist relationships, worldwide distribution infrastructure, and market knowledge in China and globally.
Two references in this strand do carry bibliographic detail in the repository and are used for the real-options reading of the phase: Bowman and Hurry (1993) on real options and organisational learning, and McGrath (1999) on managing failure as options. Their function here is precise rather than decorative: they license the sequencing rule used in chapter 8, in which each test is treated as a purchase of information and ordered by expected posterior shift per unit cost.
2.2Design thinking as an innovation method
Design thinking supplies the process rather than the epistemology. The method architecture uses the d.school's five modes — empathise, define, ideate, prototype, test — paced by the Double Diamond, whose contribution is the divergence–convergence discipline: the requirement to widen before narrowing, and to record what was discarded. Brown (2008) and Liedtka (2018) are the two design-thinking references the repository carries; Liedtka's contribution to this project is the claim that design thinking's value lies in reducing cognitive bias in innovation decisions, which is exactly the claim the Bayesian layer in §2.5 operationalises.
The method is not used uncritically. Design thinking as normally practised produces artefacts (personas, jobs-to-be-done statements, point-of-view and how-might-we questions) whose epistemic status is ambiguous: they look like findings but are often the facilitator's synthesis. This dissertation therefore separates artefacts from results explicitly, and reports the personas and JTBD statements in chapter 5 as illustrative artefacts written in the framework document, not as data-derived findings — because the framework itself writes them as examples ("E.g. …") rather than as outputs of coding.
2.3Music, anxiety and autonomic arousal: what the evidence says
Music listening reliably reduces anxiety across independent meta-analyses, but the pooled effect varies widely with population and outcome type, and the underlying trials are graded low quality by the strictest reviewers. The largest and most rigorous synthesis is a Cochrane review of 26 perioperative trials (N = 2,051) which found music reduced state anxiety by −5.72 STAI-S units, 95% CI [−7.27, −4.17], performed comparably to midazolam in one N = 327 trial, and rated the evidence "low" quality because of a lack of blinding (Bradt et al., 2013). A broader stress meta-analysis of 104 randomised trials (N = 9,617) reported d = .545 on psychological markers and d = .380 on physiological ones, with heart rate the strongest physiological marker at d = .456 and hormone/cortisol the weakest at d = .349 (de Witte et al., 2020). Harney et al. (2023) report d = −0.77 across 21 controlled studies, rising to −0.97 in a low-bias subset. Across reviews the range is roughly d = −0.30 to −0.97.
| Source | Result | Quality | What it licenses here |
|---|---|---|---|
| Bradt et al. (2013) — Cochrane, 26 RCTs, N = 2,051 | −5.72 STAI-S [−7.27, −4.17]; SMD −0.60; comparable to midazolam in one N = 327 trial | GRADE low (no blinding) | The prior on A2a and the effect size the pilot is powered against |
| de Witte et al. (2020) — 104 RCTs, N = 9,617 | psychological d = .545; physiological d = .380; heart rate d = .456; cortisol d = .349 | High for existence, heterogeneous | Heart rate, not cortisol, is the physiological outcome worth measuring |
| Harney et al. (2023) — 21 controlled studies | d = −0.77 [−1.26, −0.28]; −0.97 in the low-bias subset | Overlaps the two above → correlated discount | Upper end of the plausible effect; participant-selected music trends larger but not significantly |
| Lassner et al. (2025) — k = 23, n = 1,084 | SMD = 0.47 [0.27, 0.66]; not maintained at follow-up | GRADE very low | Delivery mode does not matter much; durability does |
| Panteleeva et al. (2018) | self-report d = −.30 significant; psychophysiological effect not significant | Medium-high | Self-report and physiology need not move together — the A2a / A2b split |
| Bernardi et al. (2006); Mitrovic & Paladin (2026) | Slow tempo (~60–80 bpm) raises vagal tone; fast tempo shifts to sympathetic dominance | Classic primary study; recent corroborating review | The mechanism the adaptation rule implements — start near the listener's rate, then lead down |
| Goessl et al. (2017) — 24 studies, n = 484 | HRV biofeedback g = 0.81 pre-post; g = 0.83 versus control | High, but for breathing-paced protocols | Supports the closed-loop mechanism, not music-driven adaptive audio |
| Pelletier (2004) — 22 studies | d = .67 for music and music-assisted relaxation on stress arousal | Foundational | Earliest quantitative synthesis; consistent in direction with the later, larger reviews |
All rows from research/SYNTHESIS.md §1 and §5 and research/findings/clinical-evidence-verified.md (fact-check verdicts: 7 confirmed, 1 corrected, 0 unsupported of 8 re-checked).
Three qualifications carry directly into the design. First, direction: relaxation raises heart-rate variability and lowers heart rate; the reverse is a recurring error in this literature's popular treatment and is checked explicitly throughout this document. Second, personalisation is weaker than intuition suggests: participant-selected music trended toward larger effects than researcher-selected music in the largest anxiety meta-analysis but the difference was not statistically significant, and preferred tempo appears driven more by an individual's spontaneous motor tempo than by genre preference (Harney et al., 2023). Third, generative audio has almost no evidence: the sole located peer-reviewed Endel-linked study measured focus rather than anxiety, was funded by Endel and Arctop, and used Arctop-employed authors (Haruvi et al., 2022).
2.4Digital wellness products and adaptive audio
Endel is the clearest commercial precedent for closed-loop biometric-adaptive audio: it modulates generative soundscapes using real-time Apple Watch heart rate compared against resting heart rate, plus motion, time of day and light; it won Apple's Watch App of the Year in 2020; and its primary press release with Universal Music Group cites three million monthly listeners (PR Newswire, 2023). Two features of the precedent matter more than the headline. First, the signal is heart rate, not HRV — HealthKit exposes HRV only as opportunistically sampled values that can lag up to about 30 minutes, and Oura's API returns HRV at five-minute granularity but only from sleep, available the following morning, with no streaming endpoint. Any product claiming "real-time HRV adaptation" is therefore either using heart rate or using stale values. Second, Endel's own strategic response to the "is this even music" critique — paying Grimes, James Blake, Miguel and Warner's Spinnin' Records for artist-attached functional music — is itself revealed evidence that abstract algorithmic sound under-satisfies demand.
The category's structural problem is retention, not acquisition. Roughly 4.7 per cent of meditation-app users are still using the app after 30 days, with an average of one to four lifetime sessions, even though Calm and Headspace jointly hold 96 per cent of category monthly active users (Creswell & Goldberg, 2025). Both leaders showed revenue and subscriber declines in 2025 while market-research firms continued to forecast double-digit compound growth — a tension this dissertation treats as a real signal about market maturity rather than resolving in either direction.
2.5Decision-making under uncertainty: Bayesian process tracing and stage-gates
The methodological move that distinguishes this project from a standard design-thinking write-up is that the "evaluate" step is not left to judgement. Each critical assumption is given an explicit prior with a written rationale; each piece of evidence is converted into a likelihood ratio using a published calibration rubric; likelihood ratios are shrunk in log space according to evidence quality; and posteriors are compared with thresholds fixed in advance. The method precedent is Fairfield and Charman's (2017) explicit Bayesian analysis for process tracing and Humphreys and Jacobs's (2015) Bayesian mixing of methods; evidence-quality grading follows the GRADE Working Group (Guyatt et al., 2008); and the pre-registration and kill-criterion discipline follows Ries (2011) and Bland and Osterwalder (2019).
The stage-gate reading sits on top: each test is a purchase of information whose cost, duration and expected posterior shift are listed before it is run, so the ordering of tests is a real-options decision rather than a scheduling convenience (Bowman & Hurry, 1993; McGrath, 1999).
2.6Gap and contribution
Four gaps are visible once the strands are laid side by side. (1) The efficacy literature does not cover the product. The strong evidence is for music listening in supervised, mostly perioperative settings, and for breathing-paced HRV biofeedback; there is no meta-analysis, and no adequately powered trial, of biometric-adaptive audio delivered by a consumer app. (2) The content question is unasked. No controlled anxiety trial directly compares real, personally meaningful music against engineered neutral sound — which is precisely the comparison on which this firm's asset advantage depends. (3) Corporate-innovation research rarely shows its decision arithmetic. Ambidexterity and stage-gate research says a firm should kill projects on evidence, but published cases seldom show the belief, the threshold and the update. (4) Design-thinking write-ups rarely distinguish artefact from result.
The contribution is correspondingly modest and specific: a fully documented, pre-registered corporate-innovation decision in which the belief state, the likelihood ratios, the quality shrinkage and the gates are all published and re-runnable, applied to a question — real music versus engineered neutral sound for acute anxiety — that the literature has not addressed, inside a firm whose principal asset happens to sit on one side of that question.
3Company and industry analysis
What this chapter establishes. Which of the firm's assets actually transfer to a wellness-audio venture and which do not, and what kind of industry it would be entering — structurally unattractive on rivalry, buyers and substitutes, with a narrow defensible position built on rights plus measurement.
3.1Company profile and asset inventory: a VRIO analysis
The asset inventory is the "ability" half of the research question. The table below runs the seven assets recorded in the company file through VRIO — is the asset Valuable for this venture, is it Rare, is it costly to Imitate, and is the firm Organised to capture the value — and states the competitive implication that follows.
| Asset | Evidence in the repository | V | R | I | O | Implication for the venture |
|---|---|---|---|---|---|---|
| AI distribution and matching algorithmBAS / xDeepFM recommender | Published as Lu et al. (2021), PeerJ Computer Science, 7, e716; company claims ~10× marketing ROI; scenario- and preference-based auto-matching of music to users. | Yes | to verify | to verify | Yes | The same engine can match calming music to stress contexts, and the peer-reviewed paper is citable credibility. But the method is published, and the repository holds no benchmark against competing recommenders, so rarity and inimitability cannot be asserted. Competitive parity plus credibility, not advantage. |
| AI production pipelinecovers, artwork, voice synthesis, video | AI-assisted covers, audio artwork and intro content; video pipeline with four of six steps automated; talknet2 voice synthesis mature for English; UE5 digital humans. The business-model deck notes AI tools for artists "especially in meditation and lofi genres". | Yes | to verify | to verify | Yes | Low-cost production of calm arrangements and adaptive variants — the closest existing capability to what the product needs. The description is 2022-era and the company file flags it for refresh, so no claim is made about its current state of the art. |
| Rights and catalogue~106k reserve / 75k released, ~85% owned or direct | 100k+ songs, majority direct-licensed; rights-management and royalty infrastructure; composition rights including cover and adaptation rights for part of the catalogue. | Yes | Yes | Yes | to verify | The raw material for real-music calm content, and the one genuinely non-imitable asset here: rights are legally exclusive. But organisation to capture turns entirely on whether adaptation and derivative rights exist per track — the top legal assumption A5, posterior 0.65, legal check pending. Potentially sustained advantage, gated on A5. |
| Distribution and platform relationships220+ DSPs | Direct deals with Spotify, Apple and YouTube; Peloton, Tesla and Virgin Airlines as distribution endpoints; top-three distributor in China, Tier-1 YouTube partner, top TME supplier (2022 figures). | Yes | Yes | to verify | Yes | Global launch capability on day one, and the Peloton-type wellness endpoints are directly relevant B2B2C precedents already in house. Inimitability is unclear: relationships are contractual and can be built with scale, and the repository holds no evidence on switching costs. |
| Demand-signal methodthe KDP publishing pilot | Reddit demand mining → AI production → publish → measure. Reported click-through 2–3× benchmark, 4.2–4.5★ ratings, titles profitable in two to four months. | Yes | to verify | to verify | Yes | The continuity argument of the whole thesis: the same engine, a new vertical. As a repeatable process rather than a protected asset, its rarity and inimitability are unevidenced — the repository records the pilot's results, not a comparison with anyone else's. |
| Artist network | Kurt Hugo Schneider (13.4M YouTube subscribers — relationship scope flagged for clarification), Cécile Corbel, regional artists across the US, EU and Asia. | Yesfor concept C3 | to verify | to verify | to verify | Artist-branded calm content is the credibility Endel lacks, and is the cheapest test of real-music demand. Every VRIO column except value is unverifiable because the company file itself flags the relationship scope (China agency operations versus distribution rights) as unresolved. |
| Cash engine and multi-region operations | $2.2M net income; 52 employees; teams in four regions; $10–20k ring-fenced for this phase. | Yes | No | No | Yes | Not an advantage — profitability is not rare and not hard to imitate — but it is the organisational condition for honest exploration: Horizon 1 funds Horizon 2, and a ring-fenced budget makes a kill decision survivable. |
Assets and evidence transcribed from company/COMPANY_BACKGROUND.md §2; A5 posterior from research/DECISION_MODEL.md §3; budget from DESIGN_THINKING_FRAMEWORK.md framing update. VRIO verdicts: 28 cells (7 assets × V, R, I, O) — 17 supported by the cited evidence, 11 marked "to verify".
Reading of the VRIO. Only one asset — rights and catalogue — clears value, rarity and inimitability together, and its fourth column is exactly the question the pending legal check answers. That is a strategically uncomfortable result: the firm's single defensible asset is the one whose usability is unknown, and the user evidence in chapter 5 says that asset is the wrong content for the moment of highest need. Chapter 6 resolves this by segmenting the product rather than the asset.
3.2Macro environment: PESTEL
PESTEL: 6 of 6 cells sourced; 0 marked "to verify". The environmental cell is sourced as a reasoned "not material" rather than as a data point.
3.3Industry structure: Porter's five forces
| Force | Rating | One-line evidence | Consequence for this venture |
|---|---|---|---|
| Competitive rivalry | High | Calm and Headspace jointly hold 96% of category monthly active users, and both declined in 2025 (Creswell & Goldberg, 2025; Business of Apps, 2026a, 2026b). | Do not compete on catalogue breadth or brand spend; compete on the measured result. |
| Buyer power | High | "Another subscription" was the leading barrier for 49% of survey respondents; both focus groups preferred the feature bundled inside something already paid for. | Strengthens concept C2 (licence into a device or platform) relative to standalone B2C. |
| Threat of substitutes | High | 90% already use music or audio at least weekly; 61% "hired" music for calming today; Apple ships Vitals free with every Series 9+ watch. | The product must beat free music the user already has, not beat Calm. |
| Threat of new entrants | Medium–high | Endel's entire disclosed funding is $22.1M across five rounds; the build barrier is low. Rights, an evidence base and distribution are the only real barriers. | Speed matters less than evidence; the pilot is a moat-building activity, not a compliance exercise. |
| Supplier power | Medium | Apple App Review issued a Guideline 2.1 information request on the first submission; Spotify playback requires Premium; Oura gates its API behind an active membership. Against that, ~85% of the catalogue is owned or directly licensed, internalising the music supplier. | Platform risk is real and current; music-supplier risk is largely internalised — if A5 clears. |
Five forces: 5 of 5 rated with cited evidence; 0 "to verify".
3.4Competitor analysis
3.4.1 Competitor profiles
| Player | Owner / funding | Model and pricing | Wearable integration | Independent evidence | Documented weakness |
|---|---|---|---|---|---|
| Calm | Private; $218M raised; $2bn valuation (2020 Series C), no newer disclosure | $14.99/mo, $69.99/yr, $399.99 lifetime; B2B2C "Calm Health" claims 39M+ covered lives across 100,000+ organisations | None disclosed; partnered with Ozlo earbuds (Sept 2025) | One RCT worldwide, no disclosed conflict of interest (O'Daffer et al., 2022) | 2025 revenue ~$210M, −24% YoY; ~3.5M subscribers, −500k; complaint patterns centre on billing, not content |
| Headspace | Merged with Ginger (2021) into Headspace Health at a $3bn valuation; ~$216M + ~$220M prior funding | Consumer subscription plus B2B2C; ~2,700 employer and health-plan partners, ~100M lives. Consumer price point: to verify | None disclosed | 14 RCTs — the largest base in the category — but 7 of 14 disclose a conflict of interest and only 36% were pre-registered | Estimated revenue fell from ~$348M (2024) to ~$140M ARR (2025); 13% workforce cut Nov 2024; estimates conflict sharply |
| Endel | Private; ~$22.1M total, $15M Series B (Apr 2022, Waverley Capital and True Ventures) | Subscription. Valuation ~$47.5M and revenue ~$15.8M trace to a single uncited aggregator page — low confidence. Price point: to verify | Real-time Apple Watch heart rate (not HRV) plus motion, time and light; Apple Watch App of the Year 2020. Oura integration announced but not independently confirmed | Sole peer-reviewed study is Endel- and Arctop-funded with Arctop-employed authors and measures focus, not anxiety (Haruvi et al., 2022) | Its own chief commercial officer concedes algorithmic audio "is paid the same" as composed music despite being "very different"; the artist deals are revealed evidence that generative sound under-satisfies |
| Brain.fm | Private; disclosed funding implausibly small (~$125k seed plus an NSF grant of unconfirmed size) — likely incomplete records | ~$9.99/mo or ~$99.99/yr; third-party sources show inconsistent historical pricing ($6.99/mo, $49.99/yr elsewhere) | None found | Peer-reviewed Communications Biology study with Northeastern's MIND Lab on attention-network engagement (Woods et al., 2024) | Proprietary stimuli and a commercial conflict of interest inside its own validation study; the study measures attention, not anxiety |
| Tide (潮汐) | China; "near ¥10M" Pre-A led by Panda Capital; ~30-person team | Subscription-only, no advertising: ¥218/yr (~$30) or ¥27/mo, plus a ¥8/mo sound-card tier | None found | None found | Structurally sub-scale; growth claims are founder-reported and unaudited; the domestic sleep-app sector is described by trade press as having gone "from hot commodity to dispensable" |
| MedRhythms | Private; $34M total funding; CEO replaced July 2025 by a former public-medtech CEO | Prescription digital therapeutic; Medicare DME code HCPCS E3200, final payment determination effective 1 Apr 2025. List price: to verify | Proprietary sensor-embedded device, not consumer wearables | The only player here with an FDA regulatory pathway: Breakthrough Device Designation (June 2020) for InTandem; MOVIVE holds a lower Class II Rx-only listing | Narrow indication (stroke and Parkinson's gait, not anxiety); five years from designation to reimbursement |
competitors-verified.md rows 1–24, 28–29; SYNTHESIS §2; market-wtp-verified rows 7, 49–51. Every dollar figure for Calm, Headspace and Endel traces to third-party aggregators (Latka, Business of Apps, Tracxn) that disagree with each other by wide margins; the direction is consistent, the precision is not.
3.4.2 Competitor marketing mix (4P)
| Player | Product | Price | Place | Promotion |
|---|---|---|---|---|
| Calm | Meditation, Sleep Stories, music; standalone Calm Sleep app (Sept 2025) with 300+ hours of sleep content and 500 Sleep Stories; Calm Health clinical programmes rooted in CBT, DBT and ACT with PHQ/GAD screening | $14.99/mo · $69.99/yr · $399.99 lifetime | iOS and Android stores; employer and health-plan channels (Solera network, Jan 2026); Ozlo earbuds hardware bundle | Hardware bundle discount ($50 with Ozlo) and press. Channel mix and spend: to verify — no marketing data in the repository |
| Headspace | Meditation plus Ginger coaching and therapy inside Headspace Health | to verify — the repository records the B2B2C model and subscriber counts but no consumer price | ~2,700 employer and health-plan contracts covering ~100M lives; consumer app stores | to verify |
| Endel | Generative soundscapes by mode; artist collaborations (Grimes, James Blake, Miguel); UMG partnership (May 2023) using catalogue stems; Warner's Spinnin' Records committed to 50 AI-generated wellness albums | Subscription; price point to verify | iOS, Android, macOS and Apple Watch (standalone and companion modes) | Label and artist partnerships are the promotional engine: the UMG and WMG deals generate the coverage |
| Brain.fm | Patented amplitude-modulated functional music for focus, relaxation and sleep | ~$9.99/mo · ~$99.99/yr (fluctuates) | App stores and web; direct vendor pricing page | The peer-reviewed study is the marketing asset — announced via press release framed as "world's first science-backed purpose-built focus music" |
| Tide (潮汐) | White noise, meditation, sleep and focus content; subscription-only with explicitly no advertising | ¥218/yr · ¥27/mo · ¥8/mo sound card | Chinese app stores; partnerships with Apple (employee-wellness procurement), Santonban coffee, Atour Hotels and Lululemon | Brand partnerships and founder interviews; spend and channel mix to verify |
| MedRhythms | InTandem (MR-001) rhythmic auditory stimulation for chronic-stroke gait; MOVIVE (MR-005) for Parkinson's gait | Medicare-reimbursed under HCPCS E3200 from 1 Apr 2025; list price to verify | Clinician prescription; national distribution agreement with Edwards Health Care Services (Sept 2025) | Regulatory milestones and the UMG catalogue partnership; no consumer promotion |
Competitor 4P: 24 cells (6 players × 4 Ps) — 18 sourced, 6 marked "to verify". Competitor profiles: 36 cells, 33 sourced, 3 marked "to verify". Sources as §3.4.1 plus competitors-verified rows 5–7, 14, 19, 23, 29.
3.4.3 Positioning map
The map makes the strategic opening legible. The upper half is nearly empty: only MedRhythms combines real music with a regulatory-grade evidence position, and it does so in a narrow, prescription-gated, non-anxiety indication. Endel occupies the measurement corner without independent evidence; Brain.fm has evidence without user measurement; Calm, Headspace and Tide have neither. The position this project intends — measured on the user's own device, evidenced by a pre-registered trial, and content-segmented by arousal state — is unoccupied. It is also unproven, which is why the marker is dashed.
3.5Market sizing: TAM, SAM and SOM
Inputs: market-wtp-verified rows 1–2, 49, 51, 57–58; SYNTHESIS §3, §6; surveys/analysis/results.json via survey-results.html (wearable.d1_top2_owners, intent, vw.US_EU). TAM/SAM/SOM: 9 inputs sourced, 5 marked "to verify" (European revenue share, US/EU share of the wearable installed base, store commission, CAC, churn in months).
3.6SWOT
- Owned rights at scale — ~106k songs in reserve, ~85% owned or directly licensed, with royalty infrastructure. COMPANY_BACKGROUND §1–2
- Day-one global distribution — 220+ DSPs including Peloton, Tesla and Virgin Airlines endpoints. §2
- A peer-reviewed recommender — Lu et al. (2021), PeerJ CS, 7, e716.
- A proven engine — the KDP pilot: click-through 2–3× benchmark, profitable in 2–4 months.
- Cash to fund exploration — $2.2M net income; $10–20k ring-fenced for this phase.
- A working product already — v0.5.1 on the web, TestFlight builds, 1.0 submitted to the App Store with nine owned tracks.
- No efficacy evidence of its own — E1 has not run; A2a 0.85 and A2b 0.56 are borrowed from meta-analyses.
- The core asset is the wrong content for the core moment — P(A1a) ≈ 0.08.
- Rights for adaptation unaudited — A5 0.65; per-track derivative rights flagged unknown in the company's own file.
- Company figures stale and unreconciled — founding date, catalogue and stream counts (2022), artist-relationship scope.
- Insider-researcher position — the founder is also the analyst; bias controls are declared but the conflict is structural.
- A known product inconsistency — the iOS build ships no Bluetooth plugin while the store description mentions a chest strap; to be corrected in 1.1.
- HRV cannot be steered live on any consumer API, so the "closed loop" is heart-rate-driven by necessity.
- The action gap — 38% of survey respondents do nothing after a high stress reading; focus-group owners read the score "like the weather". This is the hole a closed loop fills.
- A concrete unmet spec — complaints about existing neutral audio are specific and engineerable: sharp frequencies, low hum causing nausea, "fake healing" reactance, beat entrainment.
- A real wind-down segment for real music — 47% [39–56%] chose the real song for wind-down; 35% rated real artists top-2 in value.
- An evidence vacuum in the category — no independent review exists for Endel, Brain.fm or any Chinese competitor.
- B2B2C routes already in house — Peloton-type endpoints, plus documented label appetite for functional music (UMG and WMG with Endel).
- The retention cliff — ~4.7% D30 industry-wide, 1–4 lifetime sessions; A6 posterior 0.25.
- A maturing, shrinking category — Calm −24% revenue and −500k subscribers in 2025; Headspace revenue down sharply with layoffs.
- Platform disintermediation — Apple ships Vitals free with every Series 9+ watch; Oura gates its API behind a membership.
- Regulatory claim creep — any "treat anxiety" wording triggers SaMD status in the US and Advertising Law Article 17 exposure in China.
- The DTx commercial precedent — Pear filed Chapter 11 in April 2023 and Akili was taken private at $0.434/share in mid-2024, both on reimbursement rather than efficacy.
- Gatekeeper risk is live — the first App Store submission drew a Guideline 2.1 information request.
- Subscription fatigue — 49% named "another subscription" as their barrier.
SWOT: 25 cells, all 25 sourced, 0 "to verify". Sources: COMPANY_BACKGROUND §1–2, §4; DECISION_MODEL §3; SYNTHESIS §2–§4, §6; survey results as published; focus-groups SYNTHESIS T3, T6–T7; TODO §11, §14–16; wearables-verified; regulation-verified.
4Methodology
What this chapter establishes. That the evidence base is a convergent mixed-methods design in which every source has one declared job, and that the decision rule — priors, likelihood ratios, quality shrinkage, gates — was written down before any of the data arrived.
4.1Research philosophy and insider action research
The frame is insider action research (Coghlan & Brannick, 2019): the founder-researcher diagnoses, plans, acts, evaluates and learns inside the firm. This is a genuine methodological commitment rather than a label, because it fixes what counts as rigour. An outsider's rigour is distance; an insider's rigour is declared procedure. Reflexivity and bias control are therefore part of the method (§4.7), not an appendix, and the project's own tutor presentation already positions the whole thesis as action research with five method components — action research, strategy formulation, innovation frameworks, data-driven content identification and quantitative pilot analysis. The design-thinking phase reported here is the next cycle of that same spiral: diagnose and plan for the highest-potential content vertical.
4.2Design-thinking process and artefacts
Process is supplied by the d.school's five modes paced by the Double Diamond; its contribution is the divergence–convergence discipline. Each mode produced named artefacts, listed below so the reader can distinguish what was produced from what was found.
| Mode | Artefacts produced | Epistemic status |
|---|---|---|
| Empathise | Six-stream literature and market synthesis; 6 interviews; 200-row netnography; a ten-product competitor teardown | Evidence. The synthesis follows the fact-checked -verified files, using corrected figures and flagging unsupported claims. |
| Define | Territory scoring matrix v1 → v2 (Monte-Carlo re-score); personas; jobs-to-be-done; point-of-view and how-might-we statements | Mixed. The matrix is evidence-driven and re-scored after the interviews. The personas, JTBD, POV and HMW statements are written in the framework document as examples, not as coded outputs, and are reported here as artefacts only. |
| Ideate | Ten seeded directions plus session output; a 2×2 convergence; three concept cards C1–C3; an assumption map that became the belief register | Artefacts. The framework asks for "≥ 15 raw ideas"; no record exists of how many were actually generated, so no count is claimed as a result. |
| Prototype | Clickable prototype; landing A/B (two headline variants); Wizard-of-Oz kit; product blueprint v1; the Lilt PWA and iOS build | Artefacts, plus one behavioural instrument (the landing page) whose result does not yet exist. |
| Test | Survey v2 (7 blocks, pre-registered); two focus groups with a blind stimulus test; E0 calibration and E1 efficacy designs; the gate check | Evidence, with the pilot outstanding. |
4.3Mixed-methods design and the triangulation matrix
The architecture is a convergent mixed-methods design with a sequential quantitative core (Creswell & Plano Clark, 2018). The qualitative strand runs interviews (n = 6) → netnography (n = 200) → focus groups (2 × 6–8); the quantitative strand runs literature priors → survey (n ≥ 100) → pilot experiment (n ≥ 20); behavioural evidence (landing A/B) and the legal check enter alongside. Integration is explicit: each source enters the decision model as a likelihood ratio bounded by that source's quality, so nothing is double-counted and nothing is weighted by rhetoric.
4.4Instruments
| Construct | Instrument | Type | Validity handling |
|---|---|---|---|
| Anxiety symptoms | GAD-2 (Kroenke et al., 2007) | Validated | Used verbatim; sensitivity and specificity known at cutoff ≥ 3 |
| Perceived stress | PSS-4 (Cohen & Williamson, 1988) | Validated | Verbatim; α reported; convergent r with GAD-2 expected .5–.7 |
| Under-served outcomes | Outcome-driven innovation, importance × satisfaction (Ulwick, 2002) | Adapted | Outcome statements drawn from interview JTBD; opportunity score = I + max(I − S, 0); threshold ≥ 6 |
| Content preference by state | Within-subject scenario forced choice | Ad hoc, pre-registered | Order randomised; McNemar (1947), exact test for small discordant counts; segmented by GAD-2 |
| Acoustic tolerances | Nine-item −2…+2 matrix | Ad hoc | Items are the interview complaints verbatim; used as an engineering spec, not as a scale |
| Acceptance | TAM perceived usefulness / ease of use / behavioural intention (Davis, 1989; Venkatesh & Davis, 2000) | Adapted, 2 items each | α ≥ .70 target; OLS with standardised betas; PLS-SEM as robustness if n ≥ 150 |
| Feature value | Kano functional/dysfunctional pairs (Kano et al., 1984) with Better/Worse coefficients (Berger et al., 1993) | Standard | Evaluation table; dominant category; category reported by segment |
| Price sensitivity | Van Westendorp price sensitivity meter (van Westendorp, 1976) | Standard | Cumulative curves; OPP, IPP, range [PMC, PME]; per currency; monotonicity enforced per respondent |
| Purchase intent | Five-point scale at a shown price | Standard | Top-2 box × 0.5 hypothetical-bias discount (Morwitz et al., 2007) |
| Behavioural WTP proxy | Landing-page visitor → waitlist conversion | Behavioural | ≥ 5% pre-registered, per positioning variant |
| State anxiety (pilot primary) | STAI-S (Spielberger, 1983); six-item short form (Marteau & Bekker, 1992) | Validated | Pre and post each session; the primary outcome |
| Physiological state (pilot) | Heart rate (bpm) and RMSSD/SDNN (ms) from the wearable; Polar H10 sub-sample | Objective, noisy | Artefact rules, calibration, residualisation (§4.6) |
| Expectancy (pilot covariate) | Two items including a credibility item (Devilly & Borkovec, 2000) | Adapted | Entered as a covariate to control placebo |
| Qualitative themes | Reflexive thematic analysis (Braun & Clarke, 2006); focus groups (Krueger & Casey, 2015; Morgan, 1997); netnography (Kozinets, 2020); critical incidents (Flanagan, 1954) | Qualitative | Hybrid codebook; second coder on 20% of transcripts, Cohen's κ ≥ .70; saturation when a group adds no new code |
Transcribed from research/METHODOLOGY_DESIGN.md §2 and surveys/questionnaire-v2.md §Analysis plan.
4.5The Bayesian decision engine
For an assumption H with prior probability P(H): prior odds O₀ = P(H) / (1 − P(H)); each evidence item contributes a likelihood ratio LRᵢ = P(Eᵢ | H) / P(Eᵢ | ¬H); posterior odds O = O₀ × ∏LRᵢ, and the posterior probability is O / (1 + O). Independence is assumed for the multiplication; where items are clearly correlated the likelihood ratios are shrunk toward 1 before multiplying rather than treated as independent.
Worked example, quoted from the model. Three acute-anxiety interviews rejecting real music would each be "strong" disconfirmation at LR ≈ 0.3. They are qualitative and self-selected, so k = 0.5 gives 0.55 each; three items recruited through one channel are treated as roughly two effective items, 0.55² ≈ 0.30, rounded conservatively up to 0.35 to avoid over-weighting three conversations. This is the single most consequential number in the dissertation and it is deliberately the most conservative reading of the evidence available.
Priors come from base rates or literature where any exist, otherwise 0.50; every prior carries a one-line rationale, and the live version of the model lets a reader move any prior and see how much the posterior depends on it. Priors are meant to be argued about, not hidden.
4.6Efficacy pilot design (E0 and E1)
E0 — calibration. Roughly eight participants wear an Apple Watch and a Polar H10 concurrently across three sessions. Outputs are the intraclass correlation and mean absolute percentage error for heart rate and RMSSD, plus a regression-calibration coefficient λ = Cov(watch, polar) / Var(watch) used to de-attenuate any model in which watch HRV is a regressor. Where HRV is an outcome, non-differential error costs power rather than causing bias, and is documented rather than corrected.
E1 — efficacy. A randomised within-subject crossover with three conditions in a Latin-square order: T1 adaptive neutral sound, heart-rate-driven; T2 an adaptive calm arrangement of a real song the participant chose; and C an active control — a generic "relaxing" playlist the participant already knows. The control is deliberately not silence, because silence changes expectancy and confounds "any audio" with "our audio".
Power, stated in advance. For a paired design at α = .05 two-sided and 80 per cent power: d_z = 0.5 needs n = 34; 0.6 needs 24; 0.65 needs 21; 0.8 needs 15. The Cochrane estimate (−5.7 STAI-S, SD ≈ 10, so d ≈ 0.55) implies n ≈ 28 for self-report, while physiological effects at d ≈ 0.4 need n ≈ 50. So n = 20 is a feasibility signal for A2a and under-powered for A2b, and this dissertation says so before the data rather than after. Two remedies are built in: three sessions per participant, and a Bayesian re-analysis with a meta-analytic prior in which the posterior, not a p-value, drives the gate.
| Threat | Direction | Control |
|---|---|---|
| Expectancy / placebo | Inflates the treatment arms | Active control; expectancy measured pre-session and entered as a covariate; the analyst is blind to condition labels; participants are told all three are "calming audio" |
| Regression to the mean | Inflates any pre→post drop | Baseline as covariate (ANCOVA form); a five-minute pre-session rest so the baseline is not the arrival spike |
| Order, carry-over, learning | Biases later sessions | Latin-square counterbalancing; Order and Session# as covariates; ≥ 24 h washout |
| Time of day / circadian HRV | Adds noise; biases if unbalanced | Fixed slot per participant; covariate |
| Movement, speech, caffeine | Artefacts in heart rate | Seated protocol; accelerometer masking; a logged two-hour rule on caffeine, alcohol and exercise |
| Hawthorne / observation | Equal across arms | Identical observation in every arm |
| Self-selection into the pilot | Limits external, not internal, validity | Within-subject design; recruitment source reported; pilot sample compared with the survey on GAD-2 and PSS-4 |
| Founder-researcher demand effects | Inflates the treatment arms | Sessions run by a research assistant from a script; the founder does not moderate and does not analyse un-blinded |
| Multiple outcomes | False positives | One primary outcome (STAI-S); secondaries reported with Holm correction |
| Small n | Low power for physiology | Bayesian analysis; three sessions per person; explicit feasibility framing |
Transcribed row for row from research/METHODOLOGY_DESIGN.md §4.3.
4.7Ethics, bias controls and limitations
| Bias | Control that is actually implemented |
|---|---|
| Confirmation — the founder wants PROCEED | Thresholds and likelihood ratios pre-registered; the model is built before the data; kill criteria written down |
| Demand effects on participants | A research assistant runs sessions and focus groups from scripts; the founder is absent from the room; the survey is anonymous |
| Analyst degrees of freedom | Analysis code written and run on simulated data first; one primary outcome; deviations logged |
| Selective reporting | Every pre-registered row is reported whether it fires or not; a negative result is a valid dissertation outcome |
| Reflexivity | A researcher diary and a reflexivity section, per Coghlan and Brannick |
Ethics. Wellness framing throughout, with no diagnosis and no medical claim in any interface copy. Consent is taken per data type (live heart rate, HRV, sleep). Data minimisation is architectural rather than promised: beat-level data are processed on device and only windowed features and session summaries leave the phone, and the user can delete everything. EU data are stored in the EU on a consent lawful basis; HealthKit terms forbid advertising use; Oura's terms gate the API behind membership. Fairness is treated as a measurement problem: photoplethysmography accuracy varies with skin tone, motion and temperature, so the quality index and calibration are reported by sub-group and no score is surfaced that would penalise a user for sensor noise. Safety: if resting heart rate exceeds 130 bpm or the user tags a panic state, the app shows a grounding message and hides all numbers, with no clinical routing beyond a general help line.
Limitations stated up front. Non-probability samples; hypothetical willingness to pay; a small pilot; consumer-grade physiology; a single coder on the netnography; the insider role; the independence assumption inside the concept joint probabilities; and subjective likelihood ratios — which are transparent and contestable, and that is exactly their merit.
5Findings
What this chapter establishes. That the need is real, frequent and state-dependent; and that the state dependence runs against the firm's principal asset in the acute moment and with it in the wind-down moment — a result that arrived in Empathise and survived quantification.
surveys/analysis/survey_pipeline.py → results.json, replaceable with fielded data via --csv; and research/focus-groups/SYNTHESIS.md, which the framework document describes as a pre-field template until transcripts exist). They are reported here exactly as the site publishes them, and fieldwork replaces them in the same pass. This is the single largest exposure in the evidence base and is stated here rather than left to be discovered.5.1Empathise: what each evidence stream said
Literature. Music lowers state anxiety, on low-quality trial evidence, with heart rate the strongest physiological marker and cortisol the weakest; slow tempo raises vagal tone; self-report and physiology need not move together; and generative audio has essentially no independent evidence (chapter 2).
Expert interview. Dr Wang, a dementia specialist, gave two findings that shaped the convergence: familiar music from a patient's youth works while purpose-made "treatment music" does not, which undercuts an AI-generated-music play for dementia; and monetisation for elderly care in China is hard.
User interviews (n = 3 acute). All three acute-anxiety interviewees independently rejected emotionally engaging music and asked for the opposite. Their language converges almost word for word:
| Participant | Stated ideal | What they reject, and why |
|---|---|---|
| I-0328, backend engineer | "A completely empty, emotionless sound, like a soft wall, that separates me from the world" | Cannot tolerate lyrical pop when anxious; late-night algorithmic "sad" recommendations deepen the anxiety; nature white noise has jarring sharp sounds (bird calls, high piano) |
| I-0426, account executive | "A sound that flattens my brainwaves — no melody, nothing that makes me picture anything" | Standard sleep playlists are ineffective or counterproductive; the sweet "healing" genre triggers reactance and feels fake; a strong rhythm makes her heart entrain to the beat — the wrong direction |
| I-0535, startup founder | "Something that fills the empty space and brings absolute calm" | Cannot tolerate silence either; sad strings trigger catastrophising, happy pop annoys, alpha-wave audio with a low hum causes physical nausea. Current substitute: an air purifier on maximum |
The recurring complaints about existing neutral audio are the opening: sharp jarring frequencies, low hums that cause nausea, "healing" playlists that feel fake, and anything with a beat. These are not vague preferences; they are an engineering specification, and they became the survey's nine-item acoustic-tolerance matrix and the focus-group stimulus specs.
Netnography (n = 200 rows). Both preferences coexist: some reviewers fault Endel and Brain.fm for being "just lofi — not the song I expected", others explicitly value the absence of lyrics ("Endel feels less distracting because no hooks or lyrics"). The correct reading is therefore not "neutral sound wins" but "there are two distinct jobs". Separately, all ten rows coded for attitudes to biometrics concern the reliability of the reading and none concerns behaviour change — the action gap.
Competitor and market scan. As chapter 3.
5.2Define: the territory matrix, the artefacts, and the core tension
Three candidate territories entered Define — dementia, ADHD and anxiety — scored on six weighted criteria (clinical evidence .25, reachable market .20, willingness to pay .20, asset fit .15, regulatory ease .10, wearable synergy .10). Version 1 produced point totals of 1.80 (dementia), 2.65 (ADHD) and 4.55 (anxiety). Version 2 turned every cell into a triangular low/mode/high distribution, perturbed the weights by ±40 per cent and re-normalised across 5,000 Monte-Carlo draws, and revised anxiety's asset fit from 4 to 3 and wearable synergy from 5 to 4 after the interview evidence and the wearable-API facts. Anxiety's mode total fell from 4.55 to 4.30 and it still ranked first in more than 99 per cent of draws.
| Criterion (weight) | Dementia | ADHD | Anxiety v1 | Anxiety v2 |
|---|---|---|---|---|
| Clinical evidence (.25) | 2 / 3 / 4 | 1 / 2 / 3 | 4 / 5 / 5 | 4 / 5 / 5 |
| Reachable market (.20) | 1 / 1 / 2 | 2 / 3 / 4 | 4 / 5 / 5 | 4 / 5 / 5 |
| Willingness to pay (.20) | 1 / 1 / 2 | 2 / 3 / 4 | 3 / 4 / 5 | 3 / 4 / 5 |
| Asset fit (.15) | 1 / 1 / 2 | 2 / 3 / 4 | 3 / 4 / 5 | 2 / 3 / 4 |
| Regulatory ease (.10) | 2 / 3 / 4 | 1 / 2 / 3 | 3 / 4 / 5 | 3 / 4 / 5 |
| Wearable synergy (.10) | 1 / 2 / 3 | 2 / 3 / 4 | 4 / 5 / 5 | 3 / 4 / 5 |
| Mode total | 1.80 | 2.65 | 4.55 | 4.30 |
research/DECISION_MODEL.md §6. Cells are low / mode / high on a 0–5 scale. Sleep, chronic pain and depression entered the funnel as extras but were not scored in full for want of evidence — stated so the reader knows the funnel's real width.
The evidence supporting the two dementia and ADHD downgrades is worth stating because it is where the funnel actually narrowed: in dementia, a 2017 meta-analysis found a medium effect of individualised music on agitation (d = 0.61 across 12 trials) but the largest pragmatic trial to date (N = 463 residents, 54 nursing homes) found no significant effect on the primary agitation outcome; in ADHD, generic background music showed no effect on core attentional networks, while the more specific amplitude-modulation result used proprietary Brain.fm stimuli, and a 2024 systematic review found zero studies of brown noise meeting inclusion criteria.
Personas, JTBD, POV and HMW. These artefacts exist and are linked in Appendix A, but the framework document writes them as illustrative examples rather than as outputs of coding, so they are reported here as design artefacts and not quoted as findings.
5.3Survey results (as published)
| Block | Result as published | What it settles |
|---|---|---|
| B · Need, timing, coping | Timing: bed 57%, work 52%, evening 33%, night 23%. "Hired today": music 61%, scrolling 57%. | Two dominant moments — bed and desk — consistent with the interview and focus-group maps. |
| B · Outcome-driven innovation | "Stop racing thoughts" 6.4 — the only outcome scoring ≥ 6. Then calm-in-minutes 5.8, not-make-it-worse 5.6, know-it-is-working 5.5. | One under-served outcome, not a list. The product's primary job is intrusive thought, not sleep. |
| C3 · Content by state | Acute scenario: real song 24% [17–31%] versus neutral 46%. Wind-down: real song 47% [39–56%] versus neutral 19%. McNemar χ² = 14.29, p < .001 (b = 16, c = 47). In the GAD-2-positive acute segment: real song 16%, neutral 49%. | The central result. State dependence is real, large, and sharper in the more anxious segment. A1a × 0.33; A1b unchanged. |
| C4 · Acoustic tolerances | Sharp −1.5, beat −1.1, lyrics −0.8, hum −0.7; very slow tempo +0.9, adaptive sound +0.8. | Converts the interview complaints into a signed engineering spec. |
| C5 · Value of real artists | Top-2 box 35%. | Below the 40% threshold pre-registered for an A1b bump — real artists matter to a minority, and mainly for wind-down. |
| D · Wearable | Owners who would connect 57% [45–69%]; value of a measured result 54%; trust in a stress/HRV score 3.06 / 5; privacy-concerned 18%. Action gap: after a high reading, nothing 38%, breathe 30%, walk 18%, more anxious 18%. | A3 × 2.92 in combination with Kano and E4. The action gap is the product opportunity, quantified. |
| E · Acceptance and form | TAM perceived usefulness 4.9, ease of use 5.2, behavioural intention 4.9 (of 7); BI top-2 25% overall and 32% among owners; βPU 0.52, βPEOU 0.34, R² 0.64; α .68 / .68 / .75. Kano: K1 attractive, K2 one-dimensional (must-be in the GAD-2-positive segment), K3 and K4 indifferent. Preferred form: wearable app 45%, standalone 32%, artist album 17%. Relative advantage top-2 60%. Leading barrier: "another subscription" 49%. | Acceptance is moderate, not enthusiastic; the measured result is the feature that changes category from attractive to expected in the anxious segment. |
| F · Price | Van Westendorp US/EU: PMC 5.19, OPP 6.28, IPP 7.37, PME 8.46 (n = 80). China: PMC 14.16, OPP 15.47, IPP 21.38, PME 28.6 (n = 41). Purchase intent at $6.99 top-2 35%, 18% after the ×0.5 discount; China at ¥30, 15%. | The $6.99 anchor sits inside the acceptable range for US/EU and outside comfortable intent for China — the overseas-first decision, quantified. |
All figures as published on site/pages/survey-results.html from surveys/analysis/results.json. Combined survey likelihood ratios applied to the belief register: A1a × 0.33 · A1b × 1.0 · A3 × 2.92 · A4 × 2.4.
5.4Focus groups (as published)
Two groups, 13 participants: FG-1 online, 82 minutes, seven participants from the US, UK and Germany (Apple Watch 4, Oura 2, none 1); FG-2 offline in a Shanghai co-working space, 78 minutes, six mainland-China office and hybrid workers. Both were moderated from a script by a research assistant with the founder absent. A blind stimulus test preceded discussion, with private ratings taken before anyone spoke.
| Clip | "Would help me settle" (1–5) | "Would irritate me" | Fits acute-at-night (n) | Fits daytime wind-down (n) | Neither (n) |
|---|---|---|---|---|---|
| S-A calm real song60–66 bpm, no percussion, −14 LUFS | 3.4 | 2.1 | 3 | 9 | 1 |
| S-B neutral adaptive soundno pitch centre or rhythm; centroid 1.2 kHz → 600 Hz; nothing below 60 Hz; transients ≤ +6 dB | 3.9 | 1.7 | 10 | 2 | 1 |
| S-C guided voice over a pad | 2.6 | 3.2 | 1 | 4 | 8 |
Eight themes emerged, of which five carry into the design. T1 Five moments, not one "anxiety" (7/7 and 6/6): rumination after the day, sleep-onset, anticipatory or performance anxiety, notification spikes, and background tension — "by 1 am I'm not anxious about work any more, I'm anxious that I'm still awake"; "I don't need to sleep, I need to not shake for 20 minutes". T2 State-dependent sound (6/7 and 5/6): melody "gives my brain something to chase"; lyrics "start a story"; the Chinese phrase for "a soft wall" surfaced unprompted in FG-2, the same metaphor as interview I-03. Two participants rejected the neutral clip as "hospital air-conditioning" — disconfirming, and recorded, because it makes engineering quality decisive rather than optional. T3 The number cuts both ways (5/7 and 4/6): owners valued a result — "92 → 68, that I would screenshot" — but four said a live heart-rate line would make it "a test I can fail". T4 One tap or nothing (6/7 and 6/6): unanimous. T7 The action gap: owners read stress scores "like the weather", informative but not actionable.
The dot vote (three dots each, 13 participants) ranked: adapts to me without asking (14), shows me the after-result (11), works from the watch or ring with one tap (9), neutral sound engineered clean (3), real songs by artists I like (2). Shrunk likelihood ratios entering the register, at k = .5 with a further ×.6 in log space for correlation with the interviews: A1a 0.6, A1b 1.4, A3 1.3, A4 0.9.
5.5The belief register: posteriors, gates and the verdict
How A1a fell from 0.50 to 0.08. Prior odds 1.00, then in order: the three acute interviews at 0.35 (shrunk from 0.3 raw, per §4.5) → 0.350; netnography, uninformative at 1.00 → 0.350; forum and Product Hunt users valuing "no hooks or lyrics" at 0.80 → 0.280; Harney's non-significant personalisation trend at 1.10 → 0.308; beat-entrainment plausibility at 0.80 → 0.246; and finally the survey's acute-scenario row at 0.33 → 0.081, giving P ≈ 0.075, reported as 8 per cent. Notice that the qualitative evidence alone took A1a only to about 0.20; it was the pre-registered survey row that crossed the pivot line. That is the shrinkage rule doing exactly what it was designed to do — qualitative evidence can move a belief, but it cannot by itself cross a gate.
The verdict is NOT YET, with the content pivot already triggered and acted on: the acute product becomes adaptive neutral sound, and the real-music catalogue becomes the wind-down mode and the C3 probe. A3 and A4 clear their gates. What remains is the efficacy pilot, which the A2 gates require by construction, and the legal check.
6Concept and business model
What this chapter establishes. How the state-dependence finding turns into a two-mode product and a business model whose cells are each traceable to a source, and where that model's arithmetic stops for want of data.
6.1Value Proposition Canvas
- Stop the racing thoughts — the only outcome scoring ≥ 6 on the opportunity index (6.4). This, not sleep, is the job.
- Get calm within minutes (5.8), without making it worse (5.6), and know it is working (5.5).
- Five separable moments, both groups: rumination after the day; sleep-onset; anticipatory or performance; notification spike; background tension.
- Do it half-asleep at 1 am with one tap, having chosen no mood, duration or genre.
- Melody "gives my brain something to chase"; lyrics "start a story" — both make the acute state worse.
- Sharp frequencies, a low hum that causes nausea, "healing" music that feels fake and triggers reactance, and any beat the heart can entrain to.
- A live number turns calming into "a test I can fail" (four owners across the two groups).
- "Another subscription" — the leading barrier, 49%.
- The action gap: after a high stress reading, 38% do nothing at all.
- A result worth keeping — "92 → 68, that I would screenshot".
- Something that adapts without asking (the top dot-vote item, 14 of 39 dots).
- Credibility without medical framing: "prove it, but don't call it therapy". Settle and unwind were welcomed; treatment and clinical were resisted.
- For wind-down only, artist identity adds meaning — "if it's their version I'd listen on the train home".
- Settle — an engineered neutral bed for the acute moment, adapting to live heart rate.
- Unwind — real, artist-linked music for wind-down, selected by tag, moment and heart-rate band, skippable at any time.
- A post-session result, a session history, and export or delete.
- Research mode for the pilot: participant code, Latin-square condition, STAI-6 and expectancy items.
- The neutral bed is specified against the complaints, not against a genre: no melody, beat, lyrics or pitch centre; spectral centroid 1.2 kHz falling to 600 Hz; nothing below 60 Hz; transients ≤ +6 dB; −16 LUFS; parameters step every 60 seconds.
- Silent adaptation, numbers after, and hide-numbers as the default in the acute mode.
- Three taps and ten seconds to sound; one action per screen.
- A signal-quality badge, and no number shown below a quality index of 0.70.
- A safety screen at resting heart rate above 130 bpm or on a panic tag: grounding message, numbers hidden.
- The sentence the user reads afterwards is the residual, not the raw number: "calmer than your usual 22:00 by 9 bpm".
- "Learning your usual, n of 14" as the baseline model fills — the shipped build shows n of 3 from session history until the 14-day model exists.
- Owned and licensed real music for Unwind, with a one-line "why this track".
- Like, Slower, Next and Shuffle — feedback that is also training signal.
6.2Business Model Canvas
- 220+ digital service providers, with direct Spotify, Apple and YouTube deals
- Apple — HealthKit, App Review, TestFlight (Team 8CP86GZ58V)
- Spotify — Web Playback SDK with PKCE auth, Premium required
- Oura — API v2, membership-gated
- Artists and rights holders
- A university lab for a proper RCT: to verify — named as an intention in the framework, no agreement exists
- Demand signal → AI production → measured library
- Running the pre-registered evidence pipeline (survey, focus groups, E0/E1)
- Rights clearance per track (masters and publishing separately)
- App Store compliance and release engineering
- Signal processing: cleaning, baseline, residualisation
- Catalogue and rights (~106k reserve, ~85% owned or direct, 2022 figures)
- The AI production pipeline
- The published recommender (Lu et al., 2021, PeerJ CS)
- 9 owned tracks ingested at −16 LUFS AAC; 69 Unwind tracks of which 53 slow piano; 4 named Settle beds
- The Lilt codebase: 70 unit tests, 11 end-to-end, 6 audio, 14 scene-engine
- Sound that adapts to your heart rate and then shows you what changed — the top two dot-vote items in both focus groups.
- Two modes, one product: an engineered neutral bed for the acute moment; real, artist-linked music for wind-down.
- Wellness, not therapy — "prove it, but don't call it therapy".
- One tap, works half-asleep — three taps and ten seconds to sound.
- Self-serve, no account — a decision taken explicitly in the App Review reply for this version
- Data stays on the phone; progressive consent per data type; user-initiated export and delete
- Copy tone: second person, short, caring, no imperatives
- App Store — Lilt Adaptive Sound 1.0, build 4, submitted 3 Sep, resubmitted 4 Sep, status waiting for review
- Installable PWA, currently v0.5.1
- Overseas-first (US/EU); China as a WeChat mini-program probe
- For C3: the firm's own 220+ DSP distribution
- For C2: an SDK or catalogue licence into a partner device
- Primary: hybrid-work professionals 28–40 with mild-to-moderate anxiety who own a smartwatch and already self-medicate with music
- Sharpest sub-segment: GAD-2-positive respondents (38%, n = 51), for whom the measured result is a must-be feature rather than an attractive one, and who choose neutral sound in the acute scenario 49% to 16%
- Out of scope: diagnosed severe generalised anxiety disorder — that belongs to clinical care
- C2 segment: wearable and hardware brands. C3 segment: streaming listeners
- Design-thinking phase: $10–20k, ≤ 3 months, ring-fenced
- Landing-page ad spend for the smoke test: ¥500–1,000
- Production cost per calm variant: to verify
- Customer acquisition cost: to verify
- Rights and royalty cost per stream: to verify
- App Store / Play Store commission: to verify — 1.0 is priced free with no in-app purchase configured
- B2C subscription at a $6.99/month test price — inside the US/EU Van Westendorp range (PMC 5.19, OPP 6.28, IPP 7.37, PME 8.46); discounted purchase intent 18%
- China probe at ¥30/month — intent 15%, against Tide's ¥218/year anchor
- C2 licence fee to a hardware or platform partner: to verify — no term sheet exists; the pre-registered evidence is "1–2 positive partner conversations"
- C3 streaming royalties on artist calm content: to verify — no per-stream rate in the repository
Business Model Canvas: 9 blocks, 42 cells, 35 sourced and 7 marked "to verify". Sources: COMPANY_BACKGROUND §1–2; DESIGN_THINKING_FRAMEWORK §0, framing update and §6; app/ARCHITECTURE.md §2, §4, §9; app/RECOMMEND_BRIEF.md; ML_RESEARCH_DESIGN §6; survey results as published; focus-groups SYNTHESIS T3–T6 and the dot vote; TODO §6–16.
6.3Marketing mix (4P) for Lilt
| P | Decision and evidence | What is still open |
|---|---|---|
| Product | Two modes, segmented by arousal state. Settle: an engineered neutral bed built to the S-B specification, adapting silently to live heart rate. Unwind: real music — 69 tracks of which 53 slow piano, plus nine owned tracks — recommended by tag, moment and heart-rate band, skippable at any time. Both end with a result sentence rather than a live readout. The choice of mode is itself the moment tag, so no separate onboarding question is needed. | Whether the AI engine can generate the four Settle beds to specification is one of the four open decisions in the product blueprint. Human listening QA gates any generated variant before it can enter the acute mode — the "hospital air-conditioning" risk, named by two focus-group participants. |
| Price | $6.99 per month as the test anchor. The Van Westendorp analysis for US/EU (n = 80) gives a point of marginal cheapness of $5.19, an optimal price point of $6.28, an indifference price of $7.37 and a point of marginal expensiveness of $8.46 — the anchor sits inside the acceptable range and just above the optimal point. Purchase intent at that price is 35% top-2, 18% after the ×0.5 hypothetical-bias discount. For China the curve is far lower in dollar terms (OPP ¥15.47, IPP ¥21.38) and intent at ¥30 is 15%, against Tide's ¥218/year. | Store commission is unmodelled (to verify); 1.0 ships free with no in-app purchase configured. Whether a bundled or employer-paid route beats standalone pricing is an open question the focus groups pushed hard: a majority in both groups preferred the feature inside something they already pay for. |
| Place | Overseas first. US and EU for willingness to pay and cleaner claims; China as a low-cost WeChat probe. Distribution is the App Store (1.0 submitted, awaiting review) plus an installable PWA that already works today. The survey's preferred form was a wearable app (45%) ahead of standalone (32%) and an artist album (17%). The B2B2C option C2 — licensing adaptive audio into a partner device — is carried as a live alternative, and the firm already distributes into Peloton-type wellness endpoints. | No partner conversation has been recorded; the pre-registered evidence for C2 is "1–2 positive partner conversations" at LR 1.8 / 0.7, and none has been logged. |
| Promotion | Positioning line: "Not meditation. Not white noise. Real music." carried identically across both landing variants, with the A/B split on the headline only: variant A "Calm, from music you actually love" (benefit framing) against variant B "Turn your watch's stress data into calm" (biometric framing). Pre-registered success is a visitor-to-waitlist conversion of ≥ 5%, worth LR 2.5 on both A3 and A4 if met and 0.4 if not. Language discipline follows the focus groups: settle, unwind and the Chinese 缓一缓 were welcomed; treatment and clinical drew resistance in both groups — which is also what the FDA general-wellness test and China's Advertising Law require. | to verify — no conversion figure exists. The hub lists the A/B as running to 20 September; the prototype's own README records that the build was verified locally and deliberately not deployed pending founder sign-off on the positioning. The tagline also predates the content pivot: "Real music" is now the wind-down half of the product only, and variant A's headline is inconsistent with the acute-mode finding. Both need rewriting before the test result can be interpreted. |
Van Westendorp and intent: survey results as published. Tagline and variants: prototype/landing/components/LandingPage.tsx and prototype/landing/README.md; product-design.html §3. Pre-registered landing LR: DECISION_MODEL §4. Language findings: focus-groups SYNTHESIS T5.
6.4Growth options: Ansoff and Three Horizons
6.5Revenue model and unit economics
The revenue model is a consumer subscription with two hedges: a licence into a partner device (C2) and a content line through the firm's own distribution (C3). What follows is the unit-economics arithmetic written out in full, with every input either sourced or flagged. It does not close, and the reason it does not close is itself a finding.
Unit economics: 7 inputs in all — 4 sourced, 3 marked "to verify"; the landing conversion is pre-registered but not yet observed. Retention base rate: Creswell & Goldberg (2025) via SYNTHESIS §6; A6 posterior: DECISION_MODEL §3.
Two structural observations follow from the arithmetic rather than from opinion. First, at a category-normal retention rate a standalone $6.99 subscription cannot support meaningful paid acquisition, which is the same conclusion the focus groups reached from the other direction when both groups preferred the feature bundled inside something they already pay for. Second, the single highest-leverage number in the whole business model is retention (A6), and it is the one number the design-thinking phase is structurally unable to test — it needs an MVP cohort. That is stated as a boundary of the phase, not solved inside it.
6.6Concept viability
| Concept | Joint probability (independence assumed) | Value | Weakest link | Reading |
|---|---|---|---|---|
| C3 — artist "calm" content line through existing distribution | A1b × A4 = 0.80 × 0.82 | 65% | A1b 0.80 | The cheapest probe of real-music calm demand, and the highest-probability concept precisely because it depends on the fewest beliefs. Needs no app, no wearable and no efficacy claim. |
| C1-neutral — adaptive neutral sound, B2C app | (1 − A1a) × A2a × A3 × A4 = 0.92 × 0.85 × 0.72 × 0.82 | 46% | A3 0.72 | The concept the content pivot points to. It does not need A5 at all, which is why the PROCEED gate has an "or the chosen variant is neutral sound" escape. |
| C2 — adaptive-audio SDK licensed to hardware brands | A2a × A4(B2B) × A5 = 0.85 × 0.82 × 0.65 | 45% | A5 0.65 | The hedge against consumer-payment friction: engagement is the partner's problem, not ours. A5 is needed only if real music is included in the licence. |
| C1-real — calm versions of real songs, B2C app | A1b × A2a × A3 × A4 × A5 = 0.80 × 0.85 × 0.72 × 0.82 × 0.65 | 26% | A5 0.65 | The original concept, and now the least likely — not because any single belief collapsed, but because it depends on five of them at once. This is what the multiplication rule is for. |
site/assets/decision.js: C3 is 0.7985 × 0.8181 = 0.653, which is 65% and not the 66% the rounded factors would suggest.The portfolio reading. The three concepts are not alternatives to be ranked once; they are a portfolio with different failure modes. C3 tests demand for real-music calm content at almost no cost and without any of the app's beliefs. C2 hedges the consumer-payment risk that the focus groups flagged and the market data corroborates. C1 is the asset-leverage bet, and the gates select which of them to fund rather than merely whether to fund anything.
7Product and technology
What this chapter establishes. That every visible design decision traces to a specific finding rather than to taste, and that the "ability" claim is demonstrable: the product exists, the algorithm is specified, and its constraints are the wearable APIs' constraints, not aspirations.
7.1From evidence to design
| Finding | Design decision | Principle and where it appears in the product |
|---|---|---|
| Acute states reject melody, lyrics and beat; wind-down accepts real music (interviews, focus groups, survey C3) | Two modes with different content classes, chosen by the user in one tap | The mode choice is the moment tag, so the app never asks a second question. Home screen: two mode cards. |
| A live number makes calming "a test I can fail" (focus-group T3) | Adapt silently; show numbers after; hide-numbers is the default in the acute mode | Weiser and Brown's calm technology — information at the periphery. Session screen has no readout; the result screen has the sentence. |
| Attention narrows under high arousal (Easterbrook, 1959); cognitive load rises (Sweller) | One action per screen; three taps and ten seconds to sound | Hick's law on choice time and Fitts's law on target size; Apple's ≥ 44 pt touch targets are met throughout, verified in the adaptation review. |
| Consumer HRV is noisy and lagged; heart rate is not | The live loop is heart-rate driven; HRV is a before/after outcome only | No number is displayed below a signal-quality index of 0.70, and a quality badge shows why. |
| Users want to know it is working (ODI 5.5) but reject medical framing (T5) | A result sentence in wellness language: "calmer than your usual 22:00 by 9 bpm" | Nielsen's heuristic on system status, expressed as a residual against the user's own baseline rather than as a score. |
| Two focus-group participants called the neutral clip "hospital air-conditioning" | A human listening QA gate before any generated variant can enter the acute mode | Engineering quality is the differentiator, so quality control is a product feature, not a process detail. |
| Privacy concern is a minority but real (18%); HealthKit terms restrict use | Progressive consent per data type; beat-level data processed on device; export and delete in settings | Data minimisation as architecture. Self-determination theory: autonomy is preserved by making every data step optional. |
| Anxiety can spike rather than settle | A safety screen at resting heart rate above 130 bpm for ten seconds, or on a "grounding" tap | Numbers are hidden and a grounding message is shown; no clinical routing beyond a general help line. |
Principles table and screen references: site/pages/product-design.html §1 and §7; safety rule: ML_RESEARCH_DESIGN §6 and app/ARCHITECTURE.md §5. Human-factors sources named in the text are excluded from the reference list where the repository gives no year or venue — see the note under the references.
7.2The app today
The product is real, not a mock-up. The flow is: home (Settle or Unwind, duration, sleep fade) → pre-rest of 60 seconds, or 300 seconds in research mode, with a 0–10 self-rating → session → a 20-second check-in → a result screen → history, export or delete → settings. In Settle, Web Audio synthesises the neutral bed and steps the low-pass cutoff, gain, density and pulse every 60 seconds according to the v0 rule table. In Unwind, playback runs through one transport over the Spotify Web Playback SDK (PKCE auth, Premium required) with an embed fallback, owned files through Web Audio with a 1.5-second crossfade, and SoundCloud; tracks are chosen by tag, moment fit and heart-rate band, and the listener can skip at any time. Research mode adds the participant code, the Latin-square condition, the five-minute pre-rest, STAI-6 before and after, the two expectancy items, and a condition column in the export. A 72-second recording of the live build is on the hub; the heart rate in it is the simulated source.
7.3The adaptation algorithm
| Rule | Value | Why |
|---|---|---|
| Starting pulse equivalent | HR ≥ 90 → 66 bpm; 75–90 → 60 bpm; < 75 → 56 bpm | Entrainment: start near the listener's own rate, then lead it down (Bernardi et al., 2006) |
| Low-pass cutoff | 600 + clamp((HR − 60) / 40, 0, 1) × 1800 Hz | Brightness tracks arousal; the ceiling keeps the bed from sounding thin |
| Every 60 s, if residual HR fell ≥ 2 bpm | Step tempo down 2 bpm; lower brightness one notch | Reward the direction of travel without a discontinuity |
| Every 60 s, if it rose ≥ 3 bpm with no motion flag | Hold; remove any transient layer; density −0.2 | Never chase a rise — the most likely cause is a startle, not a need for more stimulus |
| Guardrails | pulseEq ≥ 50 · cutoff 600–2400 Hz · density 0–1 · gain −6 to 0 dB · −16 LUFS fixed | The loudness is fixed so adaptation can never be heard as "it got louder" |
| Wizard-of-Oz equivalence | A human plays the same table from a script in E1 | The pilot tests the policy, not a particular build |
app/ARCHITECTURE.md §7; research/ML_RESEARCH_DESIGN.md §4.1.
7.4Recommendation and the learning loop
For Unwind the track choice is scored rather than shuffled:
The learning roadmap is deliberately staged and each stage has a guardrail. E2 randomly assigns a variant per session from a small allowed set and fits a hierarchical model of Δr and self-report on variant × moment with partial pooling. E3 replaces the fixed policy with a contextual bandit using Thompson sampling, with the hierarchical posterior as its prior; exploration is capped at 15 per cent of sessions, and any arm whose posterior mean reward is more than 0.3 SD worse than the default is suspended for that user. E4 lets the AI engine propose variants inside a parameter cell, each entering as an arm with a shrunk prior, promoted or retired by posterior thresholds — with a human listening QA gate before anything reaches the acute mode. Every policy change is evaluated offline first with inverse-propensity or doubly-robust estimation on logged data, then online against a hold-out default cohort; never a full rollout of an untested policy.
The reward is a composite by design: −standardised Δr on heart rate plus λ times the change in self-report, with λ starting at 0.5. Optimising on self-report alone is gameable by expectancy; optimising on heart rate alone is gameable by stillness. The composite plus the quality index guards both.
7.5Data, privacy and rights
Data. Consent is taken per data type. Beat-level data are processed on device and only windowed features and session summaries leave the phone. EU data are stored in the EU with consent as the lawful basis. HealthKit terms forbid advertising use of health data. Oura's API requires an active membership for Gen3 and later rings, so any Oura-based feature inherits a paywall the user may not have.
Rights. Catalogue entries carry an explicit rights tier — owned, embed-only or licensed — and an adaptation tier: P for parametric adaptation, permitted only on owned or cleared tracks, and S for selection-level adaptation, which is all an embedded player permits. Embeds are non-commercial and play 30-second previews unless the listener is signed in to Premium. The Settle beds need no catalogue rights at all, which is precisely why the neutral-sound concept survives a negative legal check. Track provenance for the Unwind pool is documented: Spotify editorial calm playlists, then ghost artists, then Wikidata-resolved artist identifiers, then the artists' Popular lists, with every identifier verified through oEmbed.
8Validation plan and decision
What this chapter establishes. The exact rule by which this project will be continued, redirected or stopped, and what each remaining test is worth in the currency the rule is written in.
8.1The stage-gate
8.2What the pilot will settle, and the pipeline that will settle it
| Test | Node | If positive | If null / negative | Cost and time |
|---|---|---|---|---|
| Legal check on ≥ 50 tracks | A5 | 8.0 | 0.15 | 1–2 weeks — the only pending row graded decisive, and the cheapest |
| Efficacy pilot, STAI-S | A2a | 4.0 | 0.30 | 2 weeks; n = 20–30 × 3 sessions |
| Efficacy pilot, HRV / residual HR | A2b | 4.0 | 0.30 | Same sessions |
| Pilot A/B content preference | A1a | 3.0 | 0.33 | Same sessions — the β1 − β2 contrast |
| Landing conversion ≥ 5% | A3, A4 | 2.5 each | 0.40 each | ¥500–1,000 of ads |
| Think-aloud, n = 5–8 | A3 | 1.5 | 0.60 | 1 week |
| Partner conversations | A4 (B2B) | 1.8 | 0.70 | Ongoing |
research/DECISION_MODEL.md §4. Positive for the pilot means p < .05 and d ≥ .3 in the correct direction.
The pipeline exists and has been exercised. Session exports flow through app/tools/export_to_csv.py into per-session and per-window CSVs, then through research/pilot/e1_analysis.py, which fits the pre-registered model, runs the Bayesian re-analysis, prints the likelihood-ratio table and the gate check mechanically, and writes the figures. Twenty-eight tests cover the schema, the exclusion rule, the order inference, the injected-effect direction and the gate arithmetic — including a test that the Latin square in the Python tools is still identical to the one in the app, so that a change to the app's randomisation table fails a test rather than silently mis-assigning conditions.
8.3The legal check and the landing test
The legal check (A5) is first in the queue and this is a deliberate real-options decision rather than a scheduling accident. It takes a little over a week, costs almost nothing, and is the only pending row graded decisive. If it returns negative at LR 0.15, A5 falls from 0.65 to roughly 0.22 and the real-music route closes on the spot — at which point no further engineering time should be spent on parametric adaptation of catalogue tracks, and C1-neutral becomes the only live app concept. Ordering tests by expected information per unit cost is what turns "real options" from a phrase into an execution order.
The landing test (A3, A4) supplies the only behavioural evidence in the entire design — everything else about willingness to pay is stated intent. Its pre-registered threshold is a visitor-to-waitlist conversion of at least 5 per cent per positioning variant. Two caveats attach to it, both stated in §6.3: no result exists yet, and the shared tagline and variant A's headline both predate the content pivot and should be rewritten before the test is read.
8.4Decision memo timing
The sequence is fixed: E0 calibration (25 August – 6 September) → E1 efficacy pilot (1–25 September) → Bayesian update and gate check (25 September – 2 October) → decision memo and supervision #2 (early October). The memo is written mechanically from the report: observed likelihood ratios replace the pending rows in the register with a date and a written reason, the same numbers are mirrored into the interactive page so the document and the site agree, and the gate check prints itself. Design thinking ends with a documented decision, not a pitch — and a documented KILL would be a valid outcome of this dissertation.
9Implementation roadmap and risks
What this chapter establishes. What happens between now and the decision memo, what could go wrong, who owns each risk and what is already in place to mitigate it — and the hard resource boundary the whole phase runs inside.
9.1Roadmap
9.2Risk register
| ID | Risk | Likelihood | Impact | Owner | Mitigation already in place |
|---|---|---|---|---|---|
| R1 | The efficacy pilot fails to clear its gates — A2b is under-powered by design at this sample size | High | High | Founder-researcher & RA-1 | Declared in advance: n = 20 is a feasibility signal for A2a and under-powered for A2b. Three sessions per participant for more within-person information; Bayesian re-analysis with a meta-analytic prior so the posterior rather than a p-value drives the gate; heart rate used as the accurate co-primary physiological outcome. Across 40 synthetic replications A2b fired in only 45% of runs — the expected behaviour, not a surprise. |
| R2 | Adaptation rights cannot be cleared — A5 at 0.65, with the check unfinished | Medium | High | Founder + legal review (provider not named in the repository) | Sequenced first because it is cheapest and most decisive. The neutral-sound concept needs no A5 at all and the Settle beds need no catalogue rights, so a negative result redirects rather than kills. Masters and publishing are confirmed separately; the catalogue already carries per-entry rights and adaptation tiers. |
| R3 | Retention below the category-changing threshold — A6 at 0.25 against a 4.7% D30 base rate | High | High | Product | Carried explicitly as a watch item rather than a gate, because it cannot be tested within the phase. Design response: anchor the habit to hardware the user already checks daily, show a result worth keeping, and let the watch or ring trigger the offer. Requires an MVP cohort to test at all — stated as a boundary of the phase. |
| R4 | Pipeline output mistaken for fieldwork — the survey results file carries a simulated flag and the focus-group synthesis is described as a pre-field template | Medium | High | Founder-researcher | A provenance sentence at every point of use, including chapter 5 of this document; the same statement made out loud at the viva rather than waited for; fielded data replace the figures in one pass, and if the site relabels them this document is relabelled in the same pass. This is the largest single exposure in the evidence base. |
| R5 | Regulatory claim creep — any drift toward "treat anxiety" triggers SaMD status in the US and Advertising Law exposure in China | Low | High | Founder | Wellness framing is architectural: no diagnosis, no medical claim in interface copy, and the focus groups independently confirmed that settle and unwind are welcomed while treatment and clinical are resisted. The DTx route is deferred until post-funding. Pear's Chapter 11 and Akili's $0.434 sale are the cautionary precedents: clearance did not imply viability. |
| R6 | Platform gatekeeping — App Review already issued a Guideline 2.1 information request; the iOS 26 SDK requirement forced a CI change | Medium | Medium | Engineering | Six replies written and filed with a recorded walkthrough; resubmitted 4 September with the content unchanged. A deliberate decision not to add accounts in this version, documented in the reply. CI moved to macOS 26 with Xcode 26.2 for the SDK requirement. Contingency: the installable PWA works today and is not gated by any store. |
| R7 | Measurement error swamps the physiological signal — Apple Watch HRV MAPE ≈ 29%, HealthKit lag ≤ 30 min, Oura sleep-only | High | Medium | Research | The architecture concedes the constraint rather than fighting it: heart rate is the live control signal and the co-primary outcome, HRV is a before/after measure only. E0 calibration with a Polar H10 (n ≈ 8) estimates ICC, MAPE and the reliability ratio λ, used to de-attenuate any model where HRV is a regressor. Sessions below a 0.70 quality index leave the physiological models and are counted, not dropped silently. |
| R8 | Insider bias — the founder is the researcher and wants PROCEED | High | High | Supervisor + RA | Thresholds and likelihood ratios written before the data; the research assistant runs sessions and focus groups from a script with the founder absent; the analyst is blind to condition labels; analysis code was written and run on simulated data first; every pre-registered row is reported whether it fires or not; a researcher diary and a reflexivity section. |
| R9 | Company facts are stale or unreconciled — founding date, catalogue and stream counts, artist-relationship scope | High | Low | Founder | The company file carries its own reconciliation checklist, and this dissertation labels every company figure with its vintage (2022 for catalogue and streams) rather than presenting it as current. The checklist must be closed before submission; until then no company figure is used in an arithmetic result. |
| R10 | The independence assumption inflates or deflates concept viability — A2a, A3 and A4 are plainly correlated | Certain | Medium | Founder-researcher | Stated as a limitation in the methodology and again beneath the viability table; the joint probabilities are used to order the concepts rather than as calibrated forecasts, and the weakest link is named for each concept so the reader can see what actually drives the number. |
| R11 | Platform substitution — Apple ships Vitals free with every Series 9+ watch; Spotify holds a granted patent on inferring emotional state from speech (not from a wearable) and could extend it toward live adaptation | Medium | High | Strategy | The defensible combination is rights plus a measured result plus engineered neutral-sound quality, no one of which is sufficient alone. Note the counter-signal: Spotify has not shipped live biometric adaptation despite years of documented user requests, which reads as much as a warning about demand as an opportunity. |
| R12 | Positioning copy contradicts the finding — the landing tagline and variant A headline both promise real music, which the acute-mode evidence rejects | Certain | Medium | Founder + product | Identified in §6.3 and blocking: the copy must be rewritten to the two-mode proposition before the A/B result can be interpreted, otherwise the conversion figure measures a positioning the product no longer has. The prototype README already blocks deployment pending founder sign-off on the positioning. |
Risks R1–R3 and R5–R7 from the assumption register and the methodology's threats table; R4 and R8–R11 from journey.audit.md §E and METHODOLOGY_DESIGN §7 and §10; R6 and R12 from TODO_2026-08-19.md §14–16 and prototype/landing/README.md. Likelihood and impact ratings are the author's assessment against those sources.
9.3Resources and the budget boundary
The budget boundary is a methodological instrument, not an accounting detail. Ambidexterity requires that the exploratory unit be small enough and separate enough that killing it does not threaten anyone's career or the exploitation P&L; a ring-fenced $10–20k over three months is what makes the KILL branch of the gate credible. It also explains several design choices that would otherwise look like corner-cutting: a Wizard-of-Oz pilot instead of a fully instrumented native app; an n of 20–30 declared under-powered in advance rather than an adequately powered trial; a landing test with a ¥500–1,000 budget; and a legal check on 50 tracks rather than the whole catalogue. Each is the cheapest instrument that can move the belief it targets, which is exactly what the value-of-information ordering in chapter 8 selects for.
What the boundary does not buy is retention evidence, an adequately powered physiological test, or a signed partner conversation. Those are the three things a PROCEED decision would fund next, and naming them here is part of the decision memo rather than an afterthought.
10Conclusion
What this chapter establishes. How far the evidence answers the two research questions, what the project contributes to practice and to the firm, what it cannot claim, and what should be studied next.
10.1Answers to RQ1 and RQ2 so far
RQ1 — is there a real, reachable, under-served need? Yes, with a qualification that changed the product. The need is real and frequent: 87 per cent of respondents need a calming moment at least weekly and 90 per cent already use music or audio to get it, against a global treatment gap in which 27.6 per cent of those needing care receive it. It is reachable: the primary segment self-identifies, self-pays and already owns the sensor. It is under-served in one specific way rather than generally — only one outcome, "stop the racing thoughts", scores above the opportunity threshold, and the wearable owners who already have a stress reading do nothing with it in 38 per cent of cases. And the qualification: the need is state-dependent, and so is the content that serves it. That was hypothesised by three interviews, reproduced in a blind stimulus test, and quantified in a pre-registered survey item with McNemar χ² = 14.29, p < .001.
RQ2 — can TDMusic serve it with its assets, measurably? Partly, and not yet demonstrably. Three of the five ability assumptions clear their pre-registered gates: users will connect a wearable and value the result (0.72), someone will pay at the tested price (0.82), and the rights position is more likely than not (0.65, pending the check). Two do not, and they are the two that matter most: that one session measurably lowers state anxiety (0.85 against a gate of 0.90) and that the same session moves a consumer wearable measure (0.56 against 0.70). Those gates were deliberately set above what the literature alone can reach. The honest answer to RQ2 today is therefore: the firm can build it, can distribute it, can probably clear the rights and can probably sell it — and has not yet shown that it works.
The decision. NOT YET. The content pivot has fired and been acted on; the remaining evidence is the pilot and the legal check; the memo is due in early October.
10.2Contribution to practice and to the firm
To practice. The contribution is a worked, publishable example of a corporate-innovation decision in which the belief state is explicit and re-runnable: eight assumptions, each with a prior and a written rationale; every evidence item converted to a likelihood ratio through a published rubric and shrunk by a published quality rule; thresholds fixed before the data; and the arithmetic exposed rather than summarised. The most instructive single episode is that qualitative evidence took the central belief only from 0.50 to about 0.20, and it was the pre-registered survey row that carried it across the pivot line — a concrete answer to the perennial question of how much three interviews should count.
To the firm. Three things. First, a finding the firm would otherwise have discovered after building: its principal asset is the wrong content for the moment of highest need, and the right response is to segment the product by arousal state rather than to abandon the asset or ignore the finding. Second, a portfolio rather than a bet — C3 tests real-music calm demand at almost no cost, C2 hedges consumer-payment friction, C1 is the asset-leverage play, and the gates select among them. Third, a reusable apparatus: a pre-registered survey pipeline, a pilot pipeline with 28 tests, a live decision engine, and a shipped product with research mode built in. The next vertical costs the firm the fieldwork, not the machinery.
10.3Limitations
- The survey and focus-group figures are pre-registered pipeline output as published on this site, not fieldwork — the results file carries a simulated flag and the focus-group synthesis is described in the framework as a pre-field template. This is stated wherever those figures appear and is the largest single limitation of the evidence base.
- The efficacy pilot has not run. Every efficacy statement in this document is borrowed from meta-analyses conducted in different populations, mostly perioperative, with supervised delivery and GRADE-low trial quality.
- Non-probability samples throughout, with hypothetical rather than revealed willingness to pay; the only behavioural instrument, the landing test, has no result.
- Consumer-grade physiology. Apple Watch HRV fails equivalence against a chest strap; the live loop is therefore heart-rate-driven, which measures autonomic arousal and is blind to its cause.
- Insider role. The founder is the researcher; the controls are declared and implemented, but the conflict is structural rather than removable.
- Subjective likelihood ratios and an independence assumption. The likelihood ratios are judgements made against a rubric; the concept joint probabilities multiply correlated beliefs. Both are transparent and contestable — which is their merit, not their defence.
- Company figures are dated. Catalogue and stream counts are 2022; the founding date and the artist-relationship scope are unreconciled in the company's own file.
- Single coder on the netnography, and the underlying netnography summary was not re-read for this document — only the row count and the patterns already carried in the decision model and the interview synthesis are used.
- Market sizing does not close. Two inputs required for a serviceable obtainable market do not exist in the evidence base, so no SOM point estimate is published.
10.4Further research
- Run the comparison the literature has not run. A properly powered trial of engineered neutral sound against personally selected real music, in acute anxiety, with state anxiety as the primary outcome and residual heart rate as a co-primary. The E1 contrast β1 − β2 is a first, under-powered look at exactly this question; nothing in the published literature answers it.
- Test whether closed-loop adaptation adds anything over a fixed calming stimulus. The strong biofeedback evidence is for breathing-paced protocols, not for music-driven adaptation; the generalisation gap should be named and then closed with an adaptive-versus-fixed arm.
- Study retention as the dependent variable, not as a footnote. At a 4.7 per cent thirty-day base rate, retention dominates the unit economics of the entire category; whether anchoring a habit to hardware the user already checks daily changes that rate is an empirical question with a clear design.
- Validate the residual as a user-facing quantity. "Calmer than your usual 22:00 by 9 bpm" is simultaneously an outcome, a reward signal and a piece of interface copy. Whether users read it as credible, and whether reading it changes behaviour, is a distinct study from whether the underlying quantity moves.
- Extend the decision apparatus to a second vertical. The strongest test of the methodological contribution is whether the same register, rubric and gates produce a defensible decision somewhere the answer is not already suspected.
References
What this list contains. Every work cited by author and date in the text, plus the primary sources behind the figures reported in chapters 2, 3 and 5. Only works already cited in this project's own documents appear here; no new literature has been introduced for this page.
site/pages/dissertation.references-todo.md. Frameworks and laws named in the text for which the repository gives no year or venue are deliberately excluded from this list rather than given fabricated detail: the Ansoff matrix, McKinsey's Three Horizons, O'Reilly and Tushman's ambidexterity, Sweller's cognitive load theory, Hick's law, Fitts's law, Nielsen's heuristics, Deci and Ryan's self-determination theory, and Efraimidis–Spirakis weighted sampling.- Above Avalon. (2025). Apple Watch [Analyst notes]. https://www.aboveavalon.com/notes/tag/Apple+Watch
- Advertising Law of the People's Republic of China (2015 revision), arts. 17–18, 58–59. https://extranet.who.int/fctcapps/sites/default/files/2024-02/341_%E4%B8%AD%E5%8D%8E%E4%BA%BA%E6%B0%91%E5%85%B1%E5%92%8C%E5%9B%BD%E5%B9%BF%E5%91%8A%E6%B3%95.pdf
- Agrawal, S., & Goyal, N. (2013). Thompson sampling for contextual bandits. [Proceedings, volume and pages to verify]
- Apple Inc. (n.d.-a). HealthKit [Developer documentation]. https://developer.apple.com/documentation/healthkit [Publication or retrieval date to verify]
- Apple Inc. (n.d.-b). Vitals on Apple Watch [Support documentation]. https://support.apple.com/guide/watch/vitals-apd15aa7ed96/watchos [Publication or retrieval date to verify]
- Apple Watch Series 9 and Ultra 2 heart-rate-variability validation study. (2024). Sensors. [Authors, exact title, volume and article number to verify] https://pmc.ncbi.nlm.nih.gov/articles/PMC11478500/
- Behavioral Health Business. (2024, November 20). Headspace axes 13% of workforce, transitions therapist network to part-time and contract roles. https://bhbusiness.com/2024/11/20/headspace-axes-13-of-workforce-transition-therapist-network-to-part-time-and-contract-roles/
- Berger, C., [remaining authors to verify]. (1993). Better/Worse coefficients for Kano categories. [Exact title, journal, volume and pages to verify]
- Bernardi, L., Porta, C., & Sleight, P. (2006). Cardiovascular, cerebrovascular and respiratory changes induced by music. Heart, 92, 445. [Issue and page range to verify; the repository also records this work as Circulation, 114(24), 2571–2577 — the two records conflict and must be reconciled] https://www.ahajournals.org/doi/full/10.1161/CIRCULATIONAHA.108.806174
- Blakeley, R. (2024, January 18). Welcome to the sound wellness revolution: Endel's AI-generated soundscapes and the commodification of passive listening. Musicology Now. https://musicologynow.org/welcome-to-the-sound-wellness-revolution-endels-ai-generated-soundscapes-and-the-commodification-of-passive-listening/
- Bland, D. J., & Osterwalder, A. (2019). Testing business ideas. [Publisher and place to verify]
- Bloomberg. (2026, January 5). Smart rings poised for 2026 growth, Oura set to lead. https://www.bloomberg.com/news/articles/2026-01-05/smart-rings-poised-for-2026-growth-oura-set-to-lead
- Bowman, E. H., & Hurry, D. (1993). Real options and organisational learning. [Exact title, journal, volume and pages to verify]
- Bradt, J., Dileo, C., & Shim, M. (2013). Music interventions for preoperative anxiety. Cochrane Database of Systematic Reviews, CD006908. [Issue number to verify] https://doi.org/10.1002/14651858.CD006908.pub2
- Braun, V., & Clarke, V. (2006). [Title to verify — the repository records only the journal, volume and start page]. Qualitative Research in Psychology, 3, 77. [Issue and end page to verify]
- Brown, T. (2008). Design thinking. [Journal, volume, issue and pages to verify]
- Bundesinstitut für Arzneimittel und Medizinprodukte. (2026). DiGA-Verzeichnis [Directory of digital health applications]. https://www.diga-verzeichnis.de/en
- Business of Apps. (2026a). Calm revenue and usage statistics. https://www.businessofapps.com/data/calm-statistics/
- Business of Apps. (2026b). Headspace revenue and usage statistics. https://www.businessofapps.com/data/headspace-statistics/
- Calm Health. (2026). Calm Health [Company website]. https://health.calm.com/
- Carroll, R. J., [remaining authors to verify]. (2006). Measurement error in nonlinear models. [Edition, publisher and place to verify]
- Christensen, C. M., [remaining authors to verify]. (2016). Jobs to be done: Current alternatives as the competitive set. [Exact title and publication details to verify]
- Coghlan, D., & Brannick, T. (2019). Doing action research in your own organization. [Edition, publisher and place to verify]
- Cohen, S., & Williamson, G. (1988). Perceived stress in a probability sample of the United States. In S. Spacapan & S. Oskamp (Eds.), The social psychology of health. Sage. [Page range and place of publication to verify]
- Creswell, J. D., & Goldberg, S. B. (2025). Meditation-app engagement and retention. [Exact title, journal, volume, issue and pages to verify] https://pmc.ncbi.nlm.nih.gov/articles/PMC12333550/
- Creswell, J. W., & Plano Clark, V. L. (2018). Designing and conducting mixed methods research. [Edition, publisher and place to verify]
- Davis, F. D. (1989). Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Quarterly, 13(3), 319–340.
- de Witte, M., Spruit, A., van Hooren, S., Moonen, X., & Stams, G.-J. (2020). Effects of music interventions on stress-related outcomes. [Subtitle to verify — the repository records the title without one] Health Psychology Review, 14(2), 294–324. https://doi.org/10.1080/17437199.2019.1627897
- Devilly, G. J., & Borkovec, T. D. (2000). Credibility and expectancy questionnaire. Journal of Behavior Therapy and Experimental Psychiatry, 31, 73. [Exact title, issue and end page to verify]
- Dudík, M., Langford, J., & Li, L. (2011). Doubly robust policy evaluation and learning. [Proceedings, volume and pages to verify]
- Easterbrook, J. A. (1959). The effect of emotion on cue utilisation and the organisation of behaviour. [Journal, volume, issue and pages to verify]
- Fairfield, T., & Charman, A. E. (2017). Explicit Bayesian analysis for process tracing: Guidelines, opportunities, and caveats. Political Analysis, 25(3), 363–380.
- Federal Trade Commission. (2024, April). Updated FTC Health Breach Notification Rule puts new provisions in place to protect users of health apps [Business guidance blog]. https://www.ftc.gov/business-guidance/blog/2024/04/updated-ftc-health-breach-notification-rule-puts-new-provisions-place-protect-users-health-apps
- Fierce Biotech. (2023, April). Prescription app developer Pear Therapeutics files for bankruptcy, lays off staff. https://www.fiercebiotech.com/medtech/cut-core-prescription-app-developer-pear-therapeutics-files-bankruptcy-lays-staff
- Flanagan, J. C. (1954). The critical incident technique. Psychological Bulletin, 51(4), 327–358.
- Fortune Business Insights. (2025). Mental health apps market size, share & global report. https://www.fortunebusinessinsights.com/mental-health-apps-market-109012
- Gelman, A., & Hill, J. (2007). Data analysis using regression and multilevel/hierarchical models. [Publisher and place to verify]
- Goessl, V. C., Curtiss, J. E., & Hofmann, S. G. (2017). The effect of heart rate variability biofeedback training on stress and anxiety: A meta-analysis. Psychological Medicine, 47(15). [Page range to verify] https://www.cambridge.org/core/journals/psychological-medicine/article/effect-of-heart-rate-variability-biofeedback-training-on-stress-and-anxiety-a-metaanalysis/A839E9C968E54774DF5C8FB186764EF0
- Grand View Research. (2025). Meditation management apps market report. https://www.grandviewresearch.com/industry-analysis/meditation-management-apps-market-report
- Guyatt, G. H., [remaining authors to verify] (GRADE Working Group). (2008). [Title to verify]. BMJ, 336, 924. [Issue and end page to verify]
- Harney, C., Johnson, J., Bailes, F., & Havelka, J. (2023). [Title to verify]. Musicae Scientiae. [Volume, issue and pages to verify] https://doi.org/10.1177/10298649211046979
- Haruvi, A., Kopito, R., Brande-Eilat, N., Kalev, S., Kay, E., & Furman, D. (2022). [Title to verify]. Frontiers in Computational Neuroscience, 15, 760561. https://doi.org/10.3389/fncom.2021.760561
- Healthcare Brew. (2025, April 4). Digital therapeutics Medicare coverage test begins. https://www.healthcare-brew.com/stories/2025/04/04/digital-therapeutics-medicare-coverage-test-begins
- Healthcare Dive. (2021, August). Headspace, Ginger to merge, creating $3B mental health company. https://www.healthcaredive.com/news/headspace-ginger-to-merge-create-3b-mental-health-company/605632/
- Hernando, D., Roca, S., Sancho, J., Alesanco, Á., & Bailón, R. (2018). [Title to verify]. Sensors, 18. [Issue and article number to verify] https://pmc.ncbi.nlm.nih.gov/articles/PMC6111985/
- Huang, Y., [remaining authors to verify]. (2019). Prevalence of mental disorders in China: A cross-sectional epidemiological study. The Lancet Psychiatry, 6(3), 211–224. https://pubmed.ncbi.nlm.nih.gov/30792114/
- Humphreys, M., & Jacobs, A. M. (2015). Mixing methods: A Bayesian approach. American Political Science Review, 109(4), 653–673.
- iiMedia Research. (2023). 2023–2024 China sleep economy industry development and consumer demand research report. https://www.iimedia.cn/c400/92947.html
- iiMedia Research. (2025). 2025–2029 China emotion economy consumption trends insight report. https://www.iimedia.cn/c400/106827.html
- Kano, N., Seraku, N., Takahashi, F., & Tsuji, S. (1984). Attractive quality and must-be quality. Journal of the Japanese Society for Quality Control, 14(2), 39–48.
- Kirk, R., & Timmers, R. (2025). Characterizing music for sleep: A comparison of Spotify playlists. Musicae Scientiae. [Volume, issue and pages to verify] https://doi.org/10.1177/10298649241269011
- Kozinets, R. V. (2020). Netnography. [Edition, publisher and place to verify]
- Kroenke, K., Spitzer, R. L., Williams, J. B., Monahan, P. O., & Löwe, B. (2007). [Exact title to verify — the GAD-2 validation paper]. Annals of Internal Medicine, 146, 317–325. [Issue number to verify]
- Krueger, R. A., & Casey, M. A. (2015). Focus groups. [Edition, publisher and place to verify]
- Lassner, [initials and remaining authors to verify]. (2025). [Title to verify]. BJPsych Open, 11(1), e4. [Publication year recorded in the repository as 2024/2025 — to reconcile] https://pmc.ncbi.nlm.nih.gov/articles/PMC11733488/
- Latka. (2026). Calm revenue. https://getlatka.com/blog/calm-revenue
- Liedtka, J. (2018). On the effectiveness of design thinking. Journal of Product Innovation Management. [Exact title, volume, issue and pages to verify]
- Lipponen, J. A., & Tarvainen, M. P. (2019). RR-interval artefact correction. Journal of Medical Engineering & Technology, 43, 173. [Exact title, issue and end page to verify]
- Lu, [given name to verify], [remaining authors to verify]. (2021). [Title to verify]. PeerJ Computer Science, 7, e716. [The repository records the byline only as “Lu, Qiao et al.”; whether “Lu Qiao” is one author or two must be reconciled from the article]
- Marteau, T. M., & Bekker, H. (1992). A six-item short form of the state scale of the Spielberger State–Trait Anxiety Inventory. British Journal of Clinical Psychology, 31, 301. [Exact title, issue and end page to verify]
- McCreedy, E. M., [remaining authors to verify]. (2021). [Title to verify]. Journal of the American Medical Directors Association. [Volume, issue and pages to verify] https://www.jamda.com/article/S1525-8610(21)01104-X/abstract
- McGrath, R. G. (1999). [Title to verify]. Academy of Management Review, 24, 13. [Issue and end page to verify]
- McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions. [Journal, volume, issue and pages to verify; the repository records the title without its final words]
- Mendes, [initials to verify], de Paula, [initials to verify], & Miranda, [initials to verify]. (2024). [Title to verify]. Interactive Journal of Medical Research. [Volume, issue and article number to verify] https://pmc.ncbi.nlm.nih.gov/articles/PMC11294770/
- Mitrovic, [initials to verify], & Paladin, [initials to verify]. (2026). [Title to verify]. Frontiers in Cardiovascular Medicine, 13. https://doi.org/10.3389/fcvm.2026.1841349
- Morgan, D. L. (1997). Focus groups as qualitative research. [Edition, series, publisher and place to verify]
- Morwitz, V. G., Steckel, J. H., & Gupta, A. (2007). When do purchase intentions predict sales? International Journal of Forecasting, 23, 347–364. [Issue to verify]
- Music Business Worldwide. (2023, May). Universal inks deal with generative AI startup Endel to create "AI-powered, artist-driven functional music". https://www.musicbusinessworldwide.com/universal-inks-deal-with-generative-ai-startup-endel-to-create-ai-powered-artist-driven-functional-music/
- National Institute for Health and Care Excellence. (2023). Digitally enabled therapies for adults with anxiety disorders: Early value assessment (HTE9). https://www.nice.org.uk/guidance/hte9
- National Institute of Mental Health. (n.d.). Any anxiety disorder. https://www.nimh.nih.gov/health/statistics/any-anxiety-disorder [Publication or retrieval date to verify]
- NetEase Cloud Music Inc. (2025, February 20). NetEase Cloud Music Inc. reports fiscal year 2024 financial results [Press release]. PR Newswire. https://www.prnewswire.com/news-releases/netease-cloud-music-inc-reports-fiscal-year-2024-financial-results-302381265.html
- Nigg, J. T., [remaining authors to verify]. (2024). [Title to verify]. [Journal, volume, issue and pages to verify] https://pubmed.ncbi.nlm.nih.gov/38428577/
- O'Daffer, A., Colt, S. F., Wasil, A. R., & Lau, N. (2022). Efficacy and conflicts of interest in RCTs evaluating Headspace and Calm apps: Systematic review. JMIR Mental Health, 9(9), e40924. https://pmc.ncbi.nlm.nih.gov/articles/PMC9533203/
- Oura Health. (2024). The Oura API [Support documentation]. https://support.ouraring.com/hc/en-us/articles/4415266939155-The-Oura-API
- Panteleeva, Y., Ceschi, G., Glowinski, D., Courvoisier, D. S., & Grandjean, D. (2018). [Title to verify]. Psychology of Music, 46(4), 473–487. https://doi.org/10.1177/0305735617712424
- Pedersen, S. K. A., Andersen, P. N., Lugo, R. G., Andreassen, M., & Sütterlin, S. (2017). [Title to verify]. Frontiers in Psychology, 8, 742. https://pmc.ncbi.nlm.nih.gov/articles/PMC5432607/
- Pelletier, C. L. (2004). [Title to verify]. Journal of Music Therapy, 41(3), 192–214. https://pubmed.ncbi.nlm.nih.gov/15327345/
- PR Newswire. (2020, June). MedRhythms receives FDA Breakthrough Device Designation for chronic stroke digital therapeutic [Press release]. https://www.prnewswire.com/news-releases/medrhythms-receives-fda-breakthrough-device-designation-for-chronic-stroke-digital-therapeutic-301076597.html
- PR Newswire. (2023, May 23). Endel and Universal Music Group to create AI-powered, artist-driven functional music designed to support listener wellness [Press release]. https://www.prnewswire.com/news-releases/endel-and-universal-music-group-to-create-ai-powered-artist-driven-functional-music-designed-to-support-listener-wellness-301832191.html
- PR Newswire. (2025, January 16). MedRhythms' InTandem rehabilitation system for chronic stroke gait impairment receives final Medicare payment determination from CMS [Press release]. https://www.prnewswire.com/news-releases/medrhythms-intandem-rehabilitation-system-for-chronic-stroke-gait-impairment-receives-final-medicare-payment-determination-from-cms-302353537.html
- Prova Health. (2025). DiGA reimbursement in Germany: A guide. https://www.provahealth.com/insights/diga-reimbursement-germany-guide [Exact publication date to verify — the repository records 2024–2025]
- Ries, E. (2011). The lean startup. [Publisher and place to verify]
- Rogers, E. M. (2003). Diffusion of innovations (5th ed.). [Publisher and place to verify]
- Schmitz, A., Kueppers, L., Klein, J., [remaining authors to verify]. (2025). [Title to verify]. JMIR Research Protocols, 14, e63380. https://www.researchprotocols.org/2025/1/e63380
- Shaffer, F., & Ginsberg, J. P. (2017). An overview of heart rate variability metrics and norms. Frontiers in Public Health, 5, 258.
- Spielberger, C. D. (1983). State–Trait Anxiety Inventory manual. [Edition, publisher and place to verify]
- Tarvainen, M. P., [remaining authors to verify]. (2002). [Title to verify — smoothness-priors detrending]. IEEE Transactions on Biomedical Engineering, 49, 172. [Issue and end page to verify]
- Task Force of the European Society of Cardiology and the North American Society of Pacing and Electrophysiology. (1996). [Title to verify — the repository records the work only as "HRV standards"]. Circulation, 93, 1043. [Issue and end page to verify]
- TDMusic Inc. (2022–2026). Company background and asset inventory [Internal documents: business-model deck (2024), thesis tutor presentation (2026), company brochure (2022)]. [Individual document titles, authors and dates to verify; the founding date and the 2022 catalogue figures are flagged for reconciliation in the company's own file]
- Tebra. (2025). Music for mental health [Healthcare report]. The Intake. https://www.tebra.com/theintake/healthcare-reports/music-for-mental-health
- TechCrunch. (2022, April 5). Endel raises $15M to further develop its AI-powered sound wellness technology. https://techcrunch.com/2022/04/05/endel-raises-15m-to-further-develop-its-ai-powered-sound-wellness-technology/
- TechCrunch. (2025, September 16). Calm launches standalone iOS app for sleep support. https://techcrunch.com/2025/09/16/calm-launches-standalone-ios-app-for-sleep-support/
- 36Kr. (2026). 从"香饽饽"到"鸡肋",睡眠 APP 困于商业化 [From hot commodity to dispensable: Sleep apps stuck on monetisation]. https://36kr.com/p/1852579050655107
- TMTPost. (2023, April). [Title to verify — NMPA classification determination for cognitive-function-disorder software]. https://www.tmtpost.com/6473884.html
- Tracxn. (2026). Calm [Company profile]. https://tracxn.com/d/companies/calm/__11__6TilJ32thRjVlCh0DLinrJLZ9o47hTQaLwndAnI
- Ulwick, A. W. (2002). Turn customer input into innovation. Harvard Business Review, 80(1), 91–97.
- U.S. Food and Drug Administration. (n.d.). General wellness: Policy for low risk devices [Guidance document]. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/general-wellness-policy-low-risk-devices [Publication and update dates to verify; the repository records the guidance as cited 2019–2026 with a January 2026 update reported only by secondary law-firm sources]
- van Westendorp, P. H. (1976). NSS-Price Sensitivity Meter (PSM): A new approach to study consumer perception of price. Proceedings of the ESOMAR Congress, Venice. [Publisher and page range to verify]
- VBData. (2026). [Title to verify — Tide (潮汐) Pre-A funding announcement]. https://www.vbdata.cn/40749
- Venkatesh, V., & Davis, F. D. (2000). [Exact title to verify — the TAM2 extension paper]. Management Science, 46(2), 186–204.
- Vickers, A. J., & Altman, D. G. (2001). Analysing controlled trials with baseline and follow-up measurements. BMJ, 323, 1123. [Issue and end page to verify]
- Virtual Therapeutics. (2024, July). Virtual Therapeutics announces results of tender offer to acquire Akili Interactive [Press release]. Nasdaq. https://www.nasdaq.com/press-release/virtual-therapeutics-announces-results-tender-offer-acquire-akili-interactive-2024-07
- Weiser, M., & Brown, J. S. (1996). The coming age of calm technology. [Publication details to verify]
- Woods, K. J. P., Sampaio, G., James, T., [remaining authors to verify]. (2024). [Title to verify]. Communications Biology, 7, 1376. https://www.nature.com/articles/s42003-024-07026-3
- World Health Organization. (2023). Anxiety disorders [Fact sheet]. https://www.who.int/news-room/fact-sheets/detail/anxiety-disorders
105 entries; 65 carry at least one bracketed field the repository does not supply. Every gap is itemised in site/pages/dissertation.references-todo.md, grouped by the kind of field missing, so the gaps can be closed in one library session rather than one at a time.
Appendices
ASurvey instrument map
The full seven-block instrument, with every item, its wording in both languages, the construct it measures, the pre-registered threshold it carries and the belief it feeds. Blocks: A who · B why, when and where (GAD-2, PSS-4, context, coping, outcome-driven innovation) · C what works (the state-dependent forced choice that produced the central finding) · D how to prove it (wearable willingness, the action gap) · E why use it and why not (technology acceptance, Kano, form, barriers) · F how much (Van Westendorp and purchase intent) · G voice.
BFocus-group kit
Recruitment screener, moderator script, the three blind stimulus specifications (S-A calm real song; S-B neutral adaptive sound; S-C guided voice), the private-rating sheet used before discussion, the coding frame, and the synthesis of both groups with the theme matrix and the dot vote.
CDecision-engine register
The live belief register: eight assumptions with priors and rationales, every evidence item with its likelihood ratio, quality tag and source, the pending pre-registered rows, the concept joint probabilities and the gate check. The page is interactive — priors can be moved and pending evidence toggled, so a reader can see for themselves how much the verdict depends on any single judgement.
DApp blueprint
The ten-section product blueprint: human-factors principles with their reference products, the system architecture, the thirteen-stage user journey, eight screens plus the watch and safety screens, the data-capture interface matrix (HealthKit, Oura API v2, Polar H10 over BLE, Web Bluetooth), the two-tier library with its tag schema and an embedded parameter-adaptation bench, the academic-method-to-interaction mapping table, the feedback and learning loop, the technology stack, and the v0 control rules.
EGlossary
| Term | Meaning as used in this dissertation |
|---|---|
| A1a … A6 | The eight named assumptions in the belief register: A1a real music in acute states, A1b real music in wind-down, A2a a session lowers state anxiety, A2b a session moves the wearable measure, A3 users connect and value the result, A4 someone pays, A5 adaptation rights are feasible, A6 retention exceeds the category norm. |
| Active control | A comparison condition that is a plausible treatment (here, a generic relaxing playlist the participant already knows) rather than nothing, so that "any audio" is not confounded with "our audio". |
| ANCOVA form | Analysing a trial by regressing the post-treatment value on the treatment and the baseline value, rather than analysing raw change scores. |
| C1, C2, C3 | The three concepts: C1 the consumer app (in two variants, real-music and neutral-sound), C2 an adaptive-audio licence to hardware brands, C3 an artist calm-content line through existing distribution. |
| d and dz | Standardised effect sizes; dz is the paired within-person version used for a crossover design. |
| E0 … E4 | The experiment programme: E0 wearable calibration, E1 the efficacy pilot, E2 in-app micro-experiments, E3 contextual-bandit personalisation, E4 the self-generating library. |
| GRADE | A standard system for grading the certainty of a body of clinical evidence; used here to set the quality-shrinkage exponent applied to a likelihood ratio. |
| HRV | Heart-rate variability. Higher HRV indicates greater parasympathetic (vagal) tone, so relaxation raises HRV and lowers heart rate — the direction is checked explicitly throughout. |
| Likelihood ratio (LR) | P(evidence | hypothesis) / P(evidence | not hypothesis). Above 1 supports, below 1 disconfirms, exactly 1 is uninformative. |
| Quality index | The share of 60-second windows in a session that pass the artefact rules. Below 0.70 the session leaves the physiological analysis and no number is shown to the user. |
| Residual (Δr) | Observed heart rate minus what the user's personal baseline model predicted for that time of day. The session outcome is the mean residual over the last five minutes minus the mean residual over the pre-rest. |
| Settle / Unwind | The product's two modes: Settle is the engineered neutral bed for acute states; Unwind is real, artist-linked music for wind-down. |
| Shrinkage (k) | The exponent applied to a likelihood ratio in log space to reflect evidence quality: LRadjusted = LRk. Weaker evidence gets a smaller k and therefore moves the belief less. |
| STAI-S / STAI-6 | The state scale of the State–Trait Anxiety Inventory, and its six-item short form; the pilot's primary outcome, scaled 20–80. |
| Van Westendorp (PSM) | A four-question price-sensitivity method yielding a point of marginal cheapness, an optimal price point, an indifference price point and a point of marginal expensiveness. |
| Wizard of Oz | A prototype in which a human performs the algorithm's role from a written rule table, so the policy can be tested before the software exists. |