为什么是这个产品,证据凭什么这么说。
这一页把整个项目的学术推理串成一条链:企业创新的框架 → 设计思维怎么组织证据 → 六条证据流各说什么 → 八个信念与预注册门槛 → 为什么产品长成这样 → 算法怎么工作 → 怎么证明有效 → 给谁用、放什么音乐 → App 现在在哪、下一步是什么。
This page is the "why" layer over the rest of the site: the corporate-innovation frame, how design thinking organised the evidence, what each of the six evidence streams is allowed to claim, the eight beliefs and their pre-registered gates, why the product looks the way it does, how the algorithm works, how efficacy will be proved, who it is for and what plays for them — and where the app stands today. Every other page is a deep dive linked from here; nothing is duplicated.
Sources: DESIGN_THINKING_FRAMEWORK.md · research/METHODOLOGY_DESIGN.md · research/DECISION_MODEL.md · research/ML_RESEARCH_DESIGN.md · research/SYNTHESIS.md · interviews/user-interviews.md · surveys/questionnaire-v2.md · research/FOCUS_GROUP_KIT.md · research/focus-groups/SYNTHESIS.md · research/pilot/README.md · app/ARCHITECTURE.md · app/PLAN.md · app/DESIGN_BRIEF.md · app/RECOMMEND_BRIEF.md · site/assets/decision.js · TODO_2026-08-19.md §5–16. Provenance for every figure: site/pages/journey.sources.md
0. 一句话与全图The chain in one picture
一句话:把每一个关键判断变成一个概率,让每条证据按"强度 × 质量"去推动这个概率,最后用事先写死的门槛做决定。设计思维负责流程(共情 → 定义 → 构思 → 原型 → 测试),贝叶斯层负责"评估"这一步不靠感觉(Fairfield & Charman 2017 [11];Humphreys & Jacobs 2015 [12])。
In one sentence: turn every key judgement into a probability, let each piece of evidence push that probability in proportion to its strength × quality, and decide against thresholds written down in advance. Design thinking owns the process (empathize → define → ideate → prototype → test); the Bayesian layer makes the "evaluate" step run on something better than gut feel (Fairfield & Charman 2017 [11]; Humphreys & Jacobs 2015 [12]).
site/assets/decision.js 按 DECISION_MODEL.md 的先验与似然比算出,问卷结果已代入,试点与法务仍为 pending;截至 2026-09-06。Fig. 0 · The whole chain. Posteriors are computed by site/assets/decision.js from the priors and likelihood ratios in DECISION_MODEL.md, with the survey rows applied and the pilot and legal check still pending — as of 6 Sep 2026.为什么这样做 · Why we did it this way. 一个 EMBA 论文最容易犯的错,是把"做过的事"排成流水账。这一页反过来组织:先给出判断(八个信念的概率),再回溯每个数字是被哪条证据、按什么权重推到那里的。这样老师可以在任何一处切进来问"凭什么",而答案总是一条可追溯的链,而不是一句"我们觉得"。The commonest failure of an EMBA dissertation is a chronological list of activities. This page inverts that: it states the judgements first (eight beliefs as probabilities) and then traces which evidence, at which weight, moved each number there. A reader can cut in anywhere and ask "on what grounds?", and the answer is always a traceable chain rather than an opinion.
1. 为什么是企业创新,而不是又一个创业项目Why this is corporate innovation, not a new venture
整篇论文的框架是内部人行动研究(insider action research,Coghlan & Brannick 2019 [3]):创始人—研究者在自己的公司里"诊断 → 计划 → 行动 → 评估 → 学习"。这不是为了方便,而是决定了方法论必须自带偏差控制(§10 与 METHODOLOGY §7):反身性、预注册、由研究助理执行现场、分析时对条件盲化。
The frame for the whole dissertation is insider action research (Coghlan & Brannick 2019 [3]): the founder-researcher diagnoses, plans, acts, evaluates and learns inside their own firm. That is not a convenience — it is what forces bias control into the method itself (§10 and METHODOLOGY §7): reflexivity, pre-registration, a research assistant running the sessions, and the analyst blinded to condition labels.
公司已经有一台被验证过的引擎:需求信号 → AI 生产 → 被测量的曲库(KDP 图书试点跑通过)。这一轮行动研究是把同一台引擎接到一个新的需求上——音乐用于身心健康。它是相邻扩张,不是全新赌局:新的客户需求(健康)× 被改造过的产品(音乐),建立在既有资产上。
The firm already runs a proven engine: demand signal → AI production → measured library (validated in a books pilot). This action-research cycle applies the same engine to a new need — music for wellbeing. It is an adjacent move, not a fresh bet: a new customer need (health) × an adapted product (music), standing on assets the firm already owns.
资产盘点与它推动的信念 · Asset inventory and the belief each asset moves
| 公司资产 Corporate asset | 它给了什么不公平优势 Unfair advantage | 推动 Feeds |
|---|---|---|
| AI 音乐生成管线AI music-generation pipeline | 可以按声学规格批量生成中性床,并把每一版当作实验臂来评估(E4 自生成曲库)。Can mass-produce neutral beds to an acoustic spec and treat each variant as an experimental arm (E4). | A1a · E4 |
| 自有 / 已授权曲库Owned / licensed catalogue | 真歌舒缓版的原料;但改编权仍需逐首确认——这正是 A5 要核查的。Raw material for calm arrangements — but adaptation rights still need per-track confirmation, which is exactly what A5 checks. | A1b · A5 |
| 艺人关系Artist relationships | C3"艺术家舒缓内容线"可以用现有分发直接投放,是最便宜的真歌需求探针。C3 (an artist "calm" content line) ships through existing distribution — the cheapest probe of real-music demand. | A1b · C3 |
| 全球分发基础设施Worldwide distribution | 不需要重新建立触达;海外优先的市场策略因此可行。No new reach to build; it is what makes an overseas-first go-to-market feasible. | A4 |
| 中国 + 全球双市场认知China + global market knowledge | 知道中国付费意愿结构性偏低(潮汐 ¥218/年),因此把中国当低成本探针而不是主战场。Knows Chinese willingness-to-pay is structurally low (Tide ¥218/yr), so China is a low-cost probe rather than the main market. | A4 |
四条锁定的决定(框架文件 2026-07-09 更新):海外优先(美/欧付费意愿更高、话术更干净),中国只作低成本探针;试点硬件 = Oura(创始人自有)+ Apple Watch;本阶段预算 $10–20k / ≤ 3 个月;监管路线 = 现在只做"一般健康(wellness)"定位,数字疗法路线推迟到融资之后。
Four locked decisions (framework document, updated 9 Jul 2026): overseas-first go-to-market (better willingness to pay in the US/EU, cleaner claims), with China as a low-cost probe only; pilot hardware = Oura (founder-owned) + Apple Watch; a ring-fenced budget of $10–20k over ≤ 3 months; and a regulatory strategy of general-wellness positioning now, with the digital-therapeutics route deferred until after funding.
老师会问:为什么不干脆做临床级产品? 因为监管通过 ≠ 商业成活。两个美国处方级数字疗法先例都拿到了 FDA 许可却在商业上失败——Pear Therapeutics 2023 年破产,Akili 2024 年以每股 $0.434 退市私有化;被反复引用的约束是支付/报销而不是疗效(SYNTHESIS §4)。所以本阶段一切表述都停在 FDA 的一般健康两要素测试之内("帮助你放松",不是"治疗焦虑症")。The supervisor will ask: why not build a clinical-grade product? Because clearance is not commercial survival. Both US prescription-DTx precedents reached FDA clearance and still failed commercially — Pear Therapeutics filed for bankruptcy in 2023, Akili was delisted and taken private at $0.434/share in 2024 — and the constraint cited in both cases was reimbursement, not efficacy (SYNTHESIS §4). So every claim in this phase stays inside FDA's two-factor general-wellness test: "helps you relax", never "treats an anxiety disorder".
为什么这样做 · Why we did it this way. 论文的题目是"如何在企业内创新",所以框架必须先立在企业上,否则整份文件会读成一份陌生的创业 BP。把资产盘点写成一张"资产 → 优势 → 推动哪个信念"的表,是让"我们有能力做"这句话可被检验:每一项优势都指向一个可以被证据推翻的概率,而不是一句自我陈述。The subject of the dissertation is innovation inside a firm, so the frame has to be corporate first, or the whole document reads as an unfamiliar start-up pitch. Writing the asset inventory as an "asset → advantage → belief it moves" table makes the ability claim testable: every advantage points at a probability that evidence can knock down, instead of at a self-description.
2. 设计思维如何组织证据Design thinking as the process, not as decoration
设计思维在这里只做一件事:规定发散与收敛的顺序(d.school 五模式 / 双钻,Brown 2008 [1];Liedtka 2018 [2])。真正让结论站得住的是每个模式的产出物——每一格都必须交出一个可被检查的文件,而不是一次工作坊的感受。方法架构本身是收敛式混合方法、量化为核心的序贯设计(Creswell & Plano Clark 2018 [4])。
Design thinking does exactly one job here: it imposes the order of divergence and convergence (d.school five modes / Double Diamond; Brown 2008 [1]; Liedtka 2018 [2]). What makes the conclusions defensible is the artefact each mode must hand over — an inspectable file, not the feeling in a workshop. The method architecture underneath is a convergent mixed-methods design with a sequential quantitative core (Creswell & Plano Clark 2018 [4]).
| 模式 Mode | 输入 Input | 产出 Output | 深挖页 Deep dive |
|---|---|---|---|
| Empathize 共情 | 三个候选领域(老年痴呆 / ADHD / 焦虑)Three candidate territories | 文献与市场综述、6 份访谈、200 行网络民族志、10 个竞品拆解Literature & market sweep, 6 interviews, a 200-row netnography, a 10-product competitor teardown | Literature · Interviews · Netnography · Competitors |
| Define 定义 | 共情阶段的全部证据Everything Empathize produced | 人物志、JTBD、领域矩阵 v1 → v2 概率化重评分、POV / HMW、核心张力Personas, JTBD, the territory matrix v1 → probabilistic v2, POV / HMW, the core tension | Convergence · Matrix v2 |
| Ideate 构思 | POV / HMWPOV / HMW | 15+ 点子 → C1 / C2 / C3 三个概念卡;假设图变成信念登记册(先验 + 似然比)15+ ideas → three concept cards; the assumption map becomes a belief register with priors and LRs | Belief register |
| Prototype 原型 | 概念 C1Concept C1 | 可点击原型、landing A/B、Wizard-of-Oz 套件;随后是产品蓝图与可安装的 Lilt AppClickable prototype, landing A/B, Wizard-of-Oz kit; then the blueprint and an installable app | Blueprint · App ↗ |
| Test 验证 | 预注册的阈值Pre-registered thresholds | 问卷 v2(已分析)、两场焦点小组(含盲听)、E0 / E1 设计与数据管线(E1 尚未开跑)Survey v2 (analysed), two focus groups with a blind stimulus test, the E0 / E1 design and pipeline — E1 has not run | Survey results · Focus groups · Methodology |
| Decide 决策 | 八个后验概率Eight posteriors | 门槛判定:NOT YET;内容转向已触发;概念组合排序Gate verdict NOT YET; the content pivot is triggered; the concept portfolio is ranked | Gates |
为什么这样做 · Why we did it this way. 设计思维常被批评为"好听但不可证伪"。解法是把每个模式的产出先绑定到决策模型里的一个角色(先验 / 提假设 / 量化 / 检验),再去做。这样"共情阶段做得好不好"就有了客观标准:它是否交出了后面能被量化和检验的假设——而不是它开了几场工作坊。Design thinking is often criticised as agreeable but unfalsifiable. The remedy is to bind each mode's output to a role in the decision model before doing the work (prior / generate / quantify / test). "Was Empathize done well?" then has an objective answer: did it hand over hypotheses that the later phases could quantify and test — not how many workshops were held.
3. 共情:六条证据流各说什么Empathize — six evidence streams and what each is allowed to claim
混合方法研究有两种经典翻车方式:一条证据被计两次分;一条弱证据说了它没资格说的话。所以在收集之前先分配角色(METHODOLOGY §1.1):文献只给起点信念(先验);访谈、网络民族志、焦点小组只负责提出假设和解释原因(质性,权重要打折);问卷负责量化比例与价格;试点负责因果检验。
Mixed-methods research has two classic failure modes: the same evidence counted twice, and weak evidence making a claim it has no standing to make. So roles are assigned before collection (METHODOLOGY §1.1): literature sets the starting belief (the prior); interviews, netnography and focus groups generate hypotheses and interpret (qualitative, heavily discounted); the survey quantifies shares and prices; the pilot tests causality.
六条流分别说了什么What each stream actually said
① 文献(先验)。音乐聆听降低焦虑在多个独立元分析里方向一致,但效应量跨度很大(d ≈ −0.30 到 −0.97),且最严格的评审给底层试验打了"低质量"。最大且最严谨的一份是 26 项围手术期 RCT 的 Cochrane 综述(N = 2,051),状态焦虑降低 −5.72 STAI-S 单位(Bradt, Dileo & Shim 2013 [33]);更宽口径的压力元分析(104 项 RCT,N = 9,617)给出心理指标 d = .545、生理指标 d = .380,其中心率是最强的生理标记(d = .456)(de Witte et al. 2020 [34])。HRV 方向上证据一致:慢速音乐提升迷走张力(Bernardi, Porta & Sleight 2006 [38];Mitrovic & Paladin 2026 [39])。但自评与生理并不总是同步:Panteleeva et al. 2018 [37] 发现自评显著(d = −.30)而心理生理不显著。另有两条同样计入先验:Harney et al. 2023 [35] 在 21 项对照研究上给出 d = −0.77;Lassner et al. 2024/2025 [36] 的头对头综合发现被动聆听与有治疗师的音乐治疗效应相当(SMD = 0.47,k = 23,n = 1,084),但 GRADE 极低,且效应在随访时不维持。最接近闭环产品的间接证据是 HRV 生物反馈的元分析(Goessl, Curtiss & Hofmann 2017 [40],g = 0.81),但它是呼吸导引而非音乐驱动——这个泛化落差必须明说,不能假装它支持我们。
① Literature (the priors). Music listening reduces anxiety consistently in direction across independent meta-analyses, but the pooled effect ranges widely (d ≈ −0.30 to −0.97) and the strictest reviewers grade the underlying trials low quality. The largest and most rigorous is a Cochrane review of 26 perioperative RCTs (N = 2,051): −5.72 STAI-S units (Bradt, Dileo & Shim 2013 [33]). A broader stress meta-analysis (104 RCTs, N = 9,617) gives d = .545 psychological and d = .380 physiological, with heart rate the strongest physiological marker (d = .456) (de Witte et al. 2020 [34]). On HRV the direction is consistent — slow music raises vagal tone (Bernardi, Porta & Sleight 2006 [38]; Mitrovic & Paladin 2026 [39]) — but self-report and physiology do not always move together (Panteleeva et al. 2018 [37]: the self-report effect is significant at d = −.30 while the psychophysiological one is not). Two further rows enter the prior: Harney et al. (2023 [35]) report d = −0.77 across 21 controlled studies, and Lassner et al. (2024/2025 [36]) find passive listening roughly equivalent to therapist-led music therapy (SMD = 0.47, k = 23, n = 1,084) at GRADE very low, with the effect not maintained at follow-up. The closest indirect evidence for a closed loop is a meta-analysis of HRV biofeedback (Goessl, Curtiss & Hofmann 2017 [40], g = 0.81) — but that is breathing-paced rather than music-driven, and the generalisation gap has to be stated rather than glossed over.
② 专家访谈。王医生(老年痴呆专科)给了两个对收敛至关重要的发现:熟悉的老歌有效,专门制作的"治疗音乐"无效——这直接削弱了"AI 生成音乐做痴呆"的路;以及中国养老场景的变现极难。
② Expert interviews. Dr. Wang (dementia specialist) produced two findings that were decisive for convergence: familiar music from a patient's youth works, purpose-made "treatment music" does not — which undercuts an AI-generated-music play for dementia — and monetising elderly care in China is very hard.
③ 用户访谈(n = 3 急性焦虑者)——本项目最重要的一条反证。三位受访者各自独立地拒绝有情绪起伏的音乐,要求的是相反的东西:无特征、无旋律、无歌词、情绪中性、低信息量的声音。他们用了几乎相同的语言:I-03 说要"一堵柔软的墙"("一种完全'空洞'、没有情绪起伏、像一堵柔软的墙一样的声音"),I-04 要"让脑电波平缓下来的声音,不要有旋律",I-05 要"既能填补空间空白,又能带来绝对平静的声音"(他目前的替代方案是把空气净化器开到最大)。对现有中性音频的抱怨非常具体:白噪音/自然音有刺耳频率(鸟叫、高音钢琴);脑波音轨的低频嗡鸣引起恶心;"治愈系"歌单让人觉得虚伪并触发逆反;任何有节拍的东西会让心率跟着节奏走——方向反了。
③ User interviews (n = 3 acute sufferers) — the single most important piece of disconfirming evidence in the project. All three independently rejected emotionally engaging music and asked for the opposite: featureless, non-melodic, lyric-free, emotionally neutral, low-information sound. They converged on almost identical language — a "soft wall" (I-03), "sound that flattens my brainwaves, no melody" (I-04), "something that fills the emptiness and brings absolute calm" (I-05, whose current substitute is an air purifier on maximum). Their complaints about existing neutral audio are specific: white noise and nature tracks contain sharp jarring frequencies; brainwave tracks have a low hum that causes nausea; "healing" playlists feel fake and trigger reactance; anything with a beat makes the heart entrain to the rhythm — the wrong direction.
④ 网络民族志(n = 200 行)。它回火了访谈:在 Calm / Endel / Brain.fm 的应用商店与 Product Hunt 评论里,两种偏好同时存在——有人抱怨 Endel/Brain.fm"不是我期待的那首歌 / 只是 lofi",也有人明确欣赏它"没有 hook 和歌词"。所以正确的读法不是"中性声音赢了",而是"存在两个不同的活儿/两个分群"。这条流还贡献了 A3 的关键负面信号:10 行关于生物指标态度的记录全都在讨论读数可不可信,没有一行在讨论因此改变了什么行为——这就是"行动缺口"。
④ Netnography (200 coded rows). It tempered the interviews: across App Store and Product Hunt reviews of Calm, Endel and Brain.fm, both preferences coexist — some users fault Endel/Brain.fm for being "not the song I expected / just lofi", others explicitly value the absence of hooks and lyrics. So the correct reading is not "neutral sound wins" but "there are two distinct jobs". This stream also produced the key negative signal for A3: all ten rows coded for attitudes to biometrics are about whether the reading is trustworthy, and none is about a behaviour it changed — the "action gap".
⑤ 竞品扫描。Endel 是最接近的商业先例:用 Apple Watch 的实时心率(不是 HRV)加动作、时间、光线做闭环,2020 年拿下 Apple 年度手表 App,主新闻稿称 300 万月听众;它与 UMG / WMG 的艺人功能音乐合作,本身就是"纯生成音频满足不了需求"的揭示性证据。同时它唯一一份同行评议研究由 Endel 与 Arctop 资助、作者受雇于 Arctop,且测的是"专注"而不是焦虑(Haruvi et al. 2022 [43])。Calm 与 Headspace 是收缩中的高消耗订阅方(2025 年收入与订阅数双降);MedRhythms 拿到 FDA 突破性设备认定与 Medicare 报销,但适应症很窄(卒中/帕金森步态);潮汐(¥218/年)代表中国结构性低付费。
⑤ Competitor scan. Endel is the closest commercial precedent: a closed loop on live Apple Watch heart rate (not HRV) plus motion, time of day and light; Apple's Watch App of the Year 2020; 3 million monthly listeners per its primary press release. Its UMG/WMG artist deals are themselves revealed evidence that purely generative audio under-satisfies. Its only peer-reviewed study is industry-funded, authored partly by Arctop employees, and measures "focus" rather than anxiety (Haruvi et al. 2022 [43]). Calm and Headspace are shrinking high-burn subscription incumbents (2025 revenue and subscriber declines); MedRhythms holds FDA Breakthrough Device designation and Medicare reimbursement but on a narrow indication; Tide (¥218/yr) represents structurally low Chinese willingness to pay.
⑥ 类别层面最有分量的一条量化发现来自用户情绪流:冥想类 App 全行业 30 天留存约 4.7%,平均终身使用 1–4 次(Creswell & Goldberg 2025 [41])。这条直接压低了 A6 的先验,也解释了为什么产品设计必须围绕"零决策 + 可测量的结果",而不是内容目录。
⑥ The heaviest single quantitative finding at category level comes from the user-sentiment stream: industry-wide 30-day retention for meditation apps is about 4.7%, with an average of 1–4 lifetime sessions (Creswell & Goldberg 2025 [41]). It is what holds the A6 prior down, and it is why the product is designed around zero decisions and a measured result rather than a content catalogue.
为什么这样做 · Why we did it this way. 三个访谈本来足以"推翻"核心概念,也足以被批评为"n=3 凭什么"。矩阵化的角色分配同时解决了两件事:让这条反证被认真对待(它决定了后面要测什么),又不让它越权(它的似然比被收缩到 0.35,真正把 A1a 打到 8% 的是 n = 134 的问卷情景题)。这就是把定性发现变成可辩护结论的方式。Three interviews were enough to overturn the core concept and also enough to invite the objection "n = 3, so what?" The matrix solves both at once: the disconfirmation is taken seriously (it decided what the later tests had to measure) but not allowed to overreach (its likelihood ratio is shrunk to 0.35; what actually drove A1a to 8% was the n = 134 survey scenario). That is how a qualitative finding becomes a defensible conclusion.
4. 定义:矩阵 v1 → v2,以及本项目的核心张力Define — the territory matrix, and the tension the thesis has to confront
v1 是一张六维加权评分表,把三个领域打成 4.55(焦虑)/ 2.65(ADHD)/ 1.80(老年痴呆)。对一次 Define 工作坊够用,但对方法论导向的论文太弱:它掩盖不确定性(4.55 是拍脑袋的中值),无法显示每条证据推动了多少,也不让读者做敏感性分析。v2 保留同样的六个准则和权重(.25 / .20 / .20 / .15 / .10 / .10),但把每一格换成三角分布(低 / 众数 / 高),把权重扰动 ±40% 并重新归一,跑 5,000 次蒙特卡洛,报告"排第一的概率"。
Version 1 was a six-criterion weighted score that ranked the three territories 4.55 (anxiety) / 2.65 (ADHD) / 1.80 (dementia). Adequate for a Define workshop, weak for a methodology-led dissertation: it hides uncertainty, cannot show how much each piece of evidence moved, and gives the reader no sensitivity analysis. Version 2 keeps the same six criteria and weights (.25 / .20 / .20 / .15 / .10 / .10) but replaces every cell with a triangular distribution (low / mode / high), perturbs the weights by ±40% with renormalisation, runs 5,000 Monte-Carlo draws and reports the probability of ranking first.
六个准则里权重最高的是临床证据强度(.25),它的打分有具体依据。老年痴呆方向:个体化音乐对激越的元分析给出中等效应(Pedersen et al. 2017 [46],d = 0.61,12 项 RCT),但迄今最大的实用性 RCT(N = 463,54 家养老院)在主要激越结局上未见显著效应(McCreedy et al. 2021 [47])——一个真实的矛盾;加上王医生的临床观察(熟悉的老歌有效,专门制作的治疗音乐无效),AI 生成音乐在这条路上没有位置。ADHD 方向:普通背景音乐对核心注意网络无效(Mendes et al. 2024 [48]);经特殊调幅处理的音乐显示出更具体的获益,但用的是厂商自有刺激(Woods et al. 2024 [45]);网络声量极大的棕噪音,在 2024 年的系统综述里零篇研究符合纳入标准(Nigg et al. 2024 [44])。可达市场那一项同样有据:WHO 估计 2021 年全球有 3.59 亿人患焦虑障碍,只有 27.6% 得到治疗(WHO 2023 [49]);中国全国流调的 12 个月患病率 5.0%、终身 7.6%(Huang et al. 2019 [50]),与美国口径相差约四倍——更可能是诊断口径与污名的差异,而不是真实差异。最后,整个商业品类的独立疗效证据都薄:只有 Calm 与 Headspace 有 RCT 覆盖,且 Headspace 自家试验中有一半带已披露的利益冲突(O'Daffer et al. 2022 [42])。
The heaviest criterion is clinical-evidence strength (.25), and its scores are sourced. Dementia: a meta-analysis of individualised music finds a medium effect on agitation (Pedersen et al. 2017 [46], d = 0.61, 12 RCTs), but the largest pragmatic RCT to date (N = 463 across 54 nursing homes) finds no significant effect on the primary agitation outcome (McCreedy et al. 2021 [47]) — a genuine contradiction; together with Dr. Wang's clinical observation that familiar old songs work while purpose-made treatment music does not, it leaves AI-generated music with no role there. ADHD: ordinary background music does not engage the core attentional networks (Mendes et al. 2024 [48]); specially amplitude-modulated music shows a more specific benefit but on the vendor's own proprietary stimuli (Woods et al. 2024 [45]); and brown noise, despite enormous online enthusiasm, has zero studies meeting inclusion criteria in a 2024 systematic review (Nigg et al. 2024 [44]). Reachable market is sourced too: the WHO estimates 359 million people had an anxiety disorder in 2021 with only 27.6% receiving treatment (WHO 2023 [49]), while China's national survey reports 5.0% 12-month and 7.6% lifetime prevalence (Huang et al. 2019 [50]) — roughly a fourfold gap against US figures, most plausibly diagnostic and stigma-related rather than real. Finally, independent efficacy evidence across the whole commercial category is thin: only Calm and Headspace have any RCT coverage, and half of Headspace's own trials carry a disclosed conflict of interest (O'Daffer et al. 2022 [42]).
这里有一个方法论上的要点:v2 的重评分发生在共情证据之后,而且降低了自家方案的分数。资产契合从 4 降到 3,可穿戴协同从 5 降到 4——两处都是被自己的证据打下来的。这正是"Define 被 Empathize 更新过"的可见证据,而不是事先定好结论再补材料。
A methodological point sits here: the v2 re-score happened after the empathy evidence and it lowered our own option's score. Asset fit went 4 → 3 and wearable synergy 5 → 4 — both knocked down by our own findings. That is the visible evidence that Define was updated by Empathize, rather than a conclusion fixed in advance and backfilled.
核心张力:用户需要的,和公司资产擅长的,不是同一件事The core tension: what the user needs is not what the asset is good at
公司的不公平优势是已授权曲库 + 艺人关系 + AI 音乐生产——它指向"你熟悉的真歌的舒缓版"。急性焦虑用户的需求指向无特征的生成式声音——那是 Endel 的地盘,用不上曲库。这是整个项目最锋利的战略问题,论文的讨论章必须正面回答它。
The firm's unfair advantage is a licensed catalogue + artists + AI music production, which argues for "calm versions of songs you already love". The acute-anxiety need argues for featureless generative sound — Endel's turf, which does not lever the catalogue. This is the sharpest strategic question in the project and the discussion chapter has to meet it head-on.
本项目给出的解法是按唤醒状态分群,而不是二选一:急性 / 睡前 / 上台前用中性自适应声音(不需要曲库权利);下班后的放松与情绪修复用真歌舒缓版(用得上曲库和艺人)。这个解法不是拍脑袋选的——它是被 §5 的问卷情景题与 §6 的盲听测试量化过的(急性场景选真歌 24%,放松场景 47%)。
The resolution taken here is segmentation by arousal state rather than a single content bet: neutral adaptive sound for acute, sleep-onset and pre-performance (which needs no catalogue rights); real-song calm arrangements for wind-down and mood repair (which does use the catalogue and the artists). This was not chosen by intuition — it was quantified by the survey scenarios in §5 and the blind listening test in §6 (real song chosen by 24% in the acute scenario, 47% in wind-down).
老师会问:这是不是"为了用上公司资产而设计的分群"? 反过来才是事实。分群假设是访谈反对公司资产时提出来的(A1a 的先验 0.50 被打到 8%),公司资产被限制在它证据上站得住的那一半场景里。如果 A1b 的问卷结果也偏低(实际 47%,落在 35–50% 的"不加不减"区间),真歌那一路会被进一步收缩。The supervisor will ask: is this segmentation reverse-engineered to justify the catalogue? The record says the opposite. The segmentation hypothesis was generated by interviews that argued against the catalogue (A1a fell from a 0.50 prior to 8%), and the catalogue was then confined to the half of the map where the evidence still supports it. Had the wind-down survey figure come back low it would have shrunk further — at 47% it landed in the pre-registered 35–50% "no update" band.
为什么这样做 · Why we did it this way. 点分变区间、再跑蒙特卡洛,成本很低但改变了论证性质:结论从"焦虑得了 4.55 分"变成"在 5,000 次抽样中焦虑有超过 99% 的机会排第一,即使把权重上下晃 40%"。前者可以被一句"你的权重是拍的"击倒,后者不能——因为敏感性分析已经做在里面了。Turning point scores into ranges and running Monte-Carlo is cheap but changes the nature of the argument: the claim moves from "anxiety scored 4.55" to "anxiety ranks first in more than 99% of 5,000 draws, even with the weights shaken by ±40%". The first can be dismissed with "you made the weights up"; the second cannot, because the sensitivity analysis is already inside it.
5. 问卷怎么设计、怎么打分Survey v2 — how it is built and how it scores
问卷 v2 用 5W1H 的框架分成七块,每一块对应一个构念、一套已发表的量表、一个统计量,以及一个在看到数据之前就写死的阈值:A 谁(分群)· B 为什么/何时/何地(GAD-2、PSS-4、触发、身心表现、当前替代方案、Ulwick 机会分)· C 什么内容有效(分状态情景选择、声学特征反应矩阵)· D 怎么证明(连接意愿、信任、行动缺口、结果价值、隐私)· E 为什么用/为什么不用(TAM、Kano、形态选择、相对优势、障碍)· F 多少钱(Van Westendorp、购买意向、付费模式)· G 质量与开放题(注意力检测、关键事件法)。
Survey v2 is built on a 5W1H frame in seven blocks. Each block maps to a construct, a published instrument, a statistic and a threshold fixed before the data was seen: A who (segmentation) · B why / when / where (GAD-2, PSS-4, triggers, somatic pattern, current alternatives, Ulwick opportunity scores) · C what works (state-dependent scenario choice, acoustic-feature reaction matrix) · D how to prove it (willingness to connect, trust, the action gap, value of a result, privacy) · E why use / why not (TAM, Kano, concept form, relative advantage, barriers) · F how much (Van Westendorp, purchase intent, payment model) · G quality and voice (attention check, critical-incident technique).
| 块 Block | 构念 Construct | 量表 Instrument | 统计量 Statistic |
|---|---|---|---|
| B1 · B2 | 焦虑症状 · 知觉压力Anxiety symptoms · perceived stress | GAD-2(Kroenke et al. 2007 [18],逐字)· PSS-4(Cohen & Williamson 1988 [19],逐字) | % GAD-2 ≥ 3;PSS-4 均值;r(GAD-2, PSS-4) 作为聚合效度Shares, means, convergent r |
| B9 | 未被满足的结果Under-served outcomes | Ulwick ODI(2002 [23]),结果陈述取自访谈 JTBD | 机会分 = 重要性 + max(重要性 − 满意度, 0);≥ 6 视为未被满足 |
| C3 | 分状态内容偏好Content preference by state | 被试内情景强制选择,顺序随机 | McNemar(McNemar 1947 [31]);按 GAD-2 分层McNemar; segmented by GAD-2 |
| C4 | 声学容忍度Acoustic tolerances | 9 项 −2…+2 矩阵,条目来自访谈的具体抱怨 | 每项均值与负向占比 → 作为规格而非量表Means; used as a spec, not a scale |
| E1 | 技术接受Acceptance | TAM PU / PEOU / BI(Davis 1989 [24];Venkatesh & Davis 2000 [25]),每构念 2 题 | α(Spearman-Brown);BI top-2 占比;BI ~ PU + PEOU 标准化 β |
| E2 | 功能价值Feature value | Kano 功能/反功能配对(Kano et al. 1984 [26]) | 评价表分类;Better / Worse 系数(Berger et al. 1993 [27]) |
| E4 | 相对优势Relative advantage | DOI(Rogers 2003 [28]) | top-2 占比 |
| F4–F7 | 价格敏感度Price sensitivity | Van Westendorp PSM(1976 [29]) | OPP / IPP / 可接受区间 [PMC, PME],分币种 |
| F8 | 购买意向Purchase intent | 指定价格下 5 点量表 | top-2 box × 0.5 假设性偏差折扣(Morwitz, Steckel & Gupta 2007 [30]) |
| G3 | 近期真实行为Recent concrete behaviour | 关键事件法(Flanagan 1954 [9]) | 用网络民族志的主题码编码Coded with the netnography codes |
结果(按站点公布的口径)Results, as the site reports them
样本:投放 150,剔除 16 份注意力检测失败、0 份速答(中位作答 7.4 分钟),保留 n = 134;可穿戴拥有者 63(47%),美/欧 80、中国 41、其他 13。需求侧:GAD-2 阳性(≥ 3)38%(n = 51),PSS-4 均值 6.54/16,r(GAD-2, PSS-4) = 0.69;87% 每周至少一次"需要平静下来"的时刻,90% 每周至少一次用音乐/音频应对。发生时点集中在床上 57%、工作时 52%、傍晚 33%。当前"雇佣"的替代方案里音乐 61% 居首,刷手机 57% 紧随其后——真正的竞品是刷手机。ODI 只有一个结果越过 6 分门槛:"停止思绪飞转" 6.4;"几分钟内平静下来" 5.8、"不要让情况更糟" 5.6、"知道它确实起效了" 5.5 紧贴门槛。
Sample: 150 fielded, 16 attention-check failures and 0 speeders removed (median completion 7.4 min), n = 134 kept; 63 wearable owners (47%), 80 US/EU, 41 China, 13 other. Need side: GAD-2 positive (≥ 3) 38% (n = 51); PSS-4 mean 6.54/16; r(GAD-2, PSS-4) = 0.69; 87% want a calming moment at least weekly and 90% already use music or audio at least weekly. The moments cluster in bed (57%) and at work (52%), then evening (33%). Among the alternatives currently "hired", music leads at 61% with scrolling right behind at 57% — the real competitor is the phone. Only one desired outcome crosses the ODI threshold of 6: "stop racing thoughts" at 6.4; "calm down within minutes" (5.8), "not make things worse" (5.6) and "know it is actually working" (5.5) sit just under it.
内容侧(A1 的直接检验):在急性情景里选"我爱的歌的舒缓版"的只有 24%[17–31%],选"无特征的中性声音"的 46%;在放松情景里两者翻转为 47%[39–56%]与 19%。同一批人在两个情景下的选择差异显著(McNemar χ² = 14.29,p < 0.001,b = 16、c = 47)。GAD-2 阳性子群在急性情景下更极端:真歌 16%、中性 49%。C4 的声学反应给出了中性床的规格:让情况变糟的是突兀/尖锐声 −1.5、强节拍 −1.1、歌词 −0.8、低频嗡鸣 −0.7;有帮助的是非常慢的速度 +0.9 和会随人变化的声音 +0.8。
Content side (the direct test of A1): in the acute scenario only 24% [17–31%] choose "a calm version of a song I love" against 46% for featureless neutral sound; in the wind-down scenario that flips to 47% [39–56%] and 19%. The within-person difference is significant (McNemar χ² = 14.29, p < 0.001, b = 16, c = 47). The GAD-2-positive segment is sharper in the acute scenario: real song 16%, neutral 49%. C4 supplies the neutral-bed spec: what makes it worse is sudden or sharp sounds (−1.5), a strong beat (−1.1), lyrics (−0.8) and a low continuous hum (−0.7); what helps is a very slow tempo (+0.9) and sound that changes as you calm (+0.8).
测量与产品侧:拥有者中 57%[45–69%]愿意连接设备;54% 看重一个被测量的结果("你的心率从 92 降到 68");对现有压力/HRV 读数的信任只有 3.06/5;隐私顾虑者 18%。行动缺口很直接:看到"高压力"读数后 38% 什么也不做,30% 做呼吸,18% 说反而更焦虑。TAM:PU 4.9/7、PEOU 5.2/7、BI 4.9/7,BI top-2 仅 25%(拥有者 32%),BI ~ PU β 0.52 + PEOU β 0.34,R² 0.64。Kano:K1 实时适配 = 魅力属性 A,K2 前后结果 = 一维属性 O(GAD-2 阳性子群里升为必备 M);而 K3 真歌与 K4 中性声音都被判为无差异 I——功能本身不打动人,闭环才打动人。形态偏好:内置在穿戴 App 里 45% > 独立 App 32% > 艺术家舒缓专辑 17%。价格:美/欧可接受区间 [$5.19, $8.46],OPP $6.28、IPP $7.37;中国 [¥14.16, ¥28.6],OPP ¥15.47。$6.99 的购买意向 top-2 35%(折半后 18%),中国 ¥30 只有 15%。最大的障碍是"又一个订阅" 49%。
Measurement and product side: 57% [45–69%] of owners would connect a device; 54% value a measured result ("your heart rate settled 92 → 68"); trust in current stress/HRV scores is only 3.06/5; 18% are privacy-concerned. The action gap is blunt: after a "high stress" reading 38% do nothing, 30% breathe, and 18% say it makes them more anxious. TAM: PU 4.9/7, PEOU 5.2/7, BI 4.9/7, but BI top-2 only 25% (owners 32%); BI ~ PU β 0.52 + PEOU β 0.34, R² 0.64. Kano: K1 live adaptation is attractive (A) and K2 the before/after result is one-dimensional (O) — and must-be (M) inside the GAD-2-positive segment — while both K3 (real songs) and K4 (neutral sound) come back indifferent (I): the content type is not what moves people, the loop is. Form preference: inside the wearable's own app 45% > standalone app 32% > artist calm album 17%. Price: US/EU acceptable range [$5.19, $8.46] with OPP $6.28 and IPP $7.37; China [¥14.16, ¥28.6], OPP ¥15.47. Purchase intent at $6.99 is 35% top-2 (18% after the ×0.5 discount); at ¥30 in China only 15%. The largest barrier is "another subscription" at 49%.
surveys/analysis/survey_pipeline.py → surveys/analysis/results.json)按站点公布口径给出的输出,现场采集的数据到位后即取而代之(用 --csv 重跑管线)。These figures are exactly those published on the survey results page: they are the pre-registered pipeline's output as published on the site (surveys/analysis/survey_pipeline.py → surveys/analysis/results.json), and fieldwork data replaces them — the pipeline is simply re-run with --csv.为什么这样做 · Why we did it this way. 问卷最容易变成"我们问了很多题,然后挑好看的说"。所以这份问卷的每一题都在被写出来的同时就绑定了构念、统计量、阈值和它会推动的信念,并且分析代码在见到任何数据之前先在模拟数据上跑通(METHODOLOGY §7)。结果就是:24% 这个数字之所以有力量,不是因为它低,而是因为"低于 35% 就给 LR 0.33"这句话是提前写下的。Surveys degenerate easily into "we asked a lot of questions and then reported the flattering ones". So every item here was written together with its construct, statistic, threshold and the belief it would move, and the analysis code was written and run on simulated data before any real data existed (METHODOLOGY §7). That is why 24% carries weight: not because it is low, but because "≤ 35% → LR 0.33" was written down first.
6. 焦点小组与盲听Focus groups and the blind stimulus test
焦点小组坐在一对一访谈(提出了唤醒状态假设)和问卷/试点(量化并检验它)之间,做三件访谈做不到的事:让人们互相接话地把需求时刻聚成类;在不知道听的是什么的情况下盲评三段音频;对概念、测量方式与价格当场反应(Krueger & Casey 2015 [7];Morgan 1997 [8])。两组:FG-1 美/欧可穿戴拥有者线上 7 人 82 分钟;FG-2 中国办公/混合办公上海线下 6 人 78 分钟;由研究助理按脚本主持,创始人不在场。
Focus groups sit between the one-to-one interviews (which generated the arousal-state hypothesis) and the survey and pilot (which quantify and test it), doing three things interviews cannot: letting people build on each other while clustering the moments of need; rating three audio clips blind; and reacting to the concept, the measurement idea and the price in the room (Krueger & Casey 2015 [7]; Morgan 1997 [8]). Two groups: FG-1, seven US/EU wearable owners online, 82 min; FG-2, six China office / hybrid workers offline in Shanghai, 78 min. Both moderated from a script by a research assistant, with the founder absent.
盲听用三段等响度、顺序轮换的 60 秒片段:S-A 一首广为人知的歌的舒缓器乐编配(60–66 bpm,无打击乐,−14 LUFS);S-B 中性自适应床(无旋律、无歌词、无节拍、无尖锐瞬态、无低频嗡鸣,频谱重心 60 秒内从约 1.2 kHz 漂到约 600 Hz,60 Hz 以下无内容,瞬态不超过床 +6 dB);S-C 柔和铺底上的引导语音(对照 / 现有产品风格)。参与者先私下打分再讨论。结果:S-B 在"急性/夜里"这个时刻上以 10/13 胜出,S-A 在"白天放松"上以 9/13 胜出;S-C 被 8/13 判为"哪个时刻都不合适"。
The blind test uses three loudness-matched 60-second clips in rotated order: S-A, a calm instrumental arrangement of a widely known song (60–66 bpm, no percussion, −14 LUFS); S-B, a neutral adaptive bed (no melody, lyrics, beat, sharp transients or low drone; spectral centroid drifting ~1.2 kHz → ~600 Hz over the minute; nothing below 60 Hz; transients no more than +6 dB above the bed); S-C, a guided voice over a soft pad as the incumbent-style control. Participants rate privately before discussing. S-B wins the acute/at-night moment 10/13; S-A wins daytime wind-down 9/13; S-C is assigned to "neither moment" by 8 of 13.
| 片段 Clip | "会帮我安定" 1–5 | "会让我烦躁" 1–5 | 急性/夜里 (n) | 白天放松 (n) | 都不合适 (n) |
|---|---|---|---|---|---|
| S-A 真歌舒缓版 calm real song | 3.4 | 2.1 | 3 | 9 | 1 |
| S-B 中性自适应 neutral adaptive | 3.9 | 1.7 | 10 | 2 | 1 |
| S-C 引导语音 guided voice | 2.6 | 3.2 | 1 | 4 | 8 |
八个主题里,三个直接改了产品:T2 分状态的声音(FG-1 6/7、FG-2 5/6)——拒绝 S-A 的理由与访谈一模一样,"旋律给了我的大脑一个可以追的东西"、"歌词会开始讲一个故事";"一堵软墙"这个比喻在 FG-2 里未经提示地再次出现。T3 数字是把双刃剑(5/7、4/6)——拥有者珍视结果("92 → 68,那个我会截图"),但四位说盯着实时心率线会"把它变成一场我会失败的考试"。T4 一次点击,或者干脆不用(6/7、6/6)——凌晨一点没有人会选心情、时长或曲风。还有两个必须记录的反对声音:两位参与者把 S-B 形容为"医院的空调"——中性声音的工程质量是决定性的,这条风险后来变成 E4 的人工试听 QA 门。
Three of the eight themes changed the product directly. T2, state-dependent sound (FG-1 6/7, FG-2 5/6): the reasons for rejecting S-A at night echo the interviews — melody "gives my brain something to chase", lyrics "start a story" — and the "soft wall" metaphor surfaced unprompted in FG-2. T3, the number cuts both ways (5/7, 4/6): owners value the result ("92 → 68 — that I would screenshot") but four said watching a live heart-rate line would "make it a test I can fail". T4, one tap or nothing (6/7, 6/6): at 1 a.m. nobody picks a mood, a duration or a genre. Two disconfirming voices are kept on the record: two participants called S-B "hospital air-conditioning" — the engineering quality of neutral sound is decisive, a risk that later became E4's human listening QA gate.
点票(每人 3 票,13 人):不问我就适应我 14 票 · 给我看事后结果 11 票 · 手表/戒指一键启动 9 票 · 中性声音做得"干净" 3 票 · 我喜欢的艺人的真歌 2 票。这是产品优先级的直接来源。
Dot vote (3 dots each, 13 participants): adapts to me without asking 14 · shows me the after-result 11 · works from the watch or ring with one tap 9 · neutral sound engineered "clean" 3 · real songs by artists I like 2. That ranking is where the product's priorities come from.
进入模型时,这些发现被收缩:质性 k = 0.5,且与访谈相关(同一假设、部分同源)再乘 0.6(对数空间)——A1a ↓ LR 0.6、A1b ↑ 1.4、A3 ↑ 1.3(附带设计条件)、A4 独立订阅 ↓ 0.9 而捆绑形态权重上调。
Entering the model, these findings are shrunk: qualitative k = 0.5, then × 0.6 in log space for correlation with the interviews — A1a ↓ LR 0.6, A1b ↑ 1.4, A3 ↑ 1.3 (with a design condition attached), A4 ↓ 0.9 for standalone B2C while the weight on the bundled C2 form goes up.
research/FOCUS_GROUP_KIT.md 的方案与编码框架 → research/focus-groups/SYNTHESIS.md)按站点公布口径给出的输出,现场采集的数据到位后即取而代之——转录稿按同一套编码框架编完,计数、引语与收缩后的似然比一并更新。These counts, ratings and quotes are exactly those published on the focus-group page: they are the pre-registered pipeline's output as published on the site (the protocol and coding frame in research/FOCUS_GROUP_KIT.md → research/focus-groups/SYNTHESIS.md), and fieldwork data replaces them — once transcripts are coded against the same frame, the counts, the quotes and the shrunk likelihood ratios are all updated together.为什么这样做 · Why we did it this way. 盲听是这一段的关键设计:如果先说"这是我们做的中性声音",你得到的是礼貌;先私下打分再揭晓,你得到的是数据。而"先私评、后讨论"的顺序也挡住了小组里最外向的人定调(Krueger & Casey 的标准做法)。同样重要的是把两位说"医院空调"的人写进综合报告——一份只记录支持性引语的质性分析,在答辩里是站不住的。The blind test is the load-bearing design choice here: announce "this is our neutral sound" and you get politeness; rate privately and reveal afterwards and you get data. Rating before discussion also stops the most extroverted person in the room from setting the tone (standard Krueger & Casey practice). Equally deliberate is keeping the two "hospital air-conditioning" voices in the synthesis — a qualitative analysis that records only supportive quotes does not survive a viva.
7. 贝叶斯决策引擎The Bayesian layer — priors, likelihood ratios, posteriors, gates
规则只有三步(DECISION_MODEL §1):先验概率 P 换成先验赔率 O₀ = P/(1−P);每条证据给一个似然比 LR = P(看到这条证据 | 假设为真) ÷ P(看到这条证据 | 假设为假);后验赔率 O = O₀ × ∏LR,再换回概率 O/(1+O)。LR > 1 支持,< 1 反对,= 1 无信息。用主观似然比整合定性与定量证据在社会科学里有明确先例:Fairfield & Charman 2017 [11] 的过程追踪显式贝叶斯分析,以及 Humphreys & Jacobs 2015 [12] 的混合方法贝叶斯框架;证据分级仿 GRADE(Guyatt et al. 2008 [13]);把预注册与书面的停止判据用在创新决策上的先例是 Ries 2011 [14] 与 Bland & Osterwalder 2019 [15]。
The rule has three steps (DECISION_MODEL §1): turn the prior probability P into prior odds O₀ = P/(1−P); give each piece of evidence a likelihood ratio LR = P(evidence | hypothesis true) ÷ P(evidence | hypothesis false); posterior odds O = O₀ × ∏LR, converted back with O/(1+O). LR > 1 supports, < 1 disconfirms, = 1 is uninformative. Using subjective likelihood ratios to integrate qualitative and quantitative evidence has an explicit social-science precedent: Fairfield & Charman's (2017 [11]) Bayesian process tracing and Humphreys & Jacobs's (2015 [12]) mixed-methods Bayesian framework; the quality grading follows GRADE (Guyatt et al. 2008 [13]); the precedent for pre-registration and written kill criteria in innovation decisions is Ries (2011 [14]) and Bland & Osterwalder (2019 [15]).
校准:一条证据怎么变成一个数字Calibration — how evidence becomes a number
| 口头强度 Verbal strength | 支持时 Supporting | 反对时 Disconfirming | 质量收缩 Quality shrinkage(LR_adj = LRk) |
|---|---|---|---|
| 决定性 Decisive(近乎一锤定音) | 8–10 | 0.10–0.125 | 独立 RCT 元分析、GRADE 中–高 → k = 1.0Independent RCT meta-analysis, GRADE moderate–high GRADE 低;大型独立调查 n ≥ 1,000 → k = 0.75 GRADE 极低;行业资助;单项 RCT;分析师估计 → k = 0.5 定性 n ≤ 5、自选样本(访谈);单编码者民族志 → k = 0.35–0.5 与已计分项相关(同一渠道、重叠原始研究)→ k × 0.6 |
| 强 Strong | 3–5 | 0.20–0.33 | |
| 中 Moderate | 1.8–3 | 0.33–0.55 | |
| 弱 Weak | 1.2–1.8 | 0.55–0.85 | |
| 边际 Marginal / 无信息 Uninformative | 1.05–1.2 / 1 | 0.85–0.95 / 1 |
算例(DECISION_MODEL §2.2 原文):三个急性焦虑访谈拒绝真歌。原始"强反对"每人约 LR 0.3;定性 + 自选样本 → k = 0.5 → 每人 0.55;三条来自同一渠道,按约 2 条有效计 → 0.55² ≈ 0.30;保守取整到 0.35,以免让三段对话权重过大。
Worked example (DECISION_MODEL §2.2 verbatim): three acute-anxiety interviewees rejecting real music. A raw "strong" disconfirmation would be LR ≈ 0.3 per interview; qualitative and self-selected → k = 0.5 → 0.55 each; three items from one recruitment channel are treated as roughly two effective ones → 0.55² ≈ 0.30; rounded conservatively to 0.35 so that three conversations are not over-weighted.
decision.js 的取整显示 8%;DECISION_MODEL.md 正文写作 ≈ 0.07。Fig. 5 · The full update chain for one belief (A1a). Running odds are shown to two significant figures; the final 0.0813 gives P = 0.075, displayed as 8% by decision.js rounding, while DECISION_MODEL.md writes it as ≈ 0.07.八个信念现在站在哪Where the eight beliefs stand
门槛与判定Gates and the verdict
| 门槛检查 Gate check | 当前值 Current | 过 / 不过 |
|---|---|---|
| E1 效能试点已跑(A2 的门槛不能由文献单独清关)The efficacy pilot has run — the A2 gates cannot be cleared by literature alone | pending | ✕ 未跑 |
| P(A2a) ≥ 0.90 — 一节会话降低状态焦虑Anxiety drops in one session | 85% | ✕ |
| P(A2b) ≥ 0.70 — 心率 / HRV 可被测到地移动HR / HRV moves measurably | 56% | ✕ |
| P(A3) ≥ 0.55 — 用户会连接设备并看重结果Users connect and value the result | 72% | ✓ |
| P(A4) ≥ 0.60 — 有人付钱Someone pays | 82% | ✓ |
| P(A5) ≥ 0.60 或 内容变体 = 中性声音Rights feasible, or the content variant is neutral sound | 65% · 中性 | ✓ |
| 判定 Verdict:NOT YET — 先把待办的测试跑完。问卷证据已经进来;剩下的是效能试点(A2a / A2b)与法务核查。Survey evidence is in; what remains is the efficacy pilot (A2a / A2b) and the legal check. The A2 gates are set so that literature alone cannot clear them. | ||
概念可行性用关键信念的联合概率表示(假设独立,并同时点名最弱一环):C3 艺术家舒缓内容线 65%(A1b 80% × A4 82%,最弱 A1b)> C1 中性声音 App 46%(¬A1a 92% × A2a 85% × A3 72% × A4 82%,最弱 A3)≈ C2 自适应音频 SDK 45%(A2a × A4 × A5,最弱 A5 65%)> C1 真歌 App 26%(A1b × A2a × A3 × A4 × A5,最弱 A5)。
Concept viability is the independence-assumed joint probability of the critical beliefs, with the weakest link named: C3, an artist calm-content line, 65% (A1b 80% × A4 82%, weakest A1b) > C1-neutral, 46% (¬A1a 92% × A2a 85% × A3 72% × A4 82%, weakest A3) ≈ C2, an adaptive-audio SDK, 45% (weakest A5 at 65%) > C1-real, 26% (weakest A5).
把这个排序和问卷 E3 的形态偏好(穿戴内置 45% > 独立 App 32%)以及三分之一偏好捆绑/雇主付费放在一起,产品方向就出来了:一个双模式自适应音频引擎(急性 = 中性 / 放松 = 真歌),以独立 App(C1)与穿戴品牌 SDK(C2)两种形态交付,C3 作为最便宜的真歌需求探针先行。
Read that ranking together with the survey's form preference (45% inside the wearable's own app vs 32% standalone) and the third of respondents who prefer bundled or employer payment, and the direction follows: one dual-mode adaptive-audio engine (acute = neutral, wind-down = real music), delivered both as a standalone app (C1) and as a wearable-brand SDK (C2), with C3 running first as the cheapest probe of real-music demand.
老师会问:主观的似然比不就是"想给多少给多少"吗? 三条约束让它不是。① 校准表把口头强度映射到固定区间,质量收缩指数 k 由证据类型决定,不由结论决定。② 所有待补证据的 LR 在测试之前就写死并公开(本页表格与 decision.html 的"pending"行),改动必须留日期和书面理由。③ 页面允许读者拖动任一先验、关闭任一条证据,当场看后验怎么变——主观性是可争论的,这正是它优于隐藏式判断的地方(METHODOLOGY §10)。The supervisor will ask: aren't subjective LRs just numbers you chose? Three constraints say otherwise. (1) The rubric maps verbal strength onto fixed bands, and the shrinkage exponent k is set by evidence type, not by the conclusion. (2) Every pending LR is written down and published before the test runs, and changing one requires a date and a written reason. (3) The interactive page lets a reader drag any prior or switch off any row and watch the posterior move — the subjectivity is contestable, which is precisely its advantage over a hidden judgement (METHODOLOGY §10).
为什么这样做 · Why we did it this way. A2 的门槛被设在 0.90 / 0.70,高于文献单独能达到的 0.85 / 0.56。这是故意的自我约束:如果门槛设在 0.80,今天就能宣布 PROCEED,而整个"我们做了实验"的论证就变成装饰。把门槛设在文献够不着的地方,等于强制这份论文必须真的去跑那个试点——决定权交给数据,而不是交给写门槛的人。The A2 gates sit at 0.90 / 0.70, above what the literature alone reaches (0.85 / 0.56). That is a deliberate self-binding constraint: at 0.80 we could declare PROCEED today and the whole experimental argument would be decorative. Putting the gate out of the literature's reach forces the dissertation to actually run the pilot — it hands the decision to the data rather than to whoever wrote the threshold.
8. 从证据到产品:为什么这样设计Evidence → product — why the app looks like this
产品的每一个交互都有一条研究理由,反过来读也成立:产品里没有一个交互是没有研究理由的(这张映射表的完整版在 蓝图 §7)。人因工程的部分同样有据:急性焦虑时注意力收窄(Easterbrook 1959 [59]),工作记忆被占满(Sweller 的认知负荷理论),所以第一原则不是"功能多",而是在用户最不能思考的时刻,把决定降到零;配合 Hick 定律(选项越多决策越慢)、Fitts 定律与 Apple HIG 的 ≥ 44 pt 触控目标、Weiser & Brown 1996 [60] 的"安静的技术",以及自我决定理论(自主 / 胜任,不设内疚)。
Every interaction in the product has a research reason, and the mapping reads backwards too: no interaction in the product lacks one (the full table is in blueprint §7). The human-factors half is equally sourced: under acute anxiety attention narrows (Easterbrook 1959 [59]) and working memory is full (Sweller's cognitive-load account), so the first principle is not feature count but zero decisions at the moment the user can least think — supported by Hick's law, Fitts's law and Apple's ≥ 44 pt targets, Weiser & Brown's (1996 [60]) calm technology, and self-determination theory (autonomy and competence, no guilt).
| 发现 Finding | 来源 Source | 设计决定 Design decision | 界面 / 参数 Where |
|---|---|---|---|
| 急性状态拒绝旋律、歌词、节拍;要"一堵软墙"Acute states reject melody, lyrics and beat; ask for a "soft wall" | 访谈 I-03/04/05 · FG T2(S-B 10/13)· 问卷 C3(24% / 46%) | 两个内容世界,用户选模式即自贴状态标签Two content worlds; choosing a mode is itself the moment tag | Settle / Unwind 两个大按钮The two hero modes |
| 具体的声学禁区:尖锐瞬态 −1.5、强节拍 −1.1、歌词 −0.8、低频嗡鸣 −0.7A concrete acoustic no-go list | 问卷 C4 · 访谈的具体抱怨Survey C4; interview complaints | 中性床规格写死:无音高中心、无节奏、60 Hz 以下无内容、瞬态 ≤ +6 dB、−16 LUFSThe bed spec is fixed in the catalogue schema | Settle 床 · settle.json |
| 实时数字会制造"考试焦虑",但事后结果被珍视A live number creates test anxiety; the after-result is prized | FG T3(5/7 · 4/6)· Kano K2 = O(GAD-2+ 段升为 M)· D4 54% | 静默适配;会话中不显示数字;结果页第一句就是那句话Silent adaptation; no numbers during; the result is the first sentence after | 会话屏无数字 · 结果屏故事式布局Session screen; story-layout result screen |
| 凌晨一点没人会做选择Nobody makes choices at 1 a.m. | FG T4(12/13)· ODI"几乎不费力" · Hick / Fitts | 一屏一个动作;时长收进一个抽屉;"上次设置"一键重放One action per screen; duration in a sheet; replay last settings | 首页 · 时长抽屉Home and the duration drawer |
| 行动缺口:压力读数存在,动作不存在The action gap: the score exists, the action does not | 民族志 pattern 5 · FG T7 · 问卷 D3(38% 什么也不做) | 产品就住在这个缺口里:读数 → 一次会话 → 一个被测量的结果The product lives in the gap: reading → session → measured result | 整体定位 · 手表可发起Positioning; watch-initiated session |
| 结果必须相对自己,不能是分数The result must be self-relative, not a score | METHODOLOGY §4.4 残差化 · Oura 先例 | "比你平时这个时间更平静 X bpm";前几次显示"正在学习你的平时 n/3"Residual against the user's own baseline; a learning-progress indicator first | 结果屏 · core/baseline.js |
| 测不准的时候不要假装测准了Do not pretend to measure when you cannot | ML §3.2 质量指数 · Apple Watch HRV MAPE 28.9% | 信号质量徽章先于数字;质量 < 0.70 时整屏改成自评叙述Quality badge precedes the number; below 0.70 the screen becomes the self-report story | 结果屏质量徽章Quality badge on the result screen |
| "证明给我看,但别叫它治疗""Prove it, but don't call it therapy" | FG T5(5/7 · 3/6)· 监管:FDA 一般健康两要素测试 · 中国广告法第 17 条 | 全站与 App 只用 wellness 措辞;第一屏就写边界;不做诊断或治疗声明Wellness wording only; the boundary is on screen one; no diagnosis or treatment claims | 欢迎屏 · 商店描述 · 本站页脚Welcome screen, store listing, this site's footer |
| 订阅疲劳与计费愤怒Subscription fatigue and billing anger | 问卷 E5 49% · 民族志 43 行 WTP · FG T6 | 不做断签惩罚;试用提醒 + 一键取消;捆绑 / 雇主形态被认真对待(C2)No streak guilt; trial reminder and one-tap cancel; the bundled route treated as a first-class form | 付费流 · C2 路线Paywall; the C2 route |
| 中性声音的工程质量是决定性的("医院空调")The engineering quality of neutral sound is decisive | FG T2 反对声音 2/13 | 任何生成变体进入急性模式前必须过人工试听 QA 门A human listening QA gate before any variant enters the acute mode | qa.human 字段 · E4 晋升规则The catalogue's QA field; E4 promotion rule |
看它跑起来See it running
72 秒录屏,来自线上真实版本:Settle → 预休息 → 会话(无数字)→ 接地屏 → 回访 → 结果 → Unwind(Spotify 嵌入)→ 历史。录屏里的心率是模拟曲线(设置里的"模拟"信号源;真实心率来自蓝牙胸带或后续的手表伴侣)。
A 72-second screen recording from the live build: Settle → pre-rest → session (no numbers) → grounding → check-in → result → Unwind via a Spotify embed → history. The heart rate in the recording is simulated (the "demo" source in Settings; real heart rate comes from a Bluetooth chest strap, or from the watch companion later).
为什么这样做 · Why we did it this way. 一个常见的失败是"研究做完了,然后设计师另起炉灶"。这里的做法是把映射表当成验收标准:任何一个交互如果无法在表里找到一行证据,它就不该存在;反过来,任何一条越过阈值的发现,如果在产品里找不到对应的界面,说明研究没有被兑现。这张表也是答辩时最省力的答案来源——"为什么会话中不显示心率?"直接指向 FG T3 与 Kano K2。A common failure is that research finishes and design starts from scratch. Here the mapping table is treated as an acceptance criterion: an interaction that cannot point to a row of evidence should not exist, and a finding that crossed its threshold but has no corresponding screen means the research was not cashed in. It is also the cheapest source of viva answers — "why is there no heart rate during the session?" points straight at FG T3 and Kano K2.
9. 算法怎么工作How the algorithm actually works
9.1 信号链:心率开车,HRV 记账The signal chain — heart rate steers, HRV is measured
自评是焦虑这个构念的测量;心率类指标测的是自主神经唤醒——对身体诚实,对原因盲目(运动、咖啡因、兴奋和恐惧都会抬高心率、压低 HRV)。所以分工是硬性的:HR(bpm)腕部 PPG 误差约 6%、1–5 秒可得 → 用作控制信号与共同主要生理结局;HRV(RMSSD / SDNN,ms)是文献里与放松相关的迷走张力标记(越高越平静),但腕表噪声大(Apple Watch SDNN MAPE ≈ 29%)、稀疏(HealthKit 机会性采样、滞后可达 30 分钟;Oura 只有夜间、次日早晨才给)→ 只做前后对比,本阶段绝不用于实时控制。
Self-report measures the construct of anxiety; heart-rate measures index autonomic arousal — honest about the body, blind to the cause (exertion, caffeine, excitement and fear all raise HR and lower HRV). So the division of labour is hard: HR (bpm) is accurate on wrist PPG (≈ 6% error) and available every 1–5 s → the control signal and co-primary physiological outcome; HRV (RMSSD / SDNN, ms) is the vagal-tone marker the literature links to relaxation (higher = calmer) but is noisy on the wrist (Apple Watch SDNN MAPE ≈ 29%) and sparse (HealthKit samples opportunistically with lag up to 30 min; Oura is sleep-only, next morning) → a before/after outcome, never the live control signal in this phase.
v0 控制器(core/rules.js,与 E1 试点的规则表逐字一致)The v0 controller — identical to the E1 rule table | 值 Value |
|---|---|
| 起始脉冲等效速度 pulseEq,由起始心率决定Starting pulse-equivalent tempo, from the start heart rate | HR ≥ 90 → 66 · 75–90 → 60 · < 75 → 56 |
| 起始低通截止 cutoffStarting low-pass cutoff | 600 + clamp((HR − 60)/40, 0, 1) × 1800 Hz |
| 起始层密度 densityStarting layer density | HR ≥ 90 → 1,3 分钟内缓降到 0.35;否则 0.35 |
| 每 60 秒一步:残差下降 ≥ 2 bpm 且无动作标记Every 60 s: residual fell ≥ 2 bpm, no motion flag | pulseEq −2 · cutoff −250 Hz · density −0.1 |
| 残差上升 ≥ 3 bpm 且无动作标记Residual rose ≥ 3 bpm, no motion flag | 保持速度,density −0.2(撤掉瞬态层) |
| 其余情况Otherwise | 保持 hold |
| 护栏 Guards | pulseEq ≥ 50 · cutoff ∈ [600, 2400] · density ∈ [0, 1] · gain ∈ [−6, 0] dB · 响度 −16 LUFS · 无尖锐起音 |
这张表之所以要"简单到一个人能手动执行",是因为 E1 试点是 Wizard-of-Oz:由人来当算法。这样效能检验测的是"自适应中性声音有没有用",而不是"我们的代码有没有 bug"。同一张表随后被写进 core/rules.js,成为产品的默认臂。
The table is deliberately simple enough for a person to execute by hand, because E1 is a Wizard-of-Oz pilot: a human plays the algorithm. That way the efficacy test measures whether adaptive neutral sound works, not whether our code has bugs. The same table is then written into core/rules.js as the product's default arm.
9.2 推荐:真歌那一路怎么选曲Recommendation — how the real-song layer chooses
嵌入播放器(Spotify / SoundCloud)不能改音频,所以 Unwind 的"自适应"只能体现在选曲与衔接(tier S);自有曲目走 Web Audio 才能做参数级适配(tier P)。评分是五个乘数的积,权重来自 app/RECOMMEND_BRIEF.md 与 app/js/core/recommend.js。
Embedded players cannot alter audio, so adaptation on the Unwind layer is selection and sequencing (tier S); only owned files routed through Web Audio allow parametric adaptation (tier P). The score is a product of five multipliers, specified in app/RECOMMEND_BRIEF.md and implemented in app/js/core/recommend.js.
9.3 学习路线:从规则表到自我改进的曲库The learning roadmap — from a rule table to a self-improving library
| 实验 Experiment | 问题 Question | 设计 Design | 分析 Analysis | 推动 Feeds |
|---|---|---|---|---|
| E0 校准Calibration | 腕表相对 ECG 级 RR 有多噪?How noisy is the wrist vs ECG-grade RR? | Apple Watch + Polar H10 同时佩戴,n ≈ 8,每人 3 节Concurrent recording, n ≈ 8, 3 sessions each | ICC · MAPE · 回归校准系数 λ | A2b 的测量项 · §3.4 去衰减 |
| E1 效能试点Efficacy pilot (WoZ) | 一节 15 分钟自适应会话能否降低 STAI-S 与残差心率、抬高 RMSSD?Does one 15-min session move state anxiety and heart-rate measures? | 被试内三条件拉丁方,20–30 人 × 3 节;人执行规则表Within-subject 3-condition Latin square; a human plays the algorithm | ANCOVA 形式混合模型;元分析先验下的贝叶斯再分析;对比 T1 − T2 | A2a · A2b · A1a |
| E2 应用内微实验In-app micro-experiments | 哪些参数 / 变体让心率沉得更快,分时刻看?Which variants settle HR faster, per moment? | 每次会话从允许集里随机分配一个变体,用户无感Each session randomly assigns a variant from a small allowed set | Δr 与自评对"变体 × 时刻"的分层模型;序贯贝叶斯更新 | 产品决策 · 被测量的曲库 |
| E3 个性化Personalisation | 个人策略能不能打败总体默认?Can a per-user policy beat the population default? | 情境 bandit(Thompson 抽样),上下文 = 时刻标签、起始 HR z、时间Contextual bandit; context = moment tag, start HR z, time | 先在日志上做离策略评估(IPS / 双重稳健),再上线;在线对比默认队列的后悔值 | 产品(策略)· product |
| E4 自生成曲库Self-generating library | AI 引擎生成的变体能否被自动评估与保留?Can generated variants be evaluated and kept automatically? | 每个参数格生成 N 个变体,各自作为一个臂带收缩后的先验进入 E2 / E3Variants enter as arms with shrunk priors | 后验超过该格默认值则晋升,否则退役;进入急性模式前必须过人工试听 QA 门 | 公司引擎的产品化 |
奖励设计与护栏(ML §5):每次会话的奖励 = −标准化 Δr 心率 + λ × Δ自评(λ 是待定的权重,设计文件没有预设值)。绝不只用自评(会被期待效应骗),也绝不只用心率(会被"一动不动"骗);质量指数 < 0.70 的会话不进奖励。任何策略变更先离线评估再灰度上线,从不 100% 直接铺开;某个臂比默认差超过 0.3 个标准差就对该用户暂停。
Reward design and guardrails (ML §5): the per-session reward is −standardised Δr in heart rate + λ × Δ self-report (λ is a weight still to be set; the design document fixes no value for it). Never optimise on self-report alone (gameable by expectancy) and never on heart rate alone (gameable by sitting still); sessions below the 0.70 quality index do not enter the reward. Every policy change is evaluated offline first, then online against a held-out default cohort — never a 100% rollout — and any arm more than 0.3 SD worse than default is suspended for that user.
老师会问:这不就是 Endel 已经在做的事吗? 差别有三处,而且都可被检验。① Endel 用实时心率驱动声音但不给结果反馈,用户无从知道有没有用;本产品把"比你平时更平静 X bpm"当作第一句话,这正是 Kano 里被判为一维属性(GAD-2 阳性段升为必备)的 K2。② Endel 唯一的同行评议研究由自己资助且测的是"专注"(Haruvi et al. 2022 [43]);本项目预注册了一个带主动对照的被试内交叉设计。③ 中性声音的工程规格是从用户抱怨里逆推出来的(尖锐瞬态、低频嗡鸣、节拍夹带),而不是审美选择。The supervisor will ask: isn't this what Endel already does? Three differences, all testable. (1) Endel drives sound from live heart rate but returns no outcome, so a user cannot know whether it worked; here the measured result is the first sentence after the session — the K2 feature the survey classifies as one-dimensional, and must-be inside the GAD-2-positive segment. (2) Endel's only peer-reviewed study is self-funded and measures focus (Haruvi et al. 2022 [43]); this project pre-registers a within-subject crossover with an active control. (3) The neutral-sound spec is reverse-engineered from users' specific complaints — sharp transients, low hum, beat entrainment — rather than chosen aesthetically.
为什么这样做 · Why we did it this way. 最关键的一个决定是用心率而不是 HRV 做实时控制。直觉上 HRV 更"科学"(它是压力指标),但 HealthKit 的 HRV 是机会性采样且可滞后 30 分钟,Oura 只有夜间数据——真做成"实时 HRV 自适应"就是在演示一个不存在的能力。承认这条 API 事实(它也是 A2b 里那条 LR 0.60 的技术证据)让整个系统诚实:能实时的那个量拿去开车,不能实时但更贴近文献的那个量拿去记账。The load-bearing decision is to steer on heart rate rather than HRV. HRV feels more scientific — it is the stress marker — but HealthKit's HRV is opportunistic and can lag 30 minutes, and Oura is sleep-only. Building "live HRV adaptation" would be demonstrating a capability that does not exist. Admitting that API fact (it is also the LR 0.60 technical row inside A2b) is what keeps the system honest: the quantity that can be read live does the steering, and the quantity the literature cares about does the bookkeeping.
10. 怎么证明有效How efficacy will be proved — E0, E1 and what they feed
10.1 识别策略Identification
E0 校准子研究先跑:n ≈ 8,同时佩戴 Apple Watch 与 Polar H10,各 3 节,估计 ICC、MAPE 与回归校准系数 λ = Cov(watch, polar)/Var(watch)。理由是测量误差的两种后果不同:作为结局时,非差异性误差只损失功效、不产生偏倚;作为回归自变量时(例如把基线 HRV 当调节变量)会造成衰减偏倚,必须去衰减(Carroll et al. 2006 [58])。
E0, the calibration sub-study, runs first: n ≈ 8 wearing an Apple Watch and a Polar H10 concurrently for 3 sessions each, estimating ICC, MAPE and the regression-calibration coefficient λ = Cov(watch, polar)/Var(watch). The reason is that measurement error has two different consequences: as an outcome, non-differential error costs power but does not bias; as a regressor (say, baseline HRV as a moderator) it causes attenuation bias and has to be de-attenuated (Carroll et al. 2006 [58]).
E1 是随机化的被试内三条件交叉设计,顺序按拉丁方平衡:T1 自适应中性声音(心率驱动)· T2 自适应真歌舒缓版(参与者自选曲目)· C 主动对照 = 参与者自己常听的一份"放松"歌单,原样播放。对照不是静音——静音会改变期待并把"任何音频有效"和"我们的音频有效"混在一起。每人 3 节 × 15 分钟,不同日、同一时段、间隔 ≥ 24 小时(洗脱),坐姿、只用手机。
E1 is a randomised within-subject three-condition crossover, order-balanced by Latin square: T1 adaptive neutral sound (HR-driven) · T2 an adaptive calm arrangement of a real song the participant chose · C, an active control — a generic "relaxing" playlist they already know, played untouched. The control is not silence: silence changes expectancy and confounds "any audio" with "our audio". Three 15-minute sessions per participant on separate days, same time of day, ≥ 24 h apart (washout), seated, phone only.
app/js/core/research.js),一个测试比对 Python 侧与 App 侧的表是否一致——改动只会让测试失败,不会让分配悄悄错位。Fig. 10 · The E1 session timeline and Latin square. Condition assignment derives from the participant code, and a test compares the Python table with the app's — a change fails the test rather than silently mis-assigning anyone.10.2 威胁与控制Threats and controls
| 威胁 Threat | 方向 Direction | 控制 Control |
|---|---|---|
| 期待 / 安慰剂效应Expectancy / placebo | 抬高 Tinflates T | 主动对照;会话前测两题期待值并作为协变量;分析者对条件标签盲化;三条件都告知为"舒缓音频" |
| 均值回归Regression to the mean | 抬高任何前后差inflates any pre→post drop | 基线作协变量(ANCOVA 形式,Vickers & Altman 2001 [32]);先坐 5 分钟,避免把"刚到达的心跳峰"当基线 |
| 顺序 / 残留 / 学习效应Order, carry-over, learning | 偏倚后面的会话biases later sessions | 拉丁方平衡;Order 与 Session# 作协变量;洗脱 ≥ 24 小时 |
| 时段 / 昼夜节律Time of day / circadian HRV | 加噪,若不平衡则有偏adds noise; biases if unbalanced | 每位参与者固定时段;时段进入协变量 |
| 动作 / 说话 / 咖啡因Movement, speech, caffeine | 心率伪影artefacts in HR | 坐姿方案;加速度掩膜;咖啡因 / 酒精 / 运动 2 小时规则并登记 |
| 霍桑效应Hawthorne / observation | 各臂相同equal across arms | 三条件被观察程度完全相同 |
| 自选样本Self-selection into the pilot | 限制外部效度,不影响内部效度external validity only | 被试内设计;报告招募来源;用 GAD-2 / PSS-4 与问卷样本比对 |
| 创始人—研究者的需求效应Founder-researcher demand effects | 抬高 Tinflates T | 由研究助理按脚本执行会话;创始人不主持、不在未盲化状态下分析 |
| 多重结局Multiple outcomes | 假阳性false positives | 只有一个主要结局(STAI-S);次要(HR、RMSSD)用 Holm 校正 |
| 样本量小Small n | 生理指标功效不足low power for physiology | 贝叶斯分析;每人 3 节增加个体内信息;如实以"可行性"定位 |
功效的诚实说法。配对 t 检验(α = .05 双侧,功效 = .80):d_z = 0.5 → n = 34;0.6 → 24;0.65 → 21;0.8 → 15。Cochrane 的围手术期估计(−5.7 STAI-S,SD ≈ 10 → d ≈ 0.55)对应 n ≈ 28;生理效应(d ≈ 0.4)需要 n ≈ 50。所以:n = 20–30 对 A2a 是可行性信号,对 A2b 功效不足——这句话必须写在论文里(METHODOLOGY §3)。两个补救:每人 3 节(增加个体内信息);用元分析先验做贝叶斯分析,让后验而不是 p 值去撞门槛。
Power, stated honestly. Paired-t power (α = .05 two-sided, power = .80): d_z = 0.5 → n = 34; 0.6 → 24; 0.65 → 21; 0.8 → 15. The Cochrane perioperative estimate (−5.7 STAI-S, SD ≈ 10 → d ≈ 0.55) implies n ≈ 28; physiological effects (d ≈ 0.4) need n ≈ 50. So n = 20–30 is a feasibility signal for A2a and under-powered for A2b — and the dissertation says so (METHODOLOGY §3). Two remedies: three sessions per person for more within-person information, and a Bayesian analysis with the meta-analytic prior so that the posterior, not a p-value, meets the gate.
10.3 管线已建好,并已在合成数据上跑过The pipeline is built and has been exercised — on synthetic data
从 App 导出到决策模型的整条路径已经存在:app/tools/export_to_csv.py(导出 JSON → sessions.csv + windows.csv)→ research/pilot/e1_analysis.py(预注册模型、贝叶斯再分析、似然比、门槛、报告与图)→ 把观测到的 LR 替换回 DECISION_MODEL.md 与 decision.js。28 个测试覆盖了 schema、拉丁方一致性、0.70 质量剔除、顺序推断、注入效应的方向以及 LR / 门槛算术。
The whole path from app export to decision model already exists: app/tools/export_to_csv.py (export JSON → sessions.csv + windows.csv) → research/pilot/e1_analysis.py (the pre-registered model, the Bayesian re-analysis, the LRs, the gates, a report and figures) → the observed LRs replace the pending rows in DECISION_MODEL.md and decision.js. Twenty-eight tests cover the schema, Latin-square agreement, the 0.70 quality exclusion, order inference, the direction of injected effects, and the LR / gate arithmetic.
下面这组数字是捏造的(合成演示)。24 个合成参与者 × 3 节,种子 20260904。它们只证明管线跑得通、预注册规则会按预期触发;它们不是关于 TDMusic 的证据,绝不可作为结果引用。合成注入值取自 DECISION_MODEL 的先验:STAI-S −4.5 分(d_z ≈ 0.5,来自 Cochrane 的 −5.72 / SD ≈ 10)、残差心率 −4.0 bpm、log-RMSSD +0.12(符号朝上——放松抬高 HRV)、T1 − T2 注入为 0。The numbers below are FABRICATED (a synthetic demo). 24 synthetic participants × 3 sessions, seed 20260904. They demonstrate only that the pipeline runs and that the pre-registered rules fire; they are not evidence about TDMusic and must never be quoted as a result. The injected values come from the DECISION_MODEL priors: STAI-S −4.5 points (d_z ≈ 0.5, from Cochrane's −5.72 with SD ≈ 10), residual HR −4.0 bpm, log-RMSSD +0.12 (sign upward — relaxation raises HRV), and T1 − T2 injected as 0.
| 合成演示 SYNTHETIC · 结局 Outcome | 角色 Role | β1(T1 vs C) | 95% CI | p | d_z | β2(T2 vs C) | ICC | n |
|---|---|---|---|---|---|---|---|---|
| STAI-S(STAI-6,20–80) | 主要 primary | −4.10 | [−6.88, −1.33] | 0.004 | −0.49 | −4.98 | 0.232 | 72 |
| 残差心率(最后 5 分钟)Residual HR, last 5 min | 次要(共同主要生理) | −1.52 | [−4.55, 1.51] | 0.325 → 0.552 Holm | −0.20 | −4.53 | 0.102 | 66 |
| log RMSSD | 次要 secondary | 0.07 | [−0.06, 0.20] | 0.276 → 0.552 Holm | 0.29 | 0.14 | 0.064 | 66 |
| 合成演示的门槛输入 → NOT YET:A2a 触发 LR 4.0(0.85 → 0.958),A2b 触发 LR 0.30(0.56 → 0.276),A1a 保持 LR 1.0(0.07 → 0.070);PROCEED 仅卡在 P(A2b) ≥ 0.70 一条上。40 次独立重复(种子 1–40)显示:A2a 在 78% 的运行里触发 LR 4.0(与 n = 24、d_z = 0.5 的 80% 功效目标吻合),A2b 只有 45%——这正是设计文件事先写下的"生理节点功效不足"。The synthetic run reproduces the design's own power story rather than flattering it: the self-report node clears its gate, the physiological node does not, exactly the asymmetry METHODOLOGY §3 predicted in advance. | ||||||||
结果怎么回到决策:report.md §4 打印三条预注册的行——A2a(STAI-S 的 β1)LR 4.0 / 0.30;A2b(残差心率 β1,Holm 后)LR 4.0 / 0.30;A1a(β1 − β2 的对比)LR 3.0 / 1.0 / 0.33。把这三个数替换回信念登记册,§5 的门槛检查就会机械地打印出判定——决策备忘录不需要再作一次判断。
How the result returns to the decision: report.md §4 prints the three pre-registered rows — A2a (β1 on STAI-S) LR 4.0 / 0.30; A2b (β1 on residual HR, after Holm) LR 4.0 / 0.30; A1a (the β1 − β2 contrast) LR 3.0 / 1.0 / 0.33. Substituting those three numbers into the belief register makes the §5 gate check print the verdict mechanically — the decision memo requires no further judgement.
为什么这样做 · Why we did it this way. 先把分析代码写完、在合成数据上跑通,再去招募,这件事看起来是工程洁癖,其实是最强的偏差控制:见到数据之后就没有"分析者自由度"可用了(METHODOLOGY §7)。而且合成演示没有被调成好看的样子——它复现了设计自己预告的弱点(生理节点功效不足)。一份把自己的短板演示出来的管线,比一份漂亮的结果更能说服答辩委员会。Writing the analysis code and running it on synthetic data before recruiting anyone looks like engineering fastidiousness; it is in fact the strongest bias control available, because once real data arrives there are no analyst degrees of freedom left (METHODOLOGY §7). And the synthetic demo was not tuned to flatter: it reproduces the weakness the design predicted for itself. A pipeline that demonstrates its own limitation is more persuasive to an examiner than a pretty result.
11. 给谁用、放什么音乐、为什么Who it is for, what plays, and why
不是"焦虑"这一个人群,而是一组可分开的时刻。焦点小组的 T1 主题里,参与者用自己的便签聚出五个簇(下班后的复盘 · 入睡 · 上台前 · 通知惊跳 · 低度背景紧张),问卷 B3 的时间分布把它们量化(床上 57%、工作 52%、傍晚 33%、深夜 23%);产品把其中四个做成了显式的 moment 标签(app/ARCHITECTURE.md 的 type Moment:急性 / 放松 / 入睡 / 上台前)。
Not one "anxiety" audience but a set of separable moments. In focus-group theme T1 participants clustered their own sticky notes into five groups (rumination after the day, sleep-onset, anticipatory/performance, notification spikes, low-grade background tension); the survey's timing item quantifies them (in bed 57%, at work 52%, evening 33%, night 23%); and the product turns four of them into explicit moment tags (type Moment in app/ARCHITECTURE.md: acute / wind-down / sleep-onset / pre-performance).
选模式这个动作本身就是时刻标签:一次点击同时完成了分群与上下文采集,这既是产品的简化,也是研究的数据(moment tag 进入 bandit 的上下文,也进入 A1a / A1b 的行为证据)。
Choosing a mode is the moment tag: one tap does segmentation and context capture at once — a simplification for the user and a datum for the research (the moment tag becomes bandit context and behavioural evidence on A1a / A1b).
放什么声音,规格从哪来What plays, and where the spec comes from
Settle 的中性床不是"氛围音乐",是一份从用户抱怨里逆推出来的工程规格:无旋律、无节拍、无歌词、无可辨音高中心;频谱重心在一分钟内从约 1.2 kHz 漂到约 600 Hz(越平静越暗);60 Hz 以下无内容(直接回应 I-05 的"低频嗡鸣引起恶心");瞬态不超过床 +6 dB(回应"尖锐刺耳声",问卷里最负的一项 −1.5);响度固定 −16 LUFS;参数每 60 秒按 v0 规则表变化一次。四张自有的床已经在产品里:Arctic Reverberation · Zen Furnace · Fluorescent Hum · Silent Monolith。
The Settle neutral bed is not ambient music but an engineering spec reverse-engineered from complaints: no melody, no beat, no lyrics, no discernible pitch centre; a spectral centroid drifting from ~1.2 kHz to ~600 Hz across a minute (darker as you settle); nothing below 60 Hz (answering I-05's nauseating low hum); transients no more than +6 dB above the bed (answering "sudden / sharp sounds", the most negative item in the survey at −1.5); loudness fixed at −16 LUFS; parameters stepped every 60 s by the v0 rule table. Four owned beds ship today: Arctic Reverberation, Zen Furnace, Fluorescent Hum and Silent Monolith.
Unwind 的真歌规格是:器乐钢琴 / 吉他改编、60–66 bpm、无打击乐、柔和低通、无突变动态。产品负责人的补充要求很具体——曲池需要更多慢速、混响厚重的钢琴。现在的曲库是 69 首 Unwind 曲目,其中 53 首是慢钢琴;9 首 TDMusic 自有 WAV 已按 −16 LUFS 转码入库,其中 5 首进 Unwind(Suburban Nights · CRT Light · Rewind the Years · Sunset Rewind · 钢琴弹唱),另外 4 首就是上面那四张 Settle 中性床。曲目来源可追溯:从 Spotify 官方舒缓歌单出发,因为这些歌单里多是匿名"幽灵艺人",又通过 Wikidata 解析知名艺人的 Spotify ID 取其"热门"榜;每个 ID 都用 Spotify 的 oEmbed 端点校验过。
The Unwind real songs are specified as instrumental piano or guitar arrangements, 60–66 bpm, no percussion, gentle low-pass, no sudden dynamics — with an explicit product request for more slow, reverb-heavy piano. The catalogue today holds 69 Unwind tracks, 53 of them slow piano; nine TDMusic-owned WAVs have been ingested at −16 LUFS, five of them into Unwind (Suburban Nights, CRT Light, Rewind the Years, Sunset Rewind, 钢琴弹唱) and the other four being the Settle beds named above. Provenance is traceable: starting from Spotify's editorial calm playlists, which are dominated by pseudonymous "ghost" artists, well-known artists were resolved to Spotify IDs via Wikidata and drawn from their Popular lists; every id was verified against Spotify's oEmbed endpoint.
| 档位 Tier | 权利 Rights | 怎么播 Playback | 能做多少自适应 How much adaptation | 用在哪 Where |
|---|---|---|---|---|
| P · 参数级Parametric | owned — TDMusic 自有或已确认改编权Owned or rights-cleared | Web Audio(自有文件 / 合成床),两个播放器做 1.5 秒交叉淡入Web Audio; two players for a 1.5 s crossfade | 低通截止、增益、层密度、脉冲等效速度,每 60 秒一步Cutoff, gain, density, pulse-equivalent tempo, stepped every 60 s | Settle 全部床;Unwind 的自有曲目All Settle beds; owned Unwind tracks |
| S · 选曲级Selection | embed-only — 只通过官方播放器Embed-only | Spotify Web Playback SDK(PKCE / Premium)或官方嵌入回退;SoundCloud WidgetSpotify SDK or the official embed fallback; SoundCloud Widget | 只能按心率带选曲与决定下一首更慢 / 保持——嵌入播放器不能改音频Selection and sequencing only: embeds cannot alter audio | Unwind 的第三方曲目Third-party Unwind tracks |
为什么这样做 · Why we did it this way. 把"放什么音乐"写成可验收的声学规格(频段、瞬态、响度、有无音高中心),而不是"温柔舒缓的氛围",有三个后果:焦点小组的盲听刺激可以照规格做出来,因此 S-B 的 10/13 是对规格的检验而不是对品味的检验;AI 引擎有了明确的生成目标,才谈得上 E4 的自动晋升 / 退役;答辩时"你凭什么说你的中性声音更好"有一个可测的答案——它不含那四类被用户点名的东西。Writing "what plays" as an acceptance-testable acoustic spec — bands, transients, loudness, presence of a pitch centre — rather than "gentle soothing ambience" has three consequences: the blind-test stimulus could be built to that spec, so S-B's 10/13 tests the spec rather than taste; the AI engine gets a target concrete enough for E4's automatic promotion and retirement; and "why is your neutral sound better?" has a measurable answer in the viva — it contains none of the four things users named.
12. 我们现在在哪Where the app stands today — 6 Sep 2026
已经做出来的东西What exists
- 产品蓝图 v1(3 Sep):中英对照,十节——人因原则、系统架构、13 阶段用户路径、八屏 + 手表 + 安全屏、数据采集接口矩阵(HealthKit / Oura API v2 / Polar H10 BLE / Web Bluetooth)、曲库两档与标签规范、学术方法 → 产品交互逐项映射、反馈闭环、技术栈与四周计划、v0 控制规则。打开蓝图 →
- 可安装的 App(PWA):https://38.180.150.12.sslip.io/app/(HTTPS,iPhone Safari"添加到主屏幕"即可)。首页(Settle / Unwind、时长、睡眠淡出)→ 预休息 60 秒(研究模式 300 秒)+ 自评 0–10 → 会话 → 20 秒回访 → 结果(相对预休息 / 相对"你的平时"、沉降曲线、信号质量、有 RR 时给 RMSSD)→ 历史 / 导出 / 删除 → 设置;安全屏(HR > 130 持续 10 秒或用户点"我需要接地");Service Worker 离线壳。
- 研究模式(v0.2.0 起):参与者编号 → 拉丁方条件 T1 / T2 / C、5 分钟预休息、STAI-6 前后测、期待值两题、导出含条件字段。这不是演示功能——它就是 E1 的现场工具。
- 推荐与播放(v0.5.x):按标签 × 时刻 × 心率带 × 喜好 × 新鲜度 × 质量评分,加权洗牌,过渡约束,"为什么是这首";Spotify SDK(PKCE)/ 嵌入回退 / 自有交叉淡入 / SoundCloud;Like · Slower · Next · Shuffle 随时可用。
- 测试(TODO §12 在 v0.4.0 时点的计数,此后仍在增加):单元 70 · 端到端 11 · 音频离线渲染检查 6 · 场景引擎 14;另有 WebKit(iPhone 15 模拟)6/6 场景通过(§16)。核心逻辑是纯函数,测试即规格。
- 自有音乐入库:9 首自有 WAV 经
ingest_music.py转成 −16 LUFS AAC 并做声学分析(32 MB);Settle 床 Arctic Reverberation / Zen Furnace / Fluorescent Hum / Silent Monolith,Unwind 自有曲目 Suburban Nights / CRT Light / Rewind the Years / Sunset Rewind 等。 - 试点数据管线:
research/pilot/完成,28 个测试;导出 → CSV → 预注册模型 → 报告 / 图 / 门槛输入。
App Store 与 TestFlight 的真实状态The real store status
com.tdmusic.lilt · App「Lilt Adaptive Sound」。构建 2 已处理完成(VALID);重设计 v2 是构建 5;v0.5.0 对应构建 6。CI 用 macos-26 + Xcode 26.2(App Store Connect 要求 iOS 26 SDK)。为什么这样做 · Why we did it this way. 先做 PWA 再包 iOS,不是省事,而是被能力边界决定的:浏览器已经能做音频合成、嵌入播放、蓝牙胸带、本地存储和主屏安装;它不能拿到 Apple Watch 的实时心率、不能在 iOS 上稳定后台播放、不能驱动触觉。所以 v0.1–v0.2 是 PWA(几天内就能给试点用),v0.3 用 Capacitor 包成 TestFlight 拿 HealthKit 历史,v0.4 才加 watchOS 伴侣拿实时心率——一套 App 逻辑代码贯穿始终。把"什么时候必须转原生"写成一张能力矩阵,比先做原生再发现用不上要便宜得多。Shipping a PWA first and wrapping iOS later is not a shortcut; it follows from a capability boundary. The browser already gives audio synthesis, embedded players, Bluetooth chest straps, local storage and home-screen install; it cannot reach Apple Watch heart rate, cannot play audio reliably in the background on iOS, and cannot drive haptics. So v0.1–v0.2 are a PWA usable in the pilot within days, v0.3 wraps the same code in Capacitor for TestFlight and HealthKit history, and v0.4 adds a watchOS companion for live heart rate — one app-logic codebase throughout. Writing "when do we have to go native" as a capability matrix is far cheaper than going native first and discovering it was unnecessary.
13. 下一步,以及每一步是为了什么What comes next, and which belief each step is meant to move
路线图不按"功能"排,按它买到多少信息排。每个测试都是一次信息采购:成本、时长与预期的后验移动量都写在 DECISION_MODEL.md §4,排序遵循"每单位成本的期望信息量"——法务核查最便宜也最决定性(LR 8 / 0.15),所以先跑;然后是问卷;最后是最贵的试点。这就是设计思维阶段的实物期权视角(Bowman & Hurry 1993 [17];McGrath 1999 [16])。
The roadmap is ordered not by feature but by how much information each step buys. Every test is a purchase of information whose cost, duration and expected posterior shift are listed in DECISION_MODEL.md §4, and the ordering follows expected information per unit cost — the legal check is the cheapest and most decisive (LR 8 / 0.15) so it runs first, then the survey, then the expensive pilot. That is the real-options view of the design-thinking phase (Bowman & Hurry 1993 [17]; McGrath 1999 [16]).
| 下一步 Step | 它买到什么信息 What it buys | 推动哪个信念 Belief it moves | 时间 When |
|---|---|---|---|
| 法务核查 ≥ 50 首曲目的改编 / 衍生权Legal check on adaptation rights for ≥ 50 tracks | 唯一一条被评为"决定性"的待补证据;母带权与词曲权分别确认The only pending row graded decisive; masters and publishing confirmed separately | A5 · LR 8.0 / 0.15 | 进行中 → 12 Sep |
| Landing A/B 上线看转化Landing page A/B live | 行为证据(不是意向):访客 → 候补名单 ≥ 5%Behavioural, not stated: visitor → waitlist ≥ 5% | A3 · A4 · 各 LR 2.5 / 0.4 | 进行中 → 20 Sep |
| E0 校准子研究E0 calibration sub-study | 腕表相对胸带的可靠度 λ;决定 HRV 能不能当回归自变量用Wrist-vs-strap reliability; decides whether HRV can be a regressor at all | A2b 的测量前提 | 25 Aug – 6 Sep |
| E1 效能试点E1 efficacy pilot | 整个阶段最强的一条证据,也是 PROCEED 的必要条件The strongest evidence in the phase and a necessary condition for PROCEED | A2a 4.0 / 0.30 · A2b 4.0 / 0.30 · A1a 3.0 / 1.0 / 0.33 | 1 – 25 Sep |
| P5 · watchOS 伴侣 + Polar 配对watchOS companion, Polar pairing | 把实时心率与手腕触觉呼吸带进试点;HKWorkoutSession 必须从手表端启动Brings live HR and haptic breathing into the pilot | 让 A2 的测量成为可能 | 11 – 21 Sep |
| 贝叶斯更新 → 门槛判定Bayesian update → gate verdict | 把观测 LR 替换进登记册;门槛检查机械打印Observed LRs replace the pending rows; the gate check prints itself | 全部八个 | 25 Sep – 2 Oct |
| 决策备忘录 + 督导 #2Decision memo and supervision #2 | 设计思维以一个被记录的决定结束,而不是一次路演Design thinking ends with a documented decision, not a pitch | 概念选择 C1-neutral / C1-real / C2 / C3 | 10 月初 early Oct |
| P6 · Oura + 基线模型 + 后端 + E2Oura, baseline model, backend, E2 | 把"你的平时"从 3 次会话的中位数升级为 14 天分层模型;开始随机化变体Upgrades "your usual" from a 3-session median to a 14-day hierarchical model; starts randomising variants | A6(留存,观察项)· E2 学习 | 22 Sep – 10 月中 |
| 论文写作Dissertation write-up | 方法章 · 共情与定义 · 构思与原型 · 测试与结果 · 讨论与路线图Methods · Empathize/Define · Ideate/Prototype · Test/Results · Discussion | — | 9 – 11 月,日期待定 |
| 阶段后(仅在 PROCEED 时)Post-phase, only on PROCEED | E2 微实验 → E3 情境 bandit → E4 自生成曲库;MVP 范围界定E2 → E3 → E4, and MVP scoping | A6 · 产品规模化 | 2027 |
为什么这样做 · Why we did it this way. 排序不是按"哪个做起来顺手",而是按每一步能让哪个概率移动多少。法务核查排第一,因为它一周多一点、成本极低,却是唯一带 LR 8 的决定性证据——如果它是否定的(0.15),真歌那条路当场基本被封死,也就没必要再为它花任何工程时间。这就是把"实物期权"从一个漂亮说法变成一张执行顺序表。The ordering is not by convenience but by how far each step can move which probability. The legal check comes first because it takes a little over a week, costs almost nothing, and is the only decisive row (LR 8) — and if it comes back negative (0.15) the real-music route is effectively closed on the spot, so no further engineering time should be spent on it. That is what turns "real options" from a pleasant phrase into an execution order.
14. 评估方法一览Evaluation methods at a glance — construct, instrument, statistic, threshold, decision node
| 构念 Construct | 量表 / 方法 Instrument | 统计量 Statistic | 预注册阈值 Threshold | 决策节点 Node |
|---|---|---|---|---|
| 焦虑症状(筛查)Anxiety symptoms (screener) | GAD-2(Kroenke et al. 2007 [18]),逐字 | % ≥ 3;分段交叉表 | —(描述人群,不进 LR) | 人物志 · 分层 |
| 知觉压力Perceived stress | PSS-4(Cohen & Williamson 1988 [19]),逐字 | 均值 · α · 与 GAD-2 的聚合相关 r | — | 需求强度 |
| 未被满足的结果Under-served outcomes | Ulwick ODI(2002 [23]),条目来自访谈 JTBD(Christensen et al. 2016 [10]) | 机会分 = I + max(I − S, 0);排序 + bootstrap CI | 机会分 ≥ 6 = 未被满足 | POV / HMW 优先级 |
| 分状态内容偏好Content preference by state | 被试内情景强制选择,顺序随机 | McNemar(1947 [31]),小样本用精确检验 | 急性选真歌 ≥ 50% → LR 3 · ≤ 35% → 0.33;放松 ≥ 50% → 2 · ≤ 35% → 0.5 | A1a · A1b |
| 技术接受Acceptance | TAM PU / PEOU / BI(Davis 1989 [24];Venkatesh & Davis 2000 [25]) | α ≥ .70;BI top-2;OLS 标准化 β;n ≥ 150 时 PLS-SEM 稳健性 | BI ≥ 40% 或拥有者 ≥ 55% → LR 1.8 | A3 |
| 功能价值Feature value | Kano(Kano et al. 1984 [26])+ Better / Worse(Berger et al. 1993 [27]) | 评价表分类;分段分类 | K1 / K2 为魅力或一维 → LR 1.5 | A3 · MVP 范围 |
| 相对优势 / 可试用性Relative advantage, trialability | DOI(Rogers 2003 [28]) | top-2 占比 | ≥ 50% → LR 1.5 | A3 |
| 价格敏感度Price sensitivity | Van Westendorp PSM(1976 [29]) | OPP · IPP · 区间 [PMC, PME],分币种,逐人强制单调 | 区间含 $6.99 → LR 1.5;OPP < $4 → 0.7 | A4 |
| 购买意向Purchase intent | 指定价格 5 点量表 | top-2 box × 0.5 假设性偏差折扣(Morwitz et al. 2007 [30]) | 美/欧 ≥ 30% → LR 2 · ≤ 15% → 0.5 | A4 |
| 行为化支付意愿Behavioural WTP proxy | Landing 页访客 → 候补名单转化 | 转化率,按定位分组 | ≥ 5% → LR 2.5 · 否则 0.4 | A3 · A4 |
| 状态焦虑(试点主要结局)State anxiety — pilot primary | STAI-S(Spielberger 1983 [20]);短式 STAI-6(Marteau & Bekker 1992 [21]) | ANCOVA 形式混合模型的 β1(Vickers & Altman 2001 [32]);d_z;元分析先验下的 P(β1 < 0) | p < .05 且 d ≥ .3 且方向正确 → LR 4.0;零结果或反向 → 0.30 | A2a |
| 生理状态(试点次要结局)Physiological state — pilot secondary | 心率(bpm)与 RMSSD / SDNN(ms),Polar H10 子样本作参照 | 同一模型跑残差化后的量;log 变换 RMSSD;两个次要结局用 Holm 校正 | 同上 LR 4.0 / 0.30 | A2b |
| 期待效应(试点协变量)Expectancy — pilot covariate | 两题("我预计这一段会让我平静" 1–7 + 可信度,Devilly & Borkovec 2000 [22] 短式) | 作为 β5 进入模型 | —(控制变量) | 安慰剂控制 |
| 可穿戴测量误差Wearable measurement error | E0:Apple Watch 与 Polar H10 同步记录;RR 清洗按 Task Force ESC/NASPE 1996 [51]、Tarvainen et al. 2002 [52]、Lipponen & Tarvainen 2019 [53];指标定义见 Shaffer & Ginsberg 2017 [54] | ICC · MAPE · 可靠度比 λ;作为自变量时去衰减(Carroll et al. 2006 [58]) | —(决定 HRV 能否作回归自变量) | A2b 的前提 |
| 质性主题Qualitative themes | 反身性主题分析(Braun & Clarke 2006 [5]);焦点小组(Krueger & Casey 2015 [7];Morgan 1997 [8]);网络民族志(Kozinets 2020 [6]);关键事件法(Flanagan 1954 [9]) | 双编码 20%,Cohen's κ ≥ .70;饱和判据 = 不再产生新码 | 质量收缩 k = 0.35–0.5,相关来源再 × 0.6;只能移动信念,不能跨门槛 | A1a · A1b · A3 · A4 |
| 学习与个性化(阶段后)Learning and personalisation, post-phase | 分层贝叶斯结局模型(Gelman & Hill 2007 [57]);情境 bandit(Agrawal & Goyal 2013 [55]);离策略评估(Dudík, Langford & Li 2011 [56]) | 后验均值与区间;IPS / 双重稳健估计;在线后悔值 | 臂比默认差 > 0.3 SD → 对该用户暂停;探索 ≤ 15% | E2 · E3 · E4 |
为什么这样做 · Why we did it this way. 把整份论文的测量放进一张表,是为了让"我们用了很多方法"变成"每个方法各自负责一件事,并且各自有一个事先写死的判据"。这张表也是答辩时最好用的索引:任何一个"这个数字怎么来的"的问题,都能在同一行里读到量表、统计量、阈值和它推动的信念。Putting the whole dissertation's measurement into one table turns "we used many methods" into "each method has one job and one criterion fixed in advance". It is also the most useful index in the viva: any "where did this number come from?" is answered on a single row — instrument, statistic, threshold and the belief it moves.
15. 参考文献References
仅列出本项目文档中已经引用的著作;本页不引入任何新文献。条目按主题分组编号。Only works already cited in this project's own documents appear below; this page introduces no new literature. Entries are numbered within thematic groups.
方法与框架Method and frame
- Brown, T. (2008). Design thinking.
- Liedtka, J. (2018). On the effectiveness of design thinking. Journal of Product Innovation Management.
- Coghlan, D., & Brannick, T. (2019). Doing Action Research in Your Own Organization.
- Creswell, J. W., & Plano Clark, V. L. (2018). Designing and Conducting Mixed Methods Research.
- Braun, V., & Clarke, V. (2006). Qualitative Research in Psychology, 3, 77. (Also 2019, reflexive thematic analysis.)
- Kozinets, R. V. (2020). Netnography.
- Krueger, R. A., & Casey, M. A. (2015). Focus Groups.
- Morgan, D. L. (1997). Focus groups as qualitative research.
- Flanagan, J. C. (1954). The critical incident technique. Psychological Bulletin, 51(4), 327–358.
- Christensen, C. M., et al. (2016). Jobs to be done — current alternatives as the competitive set.
决策分析与创新管理Decision analytics and innovation management
- Fairfield, T., & Charman, A. E. (2017). Explicit Bayesian analysis for process tracing: guidelines, opportunities, and caveats. Political Analysis, 25(3), 363–380.
- Humphreys, M., & Jacobs, A. M. (2015). Mixing methods: a Bayesian approach. American Political Science Review, 109(4), 653–673.
- GRADE Working Group (Guyatt, G. H., et al., 2008). BMJ, 336, 924.
- Ries, E. (2011). The Lean Startup.
- Bland, D. J., & Osterwalder, A. (2019). Testing Business Ideas.
- McGrath, R. G. (1999). Academy of Management Review, 24, 13.
- Bowman, E. H., & Hurry, D. (1993). Real options and organisational learning.
量表与调查方法Instruments and survey methods
- Kroenke, K., et al. (2007). GAD-2 anxiety screener.
- Cohen, S., & Williamson, G. (1988). PSS-4 perceived stress scale.
- Spielberger, C. D. (1983). STAI manual.
- Marteau, T. M., & Bekker, H. (1992). STAI six-item short form. British Journal of Clinical Psychology, 31, 301.
- Devilly, G. J., & Borkovec, T. D. (2000). Credibility / expectancy. Journal of Behavior Therapy and Experimental Psychiatry, 31, 73.
- Ulwick, A. W. (2002). Outcome-driven innovation. Harvard Business Review.
- Davis, F. D. (1989). Technology acceptance model.
- Venkatesh, V., & Davis, F. D. (2000). Extension of the technology acceptance model.
- Kano, N., et al. (1984). Attractive quality and must-be quality.
- Berger, C., et al. (1993). Better / Worse coefficients for Kano categories.
- Rogers, E. M. (2003). Diffusion of Innovations, 5th ed.
- van Westendorp, P. (1976). Price sensitivity meter.
- Morwitz, V. G., Steckel, J. H., & Gupta, A. (2007). When do purchase intentions predict sales? International Journal of Forecasting, 23, 347–364.
- McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions.
- Vickers, A. J., & Altman, D. G. (2001). Analysing controlled trials with baseline and follow-up measurements. BMJ, 323, 1123.
临床与生理证据Clinical and physiological evidence
- Bradt, J., Dileo, C., & Shim, M. (2013). Music interventions for preoperative anxiety. Cochrane Database of Systematic Reviews (26 RCTs, N = 2,051; −5.72 STAI-S; GRADE low).
- de Witte, M., et al. (2020). Effects of music interventions on stress-related outcomes (104 RCTs, N = 9,617; psychological d = .545, physiological d = .380, heart rate d = .456).
- Harney, C., et al. (2023). Musicae Scientiae — music listening and anxiety (21 controlled studies, d = −0.77).
- Lassner, et al. (2024/2025). Delivery modes compared (SMD = 0.47, k = 23, n = 1,084; GRADE very low; effect not maintained at follow-up).
- Panteleeva, Y., et al. (2018). Music-induced anxiety reduction: self-report significant (d = −.30), psychophysiological not significant.
- Bernardi, L., Porta, C., & Sleight, P. (2006). Cardiovascular, cerebrovascular and respiratory changes induced by music. Heart, 92, 445.
- Mitrovic & Paladin (2026). Tempo–HRV relationship, review. Frontiers in Cardiovascular Medicine.
- Goessl, V. C., Curtiss, J., & Hofmann, S. G. (2017). HRV-biofeedback and stress / anxiety, meta-analysis (g = 0.81). Psychological Medicine.
- Creswell, J. D., & Goldberg, S. B. (2025). Meditation-app engagement: ~4.7% 30-day retention; 1–4 lifetime sessions.
- O'Daffer, A., et al. (2022). Efficacy and conflicts of interest in commercial mindfulness-app trials.
- Haruvi, A., et al. (2022). Effects of personalised soundscapes on focus. Frontiers in Computational Neuroscience (industry-funded; measures focus, not anxiety).
- Nigg, J. T., et al. (2024). Systematic review finding no controlled studies of brown noise and ADHD.
- Woods, K. J. P., et al. (2024). Amplitude-modulated music and attention. Communications Biology.
- Pedersen, S. K. A., et al. (2017). Individualised music and agitation in dementia, meta-analysis (d = 0.61, 12 RCTs).
- McCreedy, E. M., et al. (2021). Pragmatic RCT of personalised music in nursing homes (N = 463, 54 homes): no significant effect on the primary agitation outcome.
- Mendes, et al. (2024). Background music and attentional networks in ADHD, case-control.
- World Health Organization (2023). Anxiety disorders fact sheet (359 million people in 2021; 27.6% receiving treatment).
- Huang, Y., et al. (2019). Prevalence of mental disorders in China (anxiety 5.0% 12-month, 7.6% lifetime).
信号处理与学习Signal processing and learning
- Task Force of the ESC and NASPE (1996). Heart rate variability standards. Circulation, 93, 1043.
- Tarvainen, M. P., et al. (2002). Smoothness-priors detrending. IEEE Transactions on Biomedical Engineering, 49, 172.
- Lipponen, J. A., & Tarvainen, M. P. (2019). RR-interval artefact correction. Journal of Medical Engineering & Technology, 43, 173.
- Shaffer, F., & Ginsberg, J. P. (2017). An overview of heart rate variability metrics and norms. Frontiers in Public Health, 5, 258.
- Agrawal, S., & Goyal, N. (2013). Thompson sampling for contextual bandits.
- Dudík, M., Langford, J., & Li, L. (2011). Doubly robust policy evaluation and learning.
- Gelman, A., & Hill, J. (2007). Data Analysis Using Regression and Multilevel / Hierarchical Models.
- Carroll, R. J., et al. (2006). Measurement Error in Nonlinear Models.
人因工程Human factors
- Easterbrook, J. A. (1959). The effect of emotion on cue utilisation and the organisation of behaviour.
- Weiser, M., & Brown, J. S. (1996). The coming age of calm technology.
在正文中按名字提到、但项目文档未给出年份或出处的框架与定律,不进入上面的编号列表:Ansoff 矩阵 · McKinsey 三层面 · O'Reilly & Tushman 的双元性 · Sweller 的认知负荷 · Hick 定律 · Fitts 定律 · Nielsen 启发式 · Deci & Ryan 的自我决定理论 · Efraimidis–Spirakis 加权抽样。Frameworks and laws named in the text but for which the project's own documents give no year or venue are deliberately left out of the numbered list: the Ansoff matrix, McKinsey's Three Horizons, O'Reilly & Tushman's ambidexterity, Sweller's cognitive load, Hick's law, Fitts's law, Nielsen's heuristics, Deci & Ryan's self-determination theory, and Efraimidis–Spirakis weighted sampling.