From "uses validated instruments"
to "has evidence."
The honest path forward, in three stages. Using UCLA-3, PHQ-2/9, and GAD-7 is table stakes — the instruments are validated, but our conversational administration of them, and any outcome from the program, are not.
This is the study program that closes that gap: first prove the measurement holds, then look for a signal in an uncontrolled pilot, then — and only then — test whether it works in a randomized controlled trial. Each stage earns exactly one honest claim and nothing more.
Three stages. Each unlocks one claim.
Evidence is earned in order. Skipping a stage does not buy the claim — it just makes the claim wrong. The honesty rule: never state a claim a completed study cannot support.
Does conversational scoring equal the validated screen?
Within-subjects, counterbalanced. The same older adults take each screen two ways — standardized self-report vs. the companion in conversation — and we test agreement (ICC, Bland–Altman, weighted κ) and equivalence (TOST) against pre-set margins.
Does loneliness/mood move — and can we run this safely?
~40 isolated older adults, 12 weeks, pre/post UCLA-3 + PHQ-9 with the Reliable Change Index, plus real engagement and safety data. No control group — by design.
Does the companion actually cause improvement vs. control?
Registered, randomized, controlled — ideally against an attention-matched control, because for a companionship product the honest question is "better than an equal amount of contact," not "better than nothing." Blinded outcome assessment; powered from the pilot's effect size.
Four IRB-grade draft documents.
Each is a full draft protocol written to the standard an IRB and a payer's clinical team would expect: objectives, design, sample-size justification, analysis plan, safety provisions, and honest limitations.
Instrument-Fidelity Study
Answers the #1 objection: a score obtained in conversation is not automatically the validated score. This proves — or disproves — equivalence.
Single-Arm Outcomes Pilot
The program's first honest data point: does loneliness and mood move over 12 weeks, and can we operate safely? Uncontrolled by design.
Randomized Controlled Trial — Outline
The eventual gold standard for a causal efficacy claim. Sequenced last, powered from the pilot. Confronts the hard attention/placebo problem head-on.
Human-Subjects Safety Plan
Governs all three studies. Suicidality escalation, adverse-event reporting, elder-abuse disclosure, consent & capacity, and data safety.
Crisis care always overrides research.
These studies enroll isolated older adults and screen for depression and suicidal ideation. Safety is not an afterthought — it is the precondition for running at all. In every study, including the pure measurement study, a positive suicidality signal triggers the standing crisis pathway immediately, independent of any research procedure.
PHQ-9 Item 9 + Suicidality
Item 9 > 0 or any expressed suicidal ideation → immediate 988 resources surfaced, real-time clinician alert, audit log. Imminent risk → emergency services per clinical director SOP.
AE / SAE Definitions & Reporting
Adverse events and serious adverse events defined per protocol. IRB reporting timelines followed. All AEs logged regardless of relatedness to the intervention.
DSMB — Staged
Pilot: mini-DSMB with pre-specified stopping rules. RCT: full independent Data Safety Monitoring Board with interim analyses and formal stopping rules.
Capacity-to-Consent Screen
All older-adult participants receive a capacity-to-consent evaluation prior to enrollment. Legally authorized representatives may provide consent where capacity is limited, per IRB approval.
Elder-Abuse Disclosure
Study staff are mandatory reporters. Any suspected elder abuse or neglect disclosed during participation is reported per applicable state law, independent of confidentiality provisions.
HIPAA Data Safety
All participant health information handled under HIPAA. De-identified data only for analysis. IRB-approved data security plan; breach notification per regulatory requirements.
The honest ceiling of each stage.
These are not caveats buried in footnotes. They are the actual epistemic bounds of a responsibly-staged evidence program.
- No results exist yet. Nothing here has been run; we claim no outcomes.
- Fidelity ≠ benefit. Proving conversational scores match validated ones says nothing about whether the program helps.
- The pilot cannot prove causation. With no control arm, change may be regression to the mean, natural course, seasonality, attention, or measurement reactivity.
- Companionship placebo is hard. Much of any benefit may be non-specific attention; only an attention-controlled RCT isolates the companion's specific effect.
- Participants cannot be blinded. Behavioral-trial limit; we blind assessment and analysis instead, and say so.
- Findings are build-specific. Any change to the production model or prompt re-opens fidelity — results are pinned to a build hash.
- Scope is bounded. English-language, capacity-intact older adults; other languages and advanced cognitive impairment require their own studies.
Trust & Evidence Center — full compliance & evidence status →