The evidence, honestly

Validated instruments. Unproven product. A roadmap to prove it.
A clinician-, IRB-, and payer-facing dossier covering instrument psychometrics, the honest strength of the AI-companion evidence base, and how we intend to build from rigorous screening toward a controlled efficacy trial.

Core stance: UCLA-3, PHQ-2/9, and GAD-7 are peer-reviewed, well-validated screeners with established psychometrics in older adults — that foundation is strong. It is not evidence that this product reduces loneliness or treats depression. The "instruments validated; product not" distinction is load-bearing throughout this dossier, and we never overstate it.

Our instruments are validated. Our product is not.

UCLA-3, PHQ-2/PHQ-9, and GAD-7 are well-established screeners with peer-reviewed psychometrics in older adults. That is a strong, defensible foundation for screening, monitoring, and routing to humans.

It is not evidence that Boojee Companion Care reduces loneliness or treats depression. No product like ours has demonstrated efficacy in a controlled trial, and the most-quoted consumer figures in this space (e.g., a "95% reduction in loneliness") come from non-randomized, self-report, response-biased data. Our path from validated instruments to a validated product runs through an IRB-reviewed controlled study — described in full below.

Instrument Psychometrics

Screening instrument profiles

Each instrument is public-domain / free-to-use, delivered using its exact validated item stems and response anchors. Each profile below covers derivation, scoring range, older-adult cutoffs, sensitivity/specificity where established, reliable-change / MID thresholds, and honest caveats. Full derivation in instrument-psychometrics.md.

UCLA-3  Three-Item Loneliness Scale Validated screener

Hughes, Waite, Hawkley & Cacioppo (2004), Research on Aging 26(6):655–672. Derived from the R-UCLA-20 (Russell 1996).

Items
Lack companionship / feel left out / feel isolated. Anchors: 1 hardly ever · 2 some of the time · 3 often.
Score range
Total 3–9 (higher = lonelier). Field convention: ≥6 = probable/meaningful loneliness (a convention, not a diagnostic cut).
Older adults
Built & tested in ~2,100 middle-aged/older adults (HRS); internal consistency α ≈ 0.72; acceptable concurrent & discriminant validity.
Sens / Spec
Loneliness is self-reported, not a diagnosis — no formal sensitivity/specificity data; only 7 possible scores yields coarse measurement.
Reliable change
No established MID or reliable-change threshold — report distributions, do not invent thresholds.
Admin mode
Designed for telephone/oral delivery — the one instrument here purpose-built for voice.
Limits & caveats
Floor/ceiling effects and only 7 possible scores produce weak change-resolution. Oral-script validation licenses fixed stems — not conversational paraphrase. Repeated measurement autocorrelation risk.
PHQ-2 & PHQ-9  Depression Validated screener

Kroenke, Spitzer & Williams — PHQ-9 (2001, J Gen Intern Med 16:606–613); PHQ-2 (2003, Med Care 41:1284–1292).

Score range
PHQ-9 0–27, bands 5/10/15/20 (mild→severe). PHQ-2 0–6; ≥3 → administer PHQ-9.
Sens / Spec
PHQ-9 ≥10: sensitivity 88%, specificity 88% for major depression (general/primary care).
Older-adult cutoff
Optimal cut runs lower in elders: ≥5 gave sens 100% / spec 81%; ≥10 spec 95% but sens fell to 71% (under-detects). Treat the 5–9 "mild" band as review-worthy, not reassurance.
Reliable change / MID
MID ≈5 points; Jacobson–Truax reliable change ≥6-point reduction; validated for treatment monitoring.
Item 9 (SI)
Safety trigger only — routes to human + 988. NOT a suicide-risk assessment; low PPV; a "0" does not rule out risk.
Limits & caveats
Somatic items overlap with normal aging/illness (score inflation). Accuracy degrades with cognitive impairment. The 2-week recall window makes frequent re-administration autocorrelated (practice/response-set effects).
GAD-7  Anxiety Validated screener

Spitzer, Kroenke, Williams & Löwe (2006), Arch Intern Med 166(10):1092–1097.

Score range
0–21, bands 5/10/15; screen-positive ≥10.
Sens / Spec
≥10: sensitivity 89%, specificity 82% for GAD (n=965); mostly unidimensional, α ≈ 0.89–0.92.
Older adults
Acceptable, widely used — less geriatric-specific validation than PHQ-9; anxiety/depression overlap lowers discriminant precision.
Reliable change / MID
No clean MID or reliable-change anchor analogous to PHQ-9 — treat trends as directional signal, not calibrated change.
Limits & caveats
Somatic items (heart racing, restlessness) overlap with cardiac/respiratory disease common in elders. No established individual-change threshold — do not invent thresholds.
Conversational Administration

The conversational-administration validation gap

Every psychometric number above was established under a standardized administration mode — paper/computer self-report for PHQ/GAD, a fixed telephone script for UCLA-3. Our companion delivers items conversationally, in dialogue. We preserve the verbatim stems and anchors, but conversational delivery is a genuine deviation that can move scores: rapport and demand effects, any paraphrase or reordering, voice pacing, and the fact that a warmly conversing AI is not a neutral administrator.

What the literature says — honestly, both directions. Early work is encouraging: a GPT-4o voice chatbot administering the PHQ-9 showed high concordance with self-administration (ICC ≈ 0.91, median absolute difference ≈ 1 point, no systematic bias), and paper/computer/mobile equivalence studies generally find rough agreement. But those studies are small, recent, not peer-reviewed at the level of our other citations, not in older adults, and not our product. Interview modes have also shown lower reliability (α ≈ 0.80) than self-report (α ≈ 0.87) in some comparisons — mode can matter.

Our position: conversational administration is plausibly near-equivalent, but it is not validated for our population and mode, and we will not claim equivalence until we run a within-subject concordance/fidelity study (ICC, mean bias, Bland–Altman limits of agreement) in older adults — reported before any efficacy claim.

Evidence Base

Does AI / tech companionship actually work?

GRADE-style read of the real literature. The problem (elder loneliness as a health risk) is strongly evidenced. Specific companion interventions range from a modest RCT signal (PARO for depression in dementia) to mixed-to-null for ICT reducing loneliness. Full review in evidence-base.md.

High — multiple consistent RCTs or systematic reviews Low–Mod — small RCTs, short follow-up, or mixed results Low → Null — null meta-analysis or significant design limitations Very Low / Untested — no RCT, response-biased self-report, or no data
Claim
Strength
Evidence & verdict
Loneliness / isolation harms elder health
High
Established. NASEM 2020 + Holt-Lunstad: isolation's all-cause mortality risk rivals smoking/obesity; ~50% higher dementia risk. Justifies screening & connection — not any specific product.
PARO robot reduces depression in dementia
Low–Mod
RCT meta-analysis: significant depression effect (SMD ≈ −0.42), may ease loneliness; but small, mostly institutional/dementia samples, short follow-up, QoL effects inconsistent. A tactile robot pet — not conversational AI.
CBT chatbots reduce short-term depression/anxiety
Low
Woebot RCT (n=70, ages 18–28, 2 weeks) cut PHQ-9 vs psychoeducation — but young adults, tiny, brief, subclinical, and not consistently replicated. Not older adults; not loneliness.
ICT / tech reduces loneliness in older adults
Low → Null
RCT meta-analysis (6 RCTs, 391 pts): pooled SMD = −0.08 (95% CI −0.33 to 0.17, ns) — "little to no difference" vs control. No harm found. This is the category we most resemble; the honest read is the burden of proof is unmet.
Consumer AI companion (ElliQ-type) reduces loneliness
Very Low
Not an RCT. ElliQ (n=173 customer survey, 62% response; NY pilot n=107) reports the "95% reduction" figure from self-report with response/selection bias. The authors themselves say RCTs are still needed. Satisfaction data, not efficacy.
Boojee Companion Care reduces loneliness / depression
Untested
No evidence yet. Requires our own IRB-reviewed controlled trial before any efficacy claim. Instrument validity does not transfer to product efficacy.

A compelling companion could displace human contact rather than bridge to it; friendly-companion demand bias can inflate our own self-report metrics (the very flaw that makes ElliQ-type numbers untrustworthy); and LLM companions can mishandle crisis disclosures — so crisis/item-9 handling must be deterministic and human-routed, never left to generative judgment.

Measurement & Study Roadmap

How we measure outcomes honestly

Endpoints, meaningful-change thresholds, confounds, and what each study design can and cannot claim. Full plan in measurement-plan.md.

Primary endpoints
Efficacy outcomes
  • UCLA-3 change — loneliness
  • PHQ-9 change — depression
Secondary endpoints
Supporting & safety
  • GAD-7 — directional anxiety signal
  • Engagement — process measure only, never efficacy
  • Item-9 events, crisis-escalation completion
  • Score deterioration — safety signal
Meaningful change thresholds
When change counts
  • PHQ-9 MID: ≈5 points
  • PHQ-9 reliable change (Jacobson–Truax): ≥6-pt reduction
  • UCLA-3 & GAD-7: no established individual threshold — report distributions only
Why pre/post is not proof
Confound control
  • Single-arm improvement fully explained by regression to mean
  • Natural history + demand bias + attention effects
  • Pre/post = feasibility & safety signal — nothing more

Study roadmap — what each stage unlocks

Full protocol at /companion-care/science/studies/ (fidelity → pilot → RCT).

Stage 1
Fidelity Study
Within-subject concordance: conversational vs. standardized delivery in older adults. Metrics: ICC, mean bias, Bland–Altman limits of agreement.
Unlocks claim Conversational delivery is equivalent to validated mode
Stage 2
IRB Feasibility
Single-arm, IRB-reviewed; labeled non-causal. Targets: recruitment rate, retention, engagement, safety-event handling.
Unlocks claim Feasible, safe to deploy at scale
Stage 3
Controlled Trial
Randomized, attention-controlled, blinded outcome assessment. Preregistered, intention-to-treat, per-participant responder analysis, all endpoints reported.
Unlocks claim Efficacy — safe to state to payers, public, or marketing
Claim Boundaries

What we can and cannot claim

We can honestly say

  • Loneliness & isolation are serious, well-evidenced health risks in older adults.
  • We use peer-reviewed, validated screening instruments (UCLA-3, PHQ-2/9, GAD-7).
  • We screen, monitor, and route positive findings to humans — a defensible measurement-based-care process.
  • Evidence for some companion technologies is promising but preliminary.

We must not claim

  • That the product "reduces loneliness by X%" — the ElliQ "95%" figure is the exact claim to avoid.
  • That it "improves mood," "treats depression," or is "clinically proven."
  • That validated instruments make the product validated — they do not.
  • Any efficacy claim before an IRB-reviewed controlled trial.
Appendix

Full analyses

Key References

Peer-reviewed sources

• Hughes, Waite, Hawkley & Cacioppo (2004), UCLA-3, Research on Agingjournals.sagepub.com · PMC2394670
• Kroenke, Spitzer & Williams (2001), PHQ-9, J Gen Intern MedWiley
• Kroenke, Spitzer & Williams (2003), PHQ-2, Med Care
• PHQ-9 in older adults (cutoff shifts) — PMC8559588 · cognitive impairment PMC3930057
• Spitzer, Kroenke, Williams & Löwe (2006), GAD-7, Arch Intern Medreference PDF · psychometrics PMC6691128
• PHQ-9 monitoring / sensitivity to change (Löwe 2004) — PMID 15550799; RCI/MID — NovoPsych · Jacobson–Truax
• Chatbot PHQ-9 administration concordance (HopeBot) — arXiv 2507.05984
• NASEM (2020), Social Isolation & Loneliness — nationalacademies.org · PMC7742588
• PARO RCT meta-analysis — Int J Nurs Stud 2023
• ElliQ progress/lessons (2024) — PMC10917141
• Woebot RCT (Fitzpatrick 2017) — JMIR Ment Health
• ICT-for-loneliness RCT meta-analysis (2021) — PMC8692663