Our clinical document, audited row by row — including the rows that did not survive.
We went back to the original research behind what this app does and read it. The register is public, and it includes the rows that did not survive. The 28 below are the ones that decided something you can see in the app — a technique it offers, a rule it follows, or something it deliberately will not do. Three of them did not hold up and one could not be checked at all. Those four are in the list with the rest.
The rest are working notes: background reading, the limits that sit on top of the rows below, and ideas the project looked at and chose not to build.
- Holds 16
- Holds, with a qualification 8
- Did not hold 3
- Unverifiable 1
Each bar is a share of the 28 rows on this page. Choose one to read only those rows.
g = 0.22
95% CI 0.10 to 0.33
Averaged across 18 trials of 22 mental-health apps: how much they moved depressive symptoms, compared with people who were doing something else rather than nothing at all.
The ceiling everything else fits inside
Apps like this one help, and they help by a small amount. That number is the size of it. It is a limit rather than a promise, and nothing claimed on this page is allowed to sit outside it.
None of the studies below is a trial of Moorwell. No outcome study of this app has been run. Every row is evidence about a technique the app draws on, which is a different and weaker thing.
The same number, drawn
g = 0.22 is the measurement. This is what it looks like. Take one person who used an app like this and one person who did something else instead, and compare how their depressive symptoms moved. Do that a hundred times.
50 of them would go the app user's way if these apps did nothing whatsoever. That is a coin toss, and it is the number to compare against — not zero.
56 go their way in fact. The 6 extra squares are the entire effect that 18 trials of 22 apps could find. Small, real, and measured.
Pooling the studies differently moves the figure between 52.8 and 59.2 — the same 95% interval as the one printed beside g = 0.22 above, said the other way. Converted with the probability of superiority, Φ(g ÷ √2), which assumes both groups are normally distributed with equal variance and treats Hedges' g as Cohen's d. It is the same measurement in different units, not a second finding: the conversion is monotone and cannot make a small effect large.
Against people doing nothing rather than something else, the same meta-analysis finds a larger effect. This page shows the smaller one on purpose: an app should be judged against the alternative a person actually has, not against an empty chair.
What d, g and SMD mean
Every number on this page is one of four things, and all four are the same kind of measurement: how far apart two groups ended up.
- d
-
Cohen’s d
The gap between the people who got the thing and the people who did not, measured in standard deviations. 0.5 means the average person who did it ended up better off than about 69 in 100 of those who did not.
- g
-
Hedges’ g
The same idea as d, corrected for small studies, which nudges the number slightly down. Treat g and d as the same scale when you compare them.
- SMD
-
Standardised mean difference
The umbrella term for both. A minus sign only means the thing being measured went down — with sleep problems or low mood, down is the good direction.
- d+
-
Pooled d
A d averaged across many studies at once, so it carries more weight than a single trial’s number.
What counts as big
0.2 small0.5 moderate0.8 large
These labels are conventions, not laws. In mental health a small effect delivered to a lot of people at no cost can matter more than a large one that needs a therapist and a waiting list — which is the argument for an app existing at all.
What the techniques measure, against the honest ceiling
Each bar is the headline effect from the paper named beside it, read from that paper's own abstract. The band behind them is where unguided self-help apps land as whole products — and it is a range rather than a line, because the same apps measure 0.56 against an inactive control and 0.22 against an active one. A strong technique does not lift an app above that band.
Every technique in the app, and the paper behind it
The three findings that did not hold up, and the one we could not check at all, are here at the same size as everything else. A list that shows only its wins is an advert.
Narrow the register
Showing all 28 rows.
-
Details for What an app of this kind can move at all
SurvivesWhat an app of this kind can move at all
Firth et al. 2017
- Study
- Meta-analysis of RCTs · 18 RCTs, 22 apps, N = 3,414
- Population
- Adults with depressive symptoms, mixed settings
- Replication
- Direction replicated by a later 176-RCT meta-analysis
This is the governing ceiling for the whole project, and every claim on this site has to fit inside it. Against an active comparator — people who were doing something else, not nothing — the pooled effect is g = 0.22 (0.10–0.33).
-
Details for The same question, asked again across 176 trials
SurvivesThe same question, asked again across 176 trials
Linardon et al. 2024
- Study
- Meta-analysis · 176 RCTs; N = 33,567 depression, 22,394 anxiety
- Population
- Adults, mixed settings
- Replication
- Consistent in direction and size with the row above
Apps containing CBT do better than apps that do not, which is why the content bank stays anchored in cognitive and behavioural material rather than in general wellness.
-
Details for Guidance changes the result more than the content does
SurvivesGuidance changes the result more than the content does
Moshe et al. 2021
- Study
- Meta-analysis · 83 studies, N = 15,530
- Population
- Adults with depression
- Replication
- —
Guided programmes outperform unguided ones, and trials run in real services underperform trials run in ideal conditions. Moorwell is unguided and free, so the bottom of that range is the honest expectation for it.
-
Details for Nerves reframed as your body getting ready
QualifiedNerves reframed as your body getting ready
Jamieson et al. 2016, 2022
- Study
- Two randomised classroom trials, both from the group that built the manipulation · 93 and 339
- Population
- Community-college students sitting real exams — the best population match in the whole register
- Replication
- A 2026 replication across twelve courses at seven institutions, against a placebo control, did not reproduce it. That study is the next row, and it is the reason this one reads Qualified rather than Survives
The part that holds is the control arm, and it decides a design question: “ignore the stress” was the placebo condition. That is why the app never tells anyone to calm down before an evaluation — a rule that rests on what the trials compared against, not on the size of what they found.
-
Details for …and the replication that did not find it
Qualified…and the replication that did not find it
Thormodsæter et al. 2026
- Study
- Replication study in real courses · 12 courses across 7 institutions, intervention against an active placebo
- Population
- Undergraduates in real courses
- Replication
- This is the failure, and it is cited beside the finding it weakens
A 2026 replication did not reproduce the effect. The project's own audit records that the claim needs a dated correction beside it naming this study. It is on this page for the same reason it is in the register: a citation that only survives when you leave out the replication is not a citation.
-
Details for Writing your worries down before a test
Did not holdWriting your worries down before a test
O'Meara & Lovett 2026
- Study
- Meta-analysis · 21 studies, N = 1,457, 30 effect sizes
- Population
- Students taking real tests — exact population, exact moment
- Replication
- The pooled result is the correction
Single-session expressive writing before an exam is one of the most repeated pieces of study advice there is. Pooled, it does not hold. The app does not do it, and this row is why.
-
Details for Lines that make a claim about the reader
QualifiedLines that make a claim about the reader
Wood, Perunovic & Lee 2009; Flynn et al. 2020
- Study
- Lab experiments, plus a replication report titled “On the failure to replicate…” · Small samples
- Population
- Undergraduates
- Replication
- Mixed — the failure is reported and cited
One study found that “you are enough” lands worst on the person who least believes it. Two later attempts did not find the same thing. Mixed evidence about exactly the people most likely to read the line is a reason to place it carefully rather than to use it or drop it, so the app holds those lines back on a bad day and never writes one about you.
-
Details for Speaking to yourself the way you would to a friend
SurvivesSpeaking to yourself the way you would to a friend
Wakelin et al. 2022
- Study
- Meta-analysis · 29 studies reviewed; the pooled effect is k = 28, N = 1,636
- Population
- People with clinical diagnoses or elevated symptom scores — the best population match for any pooled effect in the register
- Replication
- None located
This is the strongest row the register holds for a technique the app actually ships.
-
Details for Self-criticism predicting depression on its own
Did not holdSelf-criticism predicting depression on its own
Gittins & Hunt 2020
- Study
- Prospective, three waves, cross-lagged · 243 adolescents, mean age 12.08
- Population
- Early adolescents
- Replication
- The clearest direct test located, and it does not support the claim as stated
The app still answers self-criticism, because the row above supports the technique. What died is the stronger claim that self-criticism predicts later depression independently of current distress. Both halves are on this page.
-
Details for Putting the feeling into words
QualifiedPutting the feeling into words
Lieberman et al. 2007
- Study
- Single fMRI study · Small imaging sample
- Population
- Volunteers viewing standardised images
- Replication
- No meta-analysis located — a search of the literature returns zero title-level meta-analyses of affect labelling
This is a single imaging study carrying a design rule, and the register says so. What survives it is narrow and is what the app does: the naming has to be yours, so the check-in asks what is going on and never offers a feeling to pick from a menu.
-
Details for Finding a more exact word for it
SurvivesFinding a more exact word for it
Nook et al. 2018
- Study
- Developmental study · Large developmental sample
- Population
- Across adolescence into adulthood
- Replication
- —
Emotion differentiation follows a curve with its low point in adolescence — which is the age the app was built for, and a reason the vocabulary it offers matters.
-
Details for Absolute words as a signal
SurvivesAbsolute words as a signal
Al-Mosaiwi & Johnstone 2018
- Study
- Three text-analysis studies · 63 forums, 6,400+ members
- Population
- Self-selected internet forum members
- Replication
- None located
Absolutist words track severity better than negative-emotion words do. The population is people who chose to post on a forum, which is not a general sample, and the register records that rather than smoothing it.
-
Details for A plan with a when and a where in it
SurvivesA plan with a when and a where in it
Toli, Webb & Hardy 2016
- Study
- Meta-analysis · Clinical and analogue mental-health samples
- Population
- People with mental health problems
- Replication
- Supported by a later synthesis of 642 tests
“I will do more work” is a wish. “After my last lecture, at the desk by the window” is a plan. The app stores the when and the will as two separate fields, because the shape is the part doing the work.
PMID 25965276 10.1111/bjc.12086 10.1080/10463283.2024.2334563
-
Details for Doing one small thing, first
QualifiedDoing one small thing, first
So et al. 2025
- Study
- Single RCT · N = 67, ages 20–30, 8 weeks
- Population
- Young adults 20–30
- Replication
- None located
A single small trial of a behavioural-activation app. The population is close to the app's own, and one trial is one trial — the register grades it qualified for that reason, and we have kept that grade.
-
Details for Slower breathing, at about six breaths a minute
SurvivesSlower breathing, at about six breaths a minute
Fincham et al. 2023
- Study
- Meta-analysis · k = 12, N = 785
- Population
- General adults
- Replication
- Mechanism replicated, including a published null on the in-to-out ratio
The null is the interesting half: a 2024 study found that lengthening the exhale specifically made no difference to heart-rate variability during slow-paced breathing. The app's exercise still runs longer out than in, and we do not claim the ratio is what does the work.
-
Details for Letting the muscles go
QualifiedLetting the muscles go
Donato et al. 2026
- Study
- Meta-analysis · 31 RCTs, 2,277 adults
- Population
- Medically ill and inpatient adults
- Replication
- Contradicted as a component of CBT by a component analysis; loses to active comparators
The population does not transfer — the register marks this one no on that column outright, and the app treats relaxation accordingly rather than as a headline technique.
-
Details for Mindfulness, and why it is not the app's answer
SurvivesMindfulness, and why it is not the app's answer
Goyal et al. 2014, and three later syntheses
- Study
- Meta-analyses · 47 trials / 3,515; 51 RCTs; 83 studies / 6,703; 43 studies / 1,427
- Population
- Mixed; one of the four is university students
- Replication
- —
Meditation programmes are no better than any active treatment they were compared against, and adverse events are reported in a measurable minority. That is not a reason to think badly of mindfulness. It is a reason this app did not build its centre on it.
-
Details for Sleep, and which parts of the therapy carry it
SurvivesSleep, and which parts of the therapy carry it
Furukawa et al. 2024; a student trial
- Study
- Component network meta-analyses, plus a null RCT · 241 trials / 31,452; 80 / 15,351; 195 students
- Population
- Mean age 45.4 in the largest; the student trial is the null
- Replication
- Two independent network analyses converge on the same two components
The two components that survive are the two the app implements: a wake time you keep and a bed kept for sleeping. The population gap is real and the project's clinical file states it — the big evidence is middle-aged, and the trial closest to a student sample is the null one.
-
Details for Digital CBT reaching student anxiety
SurvivesDigital CBT reaching student anxiety
Oliveira et al. 2023
- Study
- Meta-analysis of RCTs · 15 studies, N = 1,619 university students
- Population
- University students — exact population
- Replication
- Largest synthesis located for this population
This is the closest the register comes to evidence for this kind of programme in this kind of person. It is still a category result, not a result about Moorwell.
-
Details for Worry and rumination are one process
SurvivesWorry and rumination are one process
Stenzel et al. 2025
- Study
- Meta-analysis · Multi-study synthesis
- Population
- Mixed
- Replication
- —
Treatment does not care which register the thought is in. That is why the app aims at the loop rather than at whether someone calls it worrying or going over things.
-
Details for Postponing a worry, tested against an active control
SurvivesPostponing a worry, tested against an active control
Versluis et al. 2016; a 2026 replication
- Study
- Two RCTs against an active thought-log control · 51 and 117 adults with clinical GAD
- Population
- Adults with diagnosed GAD, not students
- Replication
- Replicated, and against an active comparator rather than a waiting list
The register marks the population transfer no as evidenced: this is a clinical GAD sample. What does transfer is the delivery — a phone, a short window, repeated prompts.
-
Details for Media portrayal, and the rule it sets for the app
SurvivesMedia portrayal, and the rule it sets for the app
Domaradzki 2021
- Study
- Literature review · 108 papers
- Population
- General media audiences; effects vary by age
- Replication
- Long-standing and repeatedly observed
How a thing is described changes what happens next. This is the row underneath the app's safe-messaging rules, and it is the reason the crisis pathway is fixed wording rather than anything generated.
-
Details for Gratitude, across 28 countries
SurvivesGratitude, across 28 countries
Choi et al. 2025
- Study
- Meta-analysis of RCTs · 25 RCTs, 6,745 participants in the earlier synthesis
- Population
- Mixed adults, many countries
- Replication
- Direction supported across countries, with significant between-country variation and no moderator that explains it
It works, it is small, and nobody can yet say why it varies by country. For an app whose first population is Indian students, an unexplained cross-country variance is exactly the kind of thing worth printing.
-
Details for How long people actually use an app like this
QualifiedHow long people actually use an app like this
Baumel et al. 2019
- Study
- Panel usage analysis of real installs · 93 apps, 10,000+ installs each
- Population
- Android app users in the wild
- Replication
- —
Real-world engagement with mental-health apps is far below what trials report. Moorwell is designed around that rather than against it — nothing resets, nothing is lost by a gap, and the app never mentions that you were away.
-
Details for Choice overload
Did not holdChoice overload
Iyengar & Lepper 2000
- Study
- Three studies, including the jam field study · Small
- Population
- Shoppers and students
- Replication
- Failed. A 2010 meta-analysis found the mean effect to be about zero
The app still offers few options at a time, and the reason is not this study. It is that a long menu is more work to read at 2am. The behaviour survived; the citation under it did not, and swapping the reason is cheaper than pretending.
-
Details for Growth mindset, and the size it actually is
SurvivesGrowth mindset, and the size it actually is
Sisk et al. 2018; Macnamara & Burgoyne 2023
- Study
- Two meta-analyses plus a later sceptical synthesis · Large, multi-study
- Population
- Students
- Replication
- The 2023 synthesis is the sceptical follow-up and it agrees with the 2018 one
The claim that survives here is the cautious one: brief mindset interventions run to a very small effect, and the later literature moved further that way rather than back. This is a case study in a finding that reached the whole world before the pooling did, and the app builds none of it.
-
Details for Evidence about Indian students specifically
UnverifiableEvidence about Indian students specifically
Louis 2016; Kumar et al. 2023
- Study
- Cross-sectional descriptive studies · 500 in one; not stated in the other
- Population
- Indian students and professionals — the app's own population
- Replication
- Searches run and recorded; no outcome or intervention study located
This is the most important row on the page. For the population this app was built for, the literature that exists is descriptive — it measures how common something is, not whether anything helps. The register grades that unverifiable rather than leaving the column blank, and the site is not going to imply otherwise.
-
Details for The one Indian outcome trial the search did locate
QualifiedThe one Indian outcome trial the search did locate
Michelson et al. 2020; Gonsalves et al. 2022
- Study
- Outcome trial with 12-month follow-up, plus a pilot RCT of the app version · School students aged 12–20 in New Delhi
- Population
- Delivered in Hindi by a human lay counsellor, against a printed booklet
- Replication
- The app version is a pilot
Structured problem-solving is the only candidate in the whole research pass with an outcome trial in this population. It was delivered by a person, and the project's own rule for that pass says no row in it may be described to a reviewer as evidence for an unguided app.
How to read a verdict
- Survives
- We found the paper. It says what we said it says, and the study behind it is strong enough for the use we put it to.
- Qualified
- We found the paper, and our claim about it is too broad — the wrong group of people, a weaker kind of study, or one paper asked to carry more than it can.
- The feature may still be right. The reason we gave for it is not good enough as written.
- Did not hold
- Someone ran the study again and did not get the same answer, or better evidence points the other way.
- Unverifiable
- We could not find the paper at all, or we could not trace the number back to the source it was credited to.
- A claim that fails stays on the list, with the date and the correction next to it. Delete it, and the next person works it out from scratch and believes it again.
A psychologist has read the app’s own words
Everything above is this project auditing the research it leans on. Reading what the app actually says to you is a different job, and a psychologist has done it — the wording changed where she asked for it.