Somewhere you picked up the idea that the feeling comes first and the doing follows. Behavioural activation runs it the other way round — something that matters, in the diary, done before you feel like it — and in the trials, on average, it worked. Why it works is a question that has been asked very carefully, and not answered.
Somewhere along the way you signed up to a rule, although you cannot remember signing. The rule says: first you feel like it, then you do it. It sounds fair. It sounds like how people work. It is also, in a certain light, a small trap with very good manners, because it hands the keys to the one part of you that is not currently answering the door. You wait for the wanting-to. The wanting-to waits for something to have gone well. Nothing goes well while you are waiting. Round it goes, very politely, and the email stays unsent, and the walk stays unwalked, and the call to the friend stays a thing you will definitely do tomorrow.
This is not laziness. It is not a verdict on you at all. This is a loop that ran on its own for years, so quietly you never knew it was there — until the restructuring reached your floor, and people who have never heard your name started redrawing your job without asking you, and the ground under the loop moved, and one morning the engine simply would not start. It is still there. It is just not running.
There is a way out of the rule. It was tested on more people than you would guess. And the strange part is that nobody can say why it works.
Does behavioural activation work? — While you were deciding about Thursday
The thing has a name, which is a relief, because things with names have usually been looked at by someone. It is called behavioural activation. The Ledger — the register behind this site, every practice graded twice, once for whether it works and once for why — describes it as scheduling and then doing something that matters to you, before the motivation to do it arrives. That is the rule from the first paragraph, turned round and stood back on its feet. The theory says waiting to feel like it is what keeps people stuck; whether that is the mechanism, or only the explanation on offer, is the second half of this article.
The first half is simpler. Does it work? On average, yes.
Here is what “on average” looks like from the inside. While you are deciding whether to keep Thursday’s walk, picture the waiting rooms of ordinary British general practices, the family doctor’s — the plastic chairs, the poster about flu jabs — and in them 440 people with a diagnosis of depression. They were sorted by chance between the standard talking treatment — which is to say cognitive behavioural therapy, from accredited therapists — and this: a person across a small table with a folder, a diary and a short course of training, asking what mattered to them, writing it in the diary for a day that week, and asking at the next session whether they had done it.
By the time the trial’s clock ran out — and it ran a long while — both sets of people had arrived at the same place. Not a similar place. To one decimal place, the same score on the depression questionnaire, and the difference between them fell well inside a margin the trial had fixed before anyone sat down.
That is one trial. Stack up every randomised trial of this — the review that did so found 105 of them, involving 13,933 people — and the picture holds. People who did this ended up better than people who did not, by a margin the field calls moderate to large, and the difference was still there twelve months after they were assigned. Not a fortnight’s lift that fades by the time the next all-staff email arrives.
It is not the small print; it is the whole of the honest claim: this works for groups of people, on average, and it is not a forecast for any single person.
A separate fact, about that word “average”, because it is doing more work than it looks. The trials did not all find the same thing, which is what an average is for. Some found more than that margin, some less, and some, run perfectly well, would be expected to find nothing at all. The average holds; the trials disagree about how much. Every figure behind this — the size of the effect, how far the trials disagreed, where the next one might land — is set out in the table at the end of this article, in the section called The working.
And now the sentence the trials earn and the headlines drop. It is not the small print; it is the whole of the honest claim: this works for groups of people, on average, and it is not a forecast for any single person. Inside that average, in the trial with the waiting rooms, roughly one in five people were, in the investigators’ own word, unchanged by either treatment. They are in the average too. A map that left them off would be a prettier map and a worse one: it would show you the route and hide the people for whom it led nowhere.
How might it work? — The room where they looked for the reason
So far the story has a working engine and no idea what the fuel is. The theory has an idea, and it is a good one, which turns out to be exactly the problem.
It goes like this. Low mood narrows what a person does. The world shrinks to the sofa, the phone, and the parts of the flat that are on the way to the kettle. The narrowing cuts you off from the things that used to pay you back — the friend, the walk, the part of your job that used to feel like yours — and they stop paying. Restore the doing, says the theory, and you restore the contact; restore the contact, and mood follows. Two pathways sit at its centre: the doing itself, and what the world gives back when you turn up — environmental reward, in the theory’s terms, which is the kind of phrase that makes a park bench sound like a bank.
It is a tidy theory, and it is checkable, by a mediation analysis: not whether the treatment worked — the trials have answered that — but whether the thing the theory points at changed first, and whether the mood improved because it did. Think of the improvement as a passenger who has arrived. The theory names the route. The analysis is the inspector who asks to see the ticket, to find out where it was actually stamped.
The trial that best shows this works is the trial that best fails to show why.
Fourteen studies asked to see the ticket, testing twenty-one candidate reasons between them. For the theory’s two favourites, the stamps were faint at best. Across the board, the results were too faint to read with confidence, some of the studies were roughly built, and they had used different kinds of ticket, so they could not be laid side by side. The reviewers’ conclusion, near enough in their own words: no firm conclusion could be reached about any of the twenty-one. The passenger arrived. The ticket does not show which way. Whatever Thursday’s walk does, if it does anything, nobody can yet say which part of it did the work.
The largest of those studies — the biggest sample, the highest quality score, the only one to meet every criterion the reviewers set — was the trial with the waiting rooms. It could not confirm that doing more was what did the work. The trial that best shows this works is the trial that best fails to show why.
Other candidates remain: re-engaging with what rewards you; sidestepping less; a returning sense of being able to do things; plain momentum, the way a moving bicycle is easier to keep upright than a stopped one. Every one plausible. Not one confirmed.
What “nobody can say why” does and does not mean
Two precisions, because this is where a tired reader is entitled to draw the wrong conclusion, and it is an easy one to draw.
First: unconfirmed is not disproved. A room you have not walked into is not a room that has been shown to be empty. The theory’s two pathways may be exactly right; the studies built to catch them in the act were not, between them, built well enough to catch anything, and the reviewers said so.
Second: “no confirmed why” is not “no why”. Something is happening between the diary and the mood. It shows up in trial after trial and it lasts. The reason has simply not been identified to the standard the field sets for identifying reasons. That is why the Ledger grades the two questions separately, and why they come out differently. The grade: Outcome A — Strong. Mechanism — Plausible. Confidence — High. Look at the middle line. Not Established. Plausible. The gap between those two words is the honest size of what is not yet known.
The rule, run backwards
Now go back to the rule you signed. First you feel like it, then you do it. Every person given this, in every one of those trials, was asked to break it, in the same direction, on purpose. They were not to wait for the wanting-to. They put the thing in the diary and went, and the wanting-to, where it came at all, came afterwards, a little sheepishly, like a guest who has missed the start.
What they show, taken together, is that going without wanting to is a thing a person can do, and that the not-wanting does not, on average, cancel the going.
This is the agency point, and it is smaller and stranger than it sounds. It is not that you can make yourself want to. You cannot; the research does not say you can; anyone who says otherwise is selling something. It is that the order can be reversed. The wanting-to is not a gatekeeper. It was only ever standing near the gate.
Now watch one person in one of those trials — she is a composite, several people’s weeks folded into one, not a real person from any of the trials — who put a walk in her diary for Thursday, did not feel like going when Thursday came, went anyway, and came back a little less flat. Not restored. Less flat. That is the whole scene, and the trials are, in the end, a great many versions of it, some of which ended less flat and some of which did not.
What they show, taken together, is that going without wanting to is a thing a person can do, and that the not-wanting does not, on average, cancel the going.
What the people in the trials actually did
Because they did not, in the main, do anything grand. Over a course of weeks they were asked what they had stopped doing and what it used to give them; they kept a note of what they did with their days; they chose things that mattered and put them in the diary, tailored to them — more time in the situations they had named as good, less of the sidestepping; and when the going-over-and-over started, they worked out what they might do with their hands instead. Then they went, and came back, and said how it had gone. The full list of components, and how many sessions carried them, is in The working, at the end.
What came back with the woman from the previous section was not enjoyment. Enjoyment is a poor instrument; it goes when mood goes. What came back was quieter than that — an afternoon that ran slightly longer, a message answered, a Thursday that had a shape — and she noticed it the way you notice a room has warmed, without having looked at the thermometer.
The people across those tables were graduates with no professional psychotherapy qualifications, given five days’ training and regular supervision, and the trial could not tell their results from those of accredited therapists — at a lower cost, since someone always asks. As far as the trial can tell, whatever does the work here travels in the structure — the choosing, the diary, the going, the coming back — and not the letters after the name of the person across the table.
Where this page stops
This is not therapy, and it is not a substitute for it. The people in the trials had been diagnosed in a proper clinical interview, were inside a service, and were asked, session after session, how they were. A page cannot ask. If this week is worse than a flat Thursday, the Start Here page has a short list of people who can, and it is there for exactly that.
The trial has edges, and they are drawn rather than smoothed in The working, at the end: everyone but the assessors knew which treatment they were in; the main measure was self-reported; some of the data went missing; nobody controlled who was taking antidepressants, because that is what an ordinary service looks like. None of it undoes the finding. All of it stands next to it.
The smallest form of this that the trials tested is the one the Ledger entry describes — choosing something that matters, scheduling it, doing it — carried out by people working from materials, no therapist in the room.
That is where “smaller, still there” comes from: smaller than the clinic figure, and still there on average. Whether it does anything for you is not something a page can say. What a page can say is that this is the version people tried on their own, and that the trials did not find it to be nothing.
The Ledger lists the effort as low to moderate, with the benefits over weeks. The engine, for what it is worth, did not need to know what the fuel was in order to start. The mechanism is pending. The effect is not.
The working
For readers who want the numbers. Everything above stands without this section. Glosses — SMD, standardised mean difference: how far the treated group ended up from the untreated group, in units of the spread of the scores (that spread is a standard deviation). CI, confidence interval: the range the true average most likely sits in. Prediction interval: where the next trial would be expected to land. Heterogeneity: how much the trials disagreed. Non-inferiority: no worse than the comparison by more than a margin fixed in advance. Per-protocol: only those who received enough of the treatment. Open-label: everyone knew which treatment they had. How the two-axis grade works is set out on the Ledger.
Ledger entry 3 — behavioural activation. Every figure, with its source.
| What | Figure | Source |
|---|---|---|
| 2026 review — trial base | 105 randomised trials · 13,933 people · 40 at low risk of bias (allocated by chance; the right people unaware of who got what; little data missing) | Cuijpers 2026 |
| Adult outpatients vs control | SMD 0.67 (95% CI 0.54–0.80; 61 comparisons) — moderate-to-large: clear of the line for “moderate”, short of the line for “large”; roughly two-thirds of a standard deviation below control groups | Cuijpers 2026 |
| Disagreement between trials (heterogeneity) | High — 71% by the usual measure | Cuijpers 2026 |
| Prediction interval | −0.11 to 1.46; crosses zero — some trials, run properly, would be expected to find no effect at all | Cuijpers 2026 |
| Durability | Effect still present twelve months after randomisation | Cuijpers 2026 |
| Self-guided, no therapist | SMD 0.36 (95% CI 0.20–0.51; fifteen comparisons) — smaller, and still there | Cuijpers 2026 |
| Against other psychological therapies | SMD 0.04 (95% CI −0.10 to 0.18; 27 comparisons) — no difference detected | Cuijpers 2026 |
| Earlier reviews | 2014 (PLoS ONE): −0.74 against control conditions, 26 studies; most trials low quality, short follow-up. 2007 meta-analysis pointed the same way. Minus sign is notation, not a different finding (2014 a fall in symptoms; 2026 a positive difference; sizes comparable). What changed: the trial base, 26 → 105 | Ekers 2014 · Cuijpers 2007 |
| COBRA — design | 440 adults, diagnosed major depression · UK primary care and psychological therapy services · behavioural activation vs CBT · PHQ-9 (self-completed depression questionnaire) at twelve months · The Lancet, 2016 | Richards 2016 |
| COBRA — main result | CBT 8.4 · behavioural activation 8.4 · difference 0.1 (95% CI −1.3 to 1.5) · non-inferiority margin, fixed in advance, 1.9 · held in main and per-protocol (at least eight sessions) analyses · nine or under counted, for an individual, as recovery | Richards 2016 |
| COBRA — delivery and cost | Graduates, no professional psychotherapy qualifications · five days’ training · an hour of supervision a fortnight · indistinguishable from accredited CBT · roughly 21% lower cost | Richards 2016 |
| COBRA — who improved (twelve months) | Recovered (nine or under) or responded (score at least halved): 61–70% · unchanged by either treatment, the investigators’ word: 20–23%, one in five · ranges do not meet at 100; remainder not reported, so not stated here | Richards 2016 |
| COBRA — attendance | Roughly a third attended fewer than eight sessions (the minimally sufficient dose), averaging 2.5 — “a problem well known to routine psychological therapy services” · mean 11.5 sessions; minimum-dose group 16.1 | Richards 2016 |
| COBRA — what was offered | Up to twenty sessions of sixty minutes over sixteen weeks, up to four boosters; all core components by session eight (minimally sufficient dose). Components, per the trial: identifying depressed behaviours; analysing their triggers and consequences; monitoring activity; developing alternative, goal-orientated behaviours; activity scheduling; alternative behavioural responses to rumination. Tailored — more contact with situations named as positive, less avoidance | Richards 2016 |
| COBRA — limits | Open-label (assessors masked) · self-reported primary outcome · 17% of main-outcome data missing, more in the behavioural activation arm · antidepressant use uncontrolled (pragmatic trial). None undoes the finding; all stand next to it | Richards 2016 |
| Mechanism review | Fourteen forward-looking mediation studies · twenty-one candidate mechanisms · the theory’s two favourites (activation; environmental reward) unconfirmed · other candidates — reward re-engagement, avoidance reduction, a returning sense of mastery, behavioural momentum — every one plausible, not one confirmed · 2021 | Janssen 2021 |
| COBRA inside the mechanism review | Largest sample (440) · highest quality score · the only one of the fourteen to meet every criterion the reviewers set · could not confirm that doing more was what did the work | Janssen 2021 · Richards 2017 |
| Ledger entry 3 | Outcome A — Strong · Mechanism — Plausible · Confidence — High in the grade itself · effort low to moderate · benefits over weeks | The Ledger |
Sources
Every claim in this article traces back to published, peer-reviewed research. We’ve listed the primary sources below so you can read the original studies yourself — because taking our word for it rather defeats the purpose.
- Cuijpers P, Ciharova M, Tong L, Liu Y, Sprenger AA, Miguel C, Karyotaki E, Harrer M (2026). Behavioral activation for depression: A comprehensive systematic review and meta-analysis. Clinical Psychology Review 128:102783. doi:10.1016/j.cpr.2026.102783 — effectiveness anchor
- Richards DA, Ekers D, McMillan D, et al. (2016). Cost and Outcome of Behavioural Activation versus Cognitive Behavioural Therapy for Depression (COBRA): a randomised, controlled, non-inferiority trial. The Lancet 388(10047):871–880. doi:10.1016/S0140-6736(16)31140-0 — outcome / non-inferiority result
- Richards DA, Rhodes S, Ekers D, et al. (2017). Cost and Outcome of BehaviouRal Activation (COBRA): a randomised controlled trial of behavioural activation versus cognitive-behavioural therapy for depression. Health Technology Assessment 21(46):1–366. doi:10.3310/hta21460 — COBRA process evaluation (mediation)
- Janssen NP, Hendriks GJ, Baranelli CT, et al. (2021). How Does Behavioural Activation Work? A Systematic Review of the Evidence on Potential Mediators. Psychotherapy and Psychosomatics 90(2):85–93. doi:10.1159/000509820 — mechanism / mediation review
- Ekers D, Webster L, Van Straten A, Cuijpers P, Richards D, Gilbody S (2014). Behavioural activation for depression; an update of meta-analysis of effectiveness and sub group analysis. PLoS ONE 9(6):e100100. doi:10.1371/journal.pone.0100100 — historical support
- Cuijpers P, van Straten A, Warmerdam L (2007). Behavioral activation treatments of depression: a meta-analysis. Clinical Psychology Review 27(3):318–326. doi:10.1016/j.cpr.2006.11.001 — historical