Antidepressants are believed to work by correcting a chemical imbalance, in particular a shortage of serotonin in the brain (this matters if you want to understand the placebo effect and antidepressants). But an analysis of published and unpublished data that the drug manufacturers had kept hidden showed that most of the benefit, if not all of it, comes down to the placebo effect. Some antidepressants raise serotonin levels, others lower them, and others have no effect on serotonin at all.
All the same, they all show the same therapeutic benefit. Instead of treating depression, some common antidepressants may produce a biological vulnerability that makes people more prone to depression later on. Other treatments (psychotherapy and physical training, for example) give the same short-term benefit as antidepressants, but show better long-term effectiveness and do so without the side effects and the health risks of taking medication.
When Sapirstein and I began to analyse the data from clinical trials of antidepressants, we were not particularly interested in antidepressants. Rather, we were interested in understanding the placebo effect, as I had been throughout my academic career. I asked myself: how is it that believing you have taken a drug can produce some of the effects of that drug?
Sapirstein and I felt that depression was a good place to observe placebo effects. After all, one of the central features of depression is the hopelessness that depressed people feel. If you ask people suffering from depression what is worst in their lives, many will answer that it is their depression. The British psychologist John Teasdale called this "depression about depression". If that is so, then the promise of an effective treatment can by itself relieve depression, because hope replaces hopelessness, the hope that you will get well in the end. With that in mind, we decided to measure the placebo effect in depression.
Sapirstein and I searched the research literature for studies in which patients had been randomly given either an inert placebo or no treatment at all. The studies we found also contained data on the response to antidepressants, because this is the only place where data on the placebo response in depressed patients are reported. I was not especially interested in the effects of the drug. I assumed that antidepressants were effective.
As a psychotherapist, I sometimes referred my seriously depressed patients to doctors so that antidepressants could be prescribed for them. Sometimes my patients improved once they started taking antidepressants, sometimes they did not. If they felt better, I assumed it was because of the effect of the drug. Given my long-standing interest in the placebo effect, I ought to have guessed that it was most likely a "placebo effect", but at the time that did not seem so pressing.
When Sapirstein and I analysed the data we had gathered, we were not surprised to find a decisive placebo effect in depression. What did surprise us was how small the drug effect was. 75 % of the improvement in the treatment group also occurred when patients were given a dummy pill with no active substance in it. Needless to say, our meta-analysis proved highly contentious. Its publication led to heated argument. The critics' response was that these data could not be right. Perhaps our search had led us to analyse an unrepresentative set of clinical trials. Antidepressants, the critics said, had been assessed in many trials, and their effectiveness was well documented.
To answer the criticism, we decided to repeat our study using a different set of clinical trials (Kirsch et al., 2002). To do this we used the Freedom of Information Act (FOIA) 1) to request from the Food and Drug Administration (FDA) 2) the documents that the pharmaceutical companies had submitted during the approval process for six new-generation antidepressants, the ones most widely prescribed at the time. This FDA data set has a number of advantages. Most importantly, the FDA requires pharmaceutical companies to supply information about every clinical study they sponsor. So we had data from unpublished as well as published studies. That turned out to be very important. Almost half of all clinical trials sponsored by pharmaceutical companies had not been published. Only the pharmaceutical companies and the FDA knew the results of the unpublished trials, and almost all of them failed to find any significant advantage of the drug over placebo. A second advantage of the FDA data sets is that the same instrument for measuring depression was used for all the data collected: the Hamilton Rating Scale for Depression (HAM-D). This made it easier to grasp the clinical importance of the difference between drug and placebo. After all, the data in the FDA data set form the basis for a drug's approval, so they have a special status. If there were something wrong with these trials, the drugs would not have been approved at all. According to the data we obtained from the FDA, only 43 % of the drugs tested were statistically significant compared with placebo. The remaining 57 % were failed or negative.
Our analysis showed that the response to antidepressants and to the placebos for these antidepressants was 82 % attributable to placebo effects. As a result, my colleagues and I repeated our meta-analysis on a larger number of trials submitted to the FDA (Kirsch et al., 2008). With this extended body of data we again found that 82 % of the drug effect was duplicated by the placebo effect. More important still, in both analyses the average difference between drug and placebo was less than two points on the HAM-D. The scale consists of 17 items, and subjects can score between 0 and 53 points, depending on the severity of their depression. 6 points can come from a change in sleep patterns alone, with no change in the other symptoms of depression. So the difference of 1.8 points that we found between drug and placebo was in fact very small, small enough to have no clinically meaningful significance (see Fig. 2). But you do not have to take my word for how small that difference is. The National Institute for Health and Clinical Excellence (NICE), which draws up treatment guidelines for the National Health Service in the United Kingdom, has set a difference of 3 points between drug and placebo on the HAM-D as the criterion of clinical significance (NICE), 2004). If the published and the unpublished data are combined, there is no meaningful difference between antidepressants and inert placebos.
It has been argued that the NICE criterion was arbitrary (Turner & Rosenthal, 2008, for example), and that is true. It is just as arbitrary as choosing p < 0.5 as the criterion of statistical significance. In the meantime, Joanna Moncrieff and I have found a non-arbitrary criterion of clinical significance (Moncrieff & Kirsch, 2015). In 2013 Stefan Leucht and his colleagues (Leucht et al., 2013) compared ratings on the HAM-D with ratings on the Clinical Global Impression – Improvement scale (CGI-I) (Guy, 1976), a 7-point scale on which clinicians rate patients from 1 (very much improved) through 4 (no change) to 7 (much worse). Leucht used patient data from 43 clinical studies with 7,131 patients and found that the mean change on the HAM-D for patients who were rated as unchanged on the CGI-I was 3 points, exactly the criterion that NICE had defined as a clinically significant improvement (see Fig. 3). The problem with the NICE criterion, then, is that it is far too imprecise. A difference of 3 points on the HAM-D cannot be detected by clinicians among their patients at all. A more sensible criterion would be the change on the HAM-D that corresponds to a CGI-I rating of "minimal improvement". According to the data of Leucht et al., a rating of "minimal improvement" corresponds to a reduction of 7 points on the HAM-D.
At this point I should point out the difference between statistical and clinical significance. Statistical significance refers to the reliability of an effect. Is it a real effect, or just chance? Statistical significance says nothing about the size of the effect. Clinical significance, on the other hand, has to do with the size of the effect and with whether it matters in a person's life. Imagine, for example, that a study of 500,000 people showed that smiling increases life expectancy by five minutes. With 500,000 subjects I can pretty much guarantee that this difference will be statistically significant; clinically, however, it is meaningless.
Since then our analyses have been repeated a number of times (Fountoulakis & Möller, 2011; Fournier et al., 2010; NICE, 2004; Turner et al., 2008). Our data were used for some of the replications; others analysed different clinical trials. The FDA even carried out its own meta-analysis of all approved antidepressants (Khin et al., 2011). Despite all the differences in how the data were handled, the figures have remained surprisingly stable. The differences on the HAM-D have stayed small throughout, always below the value that corresponds to a CGI-I rating of "no change" (see Fig. 4). Thomas P. Laughren, director of the FDA's Division of Psychiatry Products, admitted as much on the American television news programme 60 Minutes: "I think we would all agree that the changes seen in the short-term trials, namely the difference in improvement between drug and placebo, are rather small."
It is not only the short-term trials that show a small, clinically insignificant difference between drug and placebo. In its meta-analysis of published clinical trials, NICE (2004) found that even in long-term studies the difference between drug and placebo was no greater than in the short-term ones. The difference between drug and placebo is small, so small that practitioners cannot detect it at all.
The severity of depression and the effectiveness of antidepressants
Criticism of our 2002 meta-analysis argued that our results were due to clinical trials on subjects who were less depressed. A more substantial difference would surely be found in severely depressed patients. That criticism was in fact what prompted my colleagues and me to analyse the 2008 FDA data again (Kirsch et al., 2008). We classified the clinical trials in the FDA database by the patients' initial severity of depression, using the traditional categories for depression. Only one study was found to have been carried out on patients with moderate depression, and there was no meaningful difference between drug and placebo in it; in fact the difference was practically zero (0.07 points on the HAM-D). All the other studies were carried out on patients whose mean baseline level of depression was "very severe", and even in these patients the difference between drug and placebo was below the level of clinical significance
Severity did matter, however. Patients with extremely severe depression, whose HAM-D score was at least 28, showed a mean difference between drug and placebo of 4.36 points. That is above the criterion of clinical significance proposed by NICE (2004), but well below the difference of 7 points that corresponds to a CGI-I rating of "minimal improvement".
To find out how many patients fell into the severely depressed group, I asked Mark Zimmerman of Brown University Medical School to give me the raw data from a study in which he and his colleagues had assessed the HAM-D scores of patients diagnosed with unipolar major depressive disorder (MDD) after presenting at an outpatient psychiatric practice (Zimmerman et al., 2005). Only about 10 % of these patients had HAM-D scores of 28 or more. This suggests that 90 % of depressed patients get no clinically meaningful effect from the antidepressants prescribed to them (the placebo effect and antidepressants).
Even so, this figure of 10 % may overstate the number of people who benefit from antidepressants. Antidepressants are also prescribed to people who are not severely depressed. My neighbour's beloved dog died, and the doctor prescribed him an antidepressant.
I have lost count of how many people have told me that they were prescribed an antidepressant for insomnia, even though insomnia is a common side effect of antidepressants. About 20 % of insomnia sufferers in the United States are prescribed antidepressants as a medication by their family doctor (Simon & VonKorff, 1997), although "this popularity of antidepressants for insomnia is by no means supported by any convincing body of data, but only by the opinion and the beliefs of the doctors who prescribe them" (Wiegand, 2008).
Attempts by other researchers to assess the relationship between baseline severity and the difference between drug and placebo have produced varying results. Some arrive at the same relationship as we did (Fournier et al., 2010; Khin et al., 2011, for example), while others find no relationship between the severity of the illness and the difference between drug and placebo (Fountoulakis et al., 2013; Locher et al., 2015, for example). But all the meta-analyses find overall differences between drug and placebo that are below the NICE criterion of clinical significance. So the question is this: is there a subgroup of severely depressed patients for whom antidepressants are clinically effective, or do they lack effectiveness at every level of severity?
Predicting the response to treatment
One of the few predictors of the response to treatment for depression is its severity. The type of antidepressant has almost no bearing on the success of the treatment. The following summary of meta-analyses from 2011 compares antidepressants:
"Based on 234 studies, no substantial differences in efficacy or effectiveness were found in the treatment of acute, continuing and chronic MDD. No differences in efficacy were found in patients with side effects, or in groups divided by age, sex, ethnicity or comorbidity [...] The current findings do not justify recommending any particular second-generation antidepressant on the basis of differences in efficacy" (Gartlehner et al., 2011)
Although the nature of the drug makes no clinically meaningful difference to the outcome, it does affect the placebo response. Almost all antidepressant trials contain a run-in phase at the start. Before the trial proper begins, all patients receive a placebo for a week or two. After this initial phase the patients are examined again, and anyone who has improved substantially is excluded from the rest of the trial. That leaves the patients who got no benefit at all from the placebo and those who got only a small benefit. These patients are then assigned at random.
The group is divided into those who receive the drug and those who receive placebo. As it turns out, the patients who show at least a small improvement in the initial phase are the ones who most often respond to the actual drug. This is confirmed not only by the clinicians' ratings but also by changes in brain function (Hunter et al., 2006; Quitkin et al., 1998).
How did these drugs get approved?
How did it come about that drugs with such low effectiveness were approved by the FDA? To understand that, you need to know the criteria the FDA uses for approval. The FDA requires two adequately conducted clinical trials that show a significant difference between drug and placebo. But there is a loophole: there is no limit on the number of trials that may be run in the search for two significant ones. Trials with a negative result simply do not count! Nor is the clinical significance of the result taken into account. The only thing that matters is the statistical significance of the findings.
The most outrageous example of these criteria at work can be seen in the FDA's approval of Viibryd in 2011. Seven controlled efficacy trials were carried out. The first five missed any significant difference on any measure of depression, and the mean difference between drug and placebo in these studies was less than half a point on the HAM-D; in two of these five trials the difference was in favour of placebo. The company ran two more trials and managed to obtain small but significant differences between drug and placebo (1.70 points). The mean difference across all seven trials was 1.01 points on the HAM-D. That was enough for the FDA to grant approval and to give doctors and patients the following information: "The efficacy of VIIBRYD was established in two 8-week, randomised, double-blind, placebo-controlled trials." The trials that preceded the two successful ones were never mentioned.
The failure to mention the unsuccessful trials was not a mere oversight; it reflects a carefully worked-out principle that has been in place for decades. As far as I know, there is only one antidepressant for which the FDA also reported the existence of negative trials. That is citalopram. This information came about because of an objection raised by Paul Leber, who was then director of the FDA's Division of Neuropharmacological Drug Products. In an internal memorandum of 4 May 1998, Leber wrote:
"One aspect of the labelling deserves special mention. [The report] not only describes the clinical trials that provide evidence of citalopram's antidepressant effect, but also mentions the adequate and well-controlled trials that failed to do so [...] The Division Director is inclined to believe that such information is of no practical use either to the patient or to the prescribing doctor. I disagree. I believe that it is useful for the doctor, for the patient and for the third-party payer, who has no direct access to the FDA's official reports, to learn that citalopram's antidepressant effect was not in fact demonstrated in every controlled clinical trial that was meant to show it. I am aware that clinical trials often fail to confirm the effectiveness of effective drugs. I doubt, however, that the public, or even most of the medical community, is aware of this fact. I am convinced that they not only have a right to know, but ought to know. Moreover, I believe that this selective labelling, which describes only the positive trials and leaves out the negative ones, can be regarded as potentially "false and misleading".
Hats off to Paul Lebert!
Hats off to Paul Leber! I have never met this man of honour and never corresponded with him, but because of that memorandum he is one of my heroes.
Translation of an article by Irving Kirsch
If you have questions, or if you would like to discuss this subject, write to me.







