← All guides

The Depression Screening Questionnaire, Explained

Nine questions, two weeks, one number — and a clear line between what that number screens for and what it cannot diagnose.

The short answer

The standard depression screening questionnaire is the PHQ-9: nine items matching the diagnostic criteria for major depression, each rated for how often it applied over the past two weeks, from not at all to nearly every day. Totals run 0 to 27, with conventional bands at 5, 10, 15 and 20 marking mild, moderate, moderately severe and severe. A score of 10 or above is the usual threshold for further assessment: pooled across 58 validation studies it detects about 88 percent of major depression cases and correctly clears about 85 percent of people without it, which still leaves a substantial number flagged who do not have it. That trade is deliberate — screening is built to avoid missing cases, not to be right about every one — which is why the result is a reason for a conversation with a clinician rather than a diagnosis. If item 9 is anything other than zero, that is worth raising with someone now, not at some later point.

What the nine items cover and why those nine

The PHQ-9 is not a list of things that seemed relevant. Its nine items map directly onto the nine diagnostic criteria for a major depressive episode, which is what separates it from the large number of online depression quizzes with no defined relationship to anything. Two items cover mood and interest — low mood or hopelessness, and loss of pleasure in things. Four cover physical symptoms: sleep in either direction, energy, appetite in either direction, and observable changes in how fast you move or speak. Two cover self-view and concentration. The ninth asks about thoughts of being better off dead or of hurting yourself.

The two-week window is part of the instrument rather than an arbitrary choice, because duration is part of what distinguishes a depressive episode from an ordinary bad stretch. Answering for how you feel today, or for how the year has gone, produces a number that does not correspond to the bands. This is the most common way the questionnaire gets misread by people taking it on their own.

The symmetry in the physical items catches something people miss about depression. Sleeping too much counts the same as not sleeping; overeating counts the same as no appetite. The criterion is a change in either direction, not a deficit, so someone sleeping ten hours and eating more than usual can score as high on that facet as someone doing the opposite.

How to read the total, and what the bands are worth

The conventional cut points are 5, 10, 15 and 20, dividing the range into minimal, mild, moderate, moderately severe and severe. Ten is the threshold that matters most in practice: at 10 or above the instrument has sensitivity of about 0.88 for major depression and specificity of about 0.85, which means a positive screen is meaningfully more likely to be a real case than not — and still wrong often enough that treating it as settled would be a mistake (Levis et al., 2019).

The bands are wider than they look precise. A total of 9 and a total of 11 differ by one item moved one step, which is well inside the noise of how you happened to read a sentence on a given day, and they sit in different bands. What is worth attending to is not the band boundary but whether the number is somewhere near the threshold at all, and whether it is moving between administrations.

The strongest use of the instrument is exactly that repeated one. Because the item set is fixed and the window is defined, the same questionnaire taken again several weeks later gives a difference score that is interpretable — a drop of five points or more is the usual marker of meaningful improvement, and it is how the PHQ-9 is used clinically to track whether a treatment is working. That comparison does not require any reference sample, which makes it the one comparison a self-administered version can support properly.

How accurate is the threshold, exactly?

This has a precise answer, from the largest study of the question. An individual participant data meta-analysis pooled raw data from 58 studies covering 17,357 people, 2,312 of whom had major depression confirmed by a structured diagnostic interview, and calculated the questionnaire's accuracy at every cut point from 5 to 15 (Levis et al., 2019). Combined accuracy was maximised at a total of 10 or above, which is where the conventional threshold comes from — it was derived from data, not chosen for tidiness.

At that threshold, against clinician-administered semistructured interviews as the reference standard, sensitivity was 0.88 and specificity was 0.85 across 29 studies and 6,725 participants. In plain terms: it catches around 88 percent of genuine major depression cases, and correctly clears around 85 percent of people who do not have it. Both of those failure rates are large enough to matter for an individual, which is the whole reason a positive screen routes to an assessment rather than to a conclusion.

The finding most worth knowing is that these numbers depend on what the questionnaire was checked against. Sensitivity ran 5 to 22 percent higher when the reference standard was a semistructured interview conducted by a clinician than when it was a fully structured interview designed for lay administration. Specificity was similar across all of them. That is a striking result: a substantial part of what looks like the instrument's accuracy is a property of how carefully the comparison diagnosis was made, and the earlier conventional meta-analyses that pooled reference standards together understated its sensitivity as a consequence.

One thing the same analysis settled usefully: a cut point of 10 works across age groups. The questionnaire appeared similarly sensitive but somewhat less specific in younger patients than in older ones, and the authors concluded the threshold does not need adjusting by age. If a site offers you an age-adjusted depression cutoff, that adjustment is not coming from this evidence base.

What screening establishes and what it does not

Screening establishes that a conversation is warranted. It cannot establish a diagnosis, and the gap between those two is not a formality. A diagnosis requires ruling out the things that produce the same nine symptoms — thyroid dysfunction, anaemia, sleep apnoea, medication effects, grief, sustained sleep deprivation — and it requires a judgment about functional impairment that no questionnaire asks about. Someone can score 15 because of untreated apnoea and be no less unwell for it not being depression, but the treatment is a different one.

The instrument is also less accurate in some populations than the headline figures suggest, and it is worth knowing which. Somatic items behave differently in people with chronic physical illness, where fatigue and sleep disruption have an obvious non-psychiatric cause and inflate the total. Cultural variation in how distress is expressed affects which items get endorsed. And self-report has the usual ceiling: it captures what you are willing to report about a period you may be remembering through the lens of how you feel right now.

What it does have, more than nearly any other psychological instrument, is evidence behind it. It has been validated across primary care, specialist and general population samples, translated and re-validated in many languages, and used as the outcome measure in a large share of depression treatment trials. When people describe it as the most widely used depression screener in clinical practice, that is a statement about tens of thousands of studies rather than a marketing line.

Item nine, and what to do if it is not zero

The ninth item asks about thoughts of being better off dead or of hurting yourself, and it is treated differently from the other eight for a reason: it is scored into the total and it also stands alone. Any response above not at all is a flag regardless of what the total came to, and a low total with a non-zero item 9 is not a reassuring result.

If that applies to you right now, contact a crisis line rather than waiting for a scheduled appointment. In the US, call or text 988 for the Suicide and Crisis Lifeline. Elsewhere, findahelpline.com lists services by country. If you are in immediate danger, contact emergency services. This is the one place where the correct response to a questionnaire result is to stop reading about measurement and talk to a person.

For the other eight items, a score in the moderate range or above is worth taking to a GP or a clinician, and taking the completed answers with you is more useful than taking the total — which items you endorsed tells them more than the sum. If your score is low but something still feels wrong, that is also worth raising. The instrument screens for one condition out of many, and a negative screen for depression is not a statement that nothing is happening.

Common questions

What is a normal score on a depression screening questionnaire?
On the PHQ-9, totals below 5 are conventionally read as minimal, and that is where most people without a current depressive episode fall. Scores of 5 to 9 indicate mild symptoms that often do not require treatment but are worth watching, and 10 is the usual threshold for further assessment. A score is not a rank against other people — it is a count of symptom frequency over two weeks, so the meaningful comparison is against the clinical cut points and against your own earlier results.
Can an online depression test diagnose depression?
No. Screening questionnaires identify people who should be assessed further; a diagnosis requires a clinician to rule out physical causes that produce identical symptoms, such as thyroid problems, anaemia or sleep disorders, and to judge whether functioning is impaired. Even at a well-chosen threshold a positive screen is wrong a meaningful fraction of the time. The right use of a high score is to bring it, and the specific items you endorsed, to a GP or clinician rather than to conclude anything from it directly.
How often should I retake a depression screening questionnaire?
Every two to four weeks is the interval that makes sense, because the questionnaire asks about the past two weeks and anything shorter overlaps the window it is measuring. That spacing is how the instrument is used clinically to track whether treatment is working, with a drop of five points or more taken as meaningful improvement. Retaking daily measures mood fluctuation rather than the thing the bands were built for, and tends to make an ordinary bad day look like deterioration.
What is the difference between the PHQ-9 and the PHQ-2?
The PHQ-2 is the first two items — low mood and loss of interest — used as an ultra-brief first pass. It is sensitive enough to rule out depression reasonably well when both answers are low, which makes it useful when time is very short, but it carries no severity information and misses presentations where the physical or cognitive symptoms dominate. If the PHQ-2 is positive, the standard next step is the full nine items, which is also what gives you a total that can be tracked over time.

Measure it on yourself

Reading about a trait and seeing your own score are different things. These assessments cover what this article describes.

Read next

Sources

  1. Kroenke, K., Spitzer, R. L., & Williams, J. B. (2001). The PHQ-9: validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9).
  2. Levis, B., Benedetti, A., & Thombs, B. D. (2019). Accuracy of the PHQ-9 for screening to detect major depression: individual participant data meta-analysis. BMJ, 365.
  3. Manea, L., Gilbody, S., & McMillan, D. (2012). Optimal cut-off score for diagnosing depression with the PHQ-9: a meta-analysis. CMAJ, 184(3).
  4. Löwe, B., Kroenke, K., Herzog, W., & Gräfe, K. (2004). Measuring depression outcome with a brief self-report instrument: sensitivity to change of the PHQ-9. Journal of Affective Disorders, 81(1).
  5. Levis, B., Benedetti, A., & Thombs, B. D. (2019). Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: individual participant data meta-analysis. BMJ, 365, l1476.
  6. Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606-613.
  7. Radloff, L. S. (1977). The CES-D scale: A self-report depression scale for research in the general population. Applied Psychological Measurement, 1(3), 385-401.
  8. Spitzer, R. L., Kroenke, K., Williams, J. B. W., & Löwe, B. (2006). A brief measure for assessing generalized anxiety disorder: the GAD-7. Archives of Internal Medicine, 166(10), 1092-1097.

Last reviewed 2026-08-27. This article is general information about psychological measurement, not medical or psychological advice.