Abstract

The Patient Health Questionnaire, which MeSH classifies under psychological tests, is a family of brief self-report instruments whose best-known member, the nine-item PHQ-9, maps each DSM criterion for major depression onto a question scored 0 to 3, so a total from 0 to 27 doubles as a severity index and a screening tool. This article treats the instrument as a case study in the psychometrics of self-report: how the same answers grade severity and screen for a probable disorder, and why the conventional cut-off of 10 is a decision threshold rather than a fact about the person. Three interactive demonstrations let a reader assemble a PHQ-9 score, slide the cut-off to trade sensitivity against specificity, and watch the positive predictive value collapse as the disorder grows rarer.

Keywords: depression screening, self-report, severity measurement, diagnostic accuracy, primary care

The Patient Health Questionnaire occupies an unusual place in clinical measurement: it is short enough to complete in a waiting room, transparent enough that a patient can see exactly what each answer contributes, and consequential enough that its total can route a person toward or away from treatment. That combination of brevity and weight is what makes it worth close study. A measure this light is easy to over-read, and much of the literature on the PHQ is really an argument about how much interpretive load a nine-item checklist can bear.

The instrument's authority rests on a chain of validation studies rather than on any single result. Each link in that chain — the original criterion validation against a structured interview, the meta-analyses that pooled dozens of such studies, and the individual-participant re-analyses that corrected them — is a step in an ongoing effort to say precisely what a PHQ score means and how far it can be trusted. Reading the instrument well means reading that chain, not just the number it produces.

Key Takeaways
  • The PHQ is a family of self-report scales; the PHQ-9 scores each of the nine DSM depression criteria from 0 to 3 for a total of 0 to 27.
  • The same nine items serve two jobs at once — grading severity and screening for a probable disorder — which is the source of most confusion about how to read them.
  • A screening cut-off is a chosen trade between catching true cases and raising false alarms, not a boundary that exists in the patient.
  • The conventional threshold of 10 was validated against structured diagnostic interviews, but later pooled analyses showed published accuracy figures were inflated by choosing cut-offs after seeing the data.
  • A positive PHQ result is a prompt for a diagnostic conversation, never a diagnosis on its own.

What the Patient Health Questionnaire Is

The Patient Health Questionnaire is the self-administered descendant of PRIME-MD, a structured diagnostic interview built in the early 1990s so that a busy primary-care physician could screen for the common mental disorders without the hour a full psychiatric interview demanded (Spitzer et al., 1999). The original PRIME-MD still needed a clinician to ask the questions. The insight behind the PHQ was that most of the interview could be handed to the patient: if the questions were written plainly enough, a person could rate their own symptoms on paper, and the physician could read the completed page in under a minute. The full PHQ covers several disorder modules, but the depression module — the PHQ-9 — became the instrument that carries the family name.

The PHQ-9 is deliberately literal about its construct. Each of its nine items restates one of the nine symptom criteria the Diagnostic and Statistical Manual lists for a major depressive episode, from depressed mood and loss of interest through sleep, appetite, and concentration disturbance to thoughts of self-harm (Kroenke et al., 2001). The respondent rates how often each symptom has troubled them over the past two weeks on a four-point frequency scale: not at all (0), several days (1), more than half the days (2), and nearly every day (3). Because the items and the diagnostic criteria are in one-to-one correspondence, the questionnaire can be read two ways from the same responses — as a severity score by summing all nine items, or as a diagnostic algorithm by counting how many items reach the more than half the days threshold.

Two shorter relatives extend the family. The PHQ-2 keeps only the first two items — depressed mood and loss of interest, the cardinal symptoms — as an ultra-brief first-stage screen that flags who should complete the full nine (Kroenke et al., 2003). The PHQ-15 leaves depression aside to tally fifteen somatic complaints, giving a parallel measure of physical symptom burden (Kroenke et al., 2002). A systematic review of the whole family found the somatic, anxiety, and depression scales each performed as self-contained measures while sharing the same waiting-room economy (Kroenke et al., 2010). Across the family the design principle is constant: keep the item count low enough that a patient will actually finish, and keep the scoring transparent enough that a clinician can trust what the total is made of.

Severity and Screening From One Set of Answers

The dual reading of the PHQ-9 is its defining feature and its most persistent trap. As a severity measure, the summed 0-to-27 total is banded into conventional ranges — minimal, mild, moderate, moderately severe, and severe — so that a change in the number tracks a change in symptom load over the course of treatment (Kroenke & Spitzer, 2002). Summing the nine items into one total is only defensible because the items hang together as a scale, a property summarised by the internal-consistency coefficient that measurement theory contributed to psychology in the middle of the last century (Cronbach, 1951). This is the use for which the instrument is least controversial: a patient whose score falls from 18 to 8 across a course of therapy has almost certainly improved, and the scale was shown to be sensitive to exactly that kind of change in treated populations (Löwe et al., 2004).

Table 1. PHQ-9 severity bands and their conventional interpretation.

Total score Severity band Screening verdict (cut-off 10) Conventional interpretation
0-4 Minimal Negative No or negligible depressive symptoms; no action indicated.
5-9 Mild Negative Subthreshold symptoms; watchful waiting or repeat measurement.
10-14 Moderate Positive Warrants a diagnostic assessment; treatment planning may begin.
15-19 Moderately severe Positive Active treatment usually indicated on assessment.
20-27 Severe Positive Prompt treatment and close follow-up indicated on assessment.

Note. Bands and interpretations follow the instrument's conventional scoring; the screening verdict reflects the standard cut-off of 10 and is a prompt for assessment, not a diagnosis. Original schematic.

As a screen, the same total is compared against a single cut-off, and everything hangs on where that line is drawn. A screening threshold partitions a continuous score into two verdicts — likely case, likely not — and any such partition trades two kinds of error against each other. Set the cut-off low and the screen catches nearly everyone with the disorder but also flags many who do not have it; set it high and the false alarms fall away but genuine cases slip through. The conventional PHQ-9 threshold of 10 was chosen as a compromise between these errors, and the early validation reported that it caught roughly seven to nine of every ten true cases while correctly clearing a similar fraction of those without the disorder (Kroenke et al., 2001).

The distinction between these two uses is not academic. A severity score describes; a screening verdict decides. When the two are collapsed — when a total of 12 is treated as a diagnosis of moderate depression rather than as a prompt to look closer — the instrument is being asked to do something its design does not support. Every careful account of the PHQ insists on the same boundary: the questionnaire identifies people who warrant a diagnostic assessment; it does not perform the assessment.

Reading a Screening Score

To read a PHQ-9 score honestly is to hold two numbers in mind that the single total hides. Sensitivity is the fraction of people who truly have the disorder that the screen correctly flags; specificity is the fraction of people who truly do not have it that the screen correctly clears. A cut-off fixes both at once, and moving the cut-off moves them in opposite directions. Reporting only one — announcing that a threshold has 88% sensitivity while staying quiet about its specificity — tells half the story and usually the more flattering half.

Because sensitivity and specificity move together as the threshold shifts, no single cut-off captures how well the instrument separates cases from non-cases. Plotting sensitivity against the false-positive rate at every possible cut-off traces the receiver operating characteristic curve, and the area under that curve condenses the entire trade-off into one cut-off-independent number — the probability that a randomly chosen true case scores higher than a randomly chosen non-case (Hanley & McNeil, 1982). It is this area, not the accuracy at any one threshold, that the diagnostic meta-analyses report as their headline measure of the PHQ-9's discriminative power.

The meta-analytic literature made this trade explicit. Pooling the diagnostic studies available by the mid-2000s, an early synthesis found the PHQ performed respectably as a screen but that its reported accuracy varied widely with setting and with the reference standard used (Gilbody et al., 2007). A later meta-analysis focused squarely on the cut-off itself, showing that thresholds from 8 to 11 all had defensible claims and that the optimal value depended on whether a service most wanted to avoid missing cases or to avoid over-referring (Manea et al., 2012). A further synthesis reframed the whole exercise as case-finding rather than diagnosis, arguing that the instrument's proper output is a shortlist for clinical attention, not a label (Moriarty et al., 2015).

There is a subtler statistic beneath sensitivity and specificity that decides what a positive screen is actually worth to the patient: the positive predictive value, the probability that someone who screens positive really has the disorder. That value depends not only on the test but on how common the disorder is in the population being screened. In a setting where depression is rare, even a good screen produces many more false positives than true ones, because there are so many more well people to misclassify. The same PHQ-9 cut-off that looks trustworthy in a psychiatric clinic can be mostly false alarms in a general population — a fact the raw sensitivity figure never shows.

The Cut-Off Was Chosen, Not Discovered

The most important lesson the PHQ taught measurement science is a cautionary one about how accuracy gets reported. When a study collects PHQ-9 scores alongside a gold-standard diagnostic interview, it can compute the sensitivity and specificity of every possible cut-off and then report the one that looks best. Choosing the threshold after seeing which threshold performs best in that particular sample inflates the reported accuracy, because the chosen number is partly fitted to the noise of that one dataset and will not perform as well on the next.

The correction came from individual-participant-data meta-analysis, which pools the raw records of many studies rather than their published summaries. Re-analysing the primary data, a large synthesis showed that when a single pre-specified cut-off of 10 was applied uniformly across all studies, the PHQ-9's accuracy was real but more modest than the data-driven figures had suggested (Levis et al., 2019). An updated and enlarged version of that work confirmed the finding on a still larger pooled sample and sharpened the case for a fixed threshold rather than a sample-optimised one (Negeri et al., 2021). A parallel systematic review in primary care reached a compatible conclusion about the instrument's real-world screening performance (Costantini et al., 2021).

The methodological moral generalises well beyond depression. Any diagnostic instrument whose threshold is chosen to maximise accuracy in the same sample used to report that accuracy will look better than it is. The PHQ became the textbook case not because it was uniquely flawed but because it was studied so heavily that the inflation could be measured and undone. That its corrected accuracy survived the correction is the strongest evidence for the instrument; that the correction was necessary at all is the strongest caution about reading any single validation study.

Figure 1. The sensitivity-specificity trade-off across PHQ-9 cut-off scores.
Sensitivity and specificity plotted against PHQ-9 cut-off score As the cut-off rises from 5 to 15, sensitivity falls and specificity rises, the two curves crossing near a cut-off of 10. PHQ-9 cut-off score proportion correct 5 8 10 12 15 1.0 0.7 0.4 cut-off 10 sensitivity specificity

Note. Sensitivity (solid) falls and specificity (dashed) rises as the cut-off increases; the curves cross near the conventional threshold of 10, where neither error is minimised alone. Curve shapes are illustrative of the published trade-off rather than fitted values. Original schematic.

The Instrument in Motion

The three demonstrations below make manipulable the parts of the instrument that the prose can only describe. The first assembles a PHQ-9 total from its nine items and shows how the same responses yield a severity band and a screening verdict at once. The second slides the cut-off along the score range and watches sensitivity and specificity move against each other. The third shows how the positive predictive value of a fixed cut-off collapses as the disorder becomes rarer in the screened population.

Score a PHQ-9: severity and screening from one page

Rate each symptom 0 (not at all) to 3 (nearly every day). The total is read two ways at once.

Little interest or pleasure (cardinal)
Feeling down or depressed (cardinal)
Sleep disturbance
Feeling tired, low energy
Appetite change
Feeling of failure
Trouble concentrating
Psychomotor change
Thoughts of self-harm
Severity total
14 / 27
Moderate
Screen (cut-off 10)
Positive
DSM algorithm
Meets
5 symptoms at 2 or 3, cardinal present

The summed severity and the symptom count are two different functions of the same nine answers, which is why the band and the algorithm can disagree about the same patient.

The scoring demonstration makes the dual reading concrete. Setting every item to nearly every day drives the total to its ceiling of 27 and lands in the severe band; a scattered profile of ones and twos can reach the same screening cut-off as a concentrated profile of a few threes, which is precisely why the severity band and the symptom-count algorithm sometimes disagree about the same patient.

Move the cut-off, watch the two errors trade

cut 10non-cases (teal) · cases (rust)
Sensitivity 0.84
true cases caught
Specificity 0.81
well people cleared

Lowering the cut-off catches more true cases but clears fewer well people; raising it does the reverse. No position maximises both at once.

The cut-off demonstration is the sensitivity-specificity trade in motion. Dragging the threshold left recovers more true cases at the cost of more false alarms; dragging it right does the reverse. There is no position that maximises both, and the demonstration makes the absence of a free lunch visible.

Why a positive screen means less when depression is rare

Positive predictive value
39%
Of 2,200 positive screens in 10,000 people, 850 are true cases and 1,350 are false alarms.

Hold the test fixed and lower the prevalence: the share of positives that are real falls steeply, because a fixed false-positive rate applied to a growing pool of well people swamps the shrinking pool of true cases.

The prevalence demonstration is the one most likely to surprise. Holding sensitivity and specificity fixed and lowering the base rate of the disorder, the share of positive screens that are true cases falls steeply, because a fixed false-positive rate applied to a growing pool of well people swamps the shrinking pool of true cases. The test has not changed; the population has.

Worked Example

Consider a patient who completes the PHQ-9 with the following nine responses, each on the 0-to-3 frequency scale: depressed mood 2, loss of interest 2, sleep disturbance 3, fatigue 2, appetite change 1, feelings of failure 2, concentration difficulty 1, psychomotor change 0, and thoughts of self-harm 1. The severity total is the simple sum of these nine values: 2 + 2 + 3 + 2 + 1 + 2 + 1 + 0 + 1 = 14. A total of 14 falls in the moderate-to-moderately-severe range and sits above the conventional screening cut-off of 10, so the screen is positive.

The diagnostic-algorithm reading counts differently. The DSM-based algorithm asks how many of the nine symptoms reached at least the more than half the days level — a rating of 2 or 3 — and requires that at least one of the two cardinal symptoms (depressed mood or loss of interest) be among them. Here the items scoring 2 or 3 are depressed mood (2), loss of interest (2), sleep disturbance (3), fatigue (2), and feelings of failure (2): five symptoms, including both cardinal ones. Five symptoms at that frequency, with a cardinal symptom present, meets the algorithm's threshold for probable major depression.

Both readings agree in this case, but notice how they could diverge. Had the same patient rated five symptoms at several days (a 1 each) and one cardinal symptom at nearly every day (a 3), the total might still approach 10 while the count of symptoms at the more than half the days level fell short of the algorithm's requirement. The summed severity and the symptom count are two different functions of the same nine numbers, and a careful reader checks both rather than trusting the total alone. The instrument's transparency is what makes this cross-check possible: every step from the nine checkboxes to the two verdicts can be recomputed by hand.

Discussion

The Patient Health Questionnaire earns its place in cognitive psychology not because it settles anything about depression but because it exposes, in unusually plain form, the reasoning that turns a self-report into a clinical decision. Its nine items are a near-verbatim translation of a diagnostic definition, so the usual gap between a construct and its measurement is narrower here than almost anywhere else in psychological testing. That narrowness is a virtue for interpretability and a hazard for over-interpretation: because the items look so much like the diagnosis, it is tempting to treat the score as the diagnosis, which is exactly the error the instrument's own authors spent two decades warning against.

The deeper contribution is methodological. The PHQ's validation history is a compressed lesson in how diagnostic accuracy should and should not be reported — how a cut-off chosen after the fact flatters the instrument, how pooling raw participant data corrects the flattery, and how an instrument can emerge from that correction still useful but honestly smaller than its first press. A field that internalises this lesson reads every new screening tool more skeptically, and reads the PHQ itself as a decision aid whose output is a conversation, not a verdict.

Current Directions

The frontier of PHQ research is the individual-participant-data program that recovered the instrument's true operating characteristics. The DEPRESSD collaboration assembled the raw records of scores of primary studies and re-ran the accuracy analysis with a single pre-specified cut-off, showing that the published, data-driven figures had overstated performance and that a fixed threshold of 10 gives a more defensible if more modest estimate (Levis et al., 2019). The enlarged update extended the pooled sample and confirmed the pattern, while also mapping how accuracy shifts across subgroups defined by age, sex, and clinical setting (Negeri et al., 2021). A contemporaneous systematic review in primary care placed those pooled estimates alongside the practical question of whether routine screening improves outcomes, a question the accuracy data alone cannot answer (Costantini et al., 2021).

Two open problems define the current work. The first is whether screening changes what happens to patients: an accurate screen is worthless if a positive result does not lead to better care, and the evidence that PHQ screening improves outcomes is thinner than the evidence that it identifies cases. The second is measurement invariance — whether a given score means the same thing across languages, cultures, and demographic groups — which the pooled data are now large enough to interrogate directly. Both questions push the instrument past the psychometrics of a single administration toward the harder question of what a score does once it enters a system of care.

Common Misconceptions

A high PHQ-9 score is a diagnosis of depression.
It is not. The score identifies a person who warrants a diagnostic assessment, and the assessment — a clinical interview that weighs history, context, and alternative explanations — is a separate step the questionnaire cannot perform (Moriarty et al., 2015).
The cut-off of 10 is the point where depression begins.
The cut-off is a chosen trade-off between missed cases and false alarms, not a natural boundary in the underlying symptom continuum. Different services legitimately use different thresholds depending on whether they most want to avoid missing cases or over-referring (Manea et al., 2012).
A validated instrument's reported accuracy is the accuracy a new setting will see.
Accuracy figures drawn from studies that chose their cut-off after seeing the data are systematically inflated; the number that survives a pre-specified, pooled re-analysis is the one to trust (Levis et al., 2019).
A high sensitivity means a positive screen is probably right.
It does not. The chance that a positive screen is a true case — the positive predictive value — depends on how common the disorder is in the screened population, and can be low even when sensitivity is high (Gilbody et al., 2007).

Glossary

Cardinal symptom.
Either of the two core depression symptoms — depressed mood and loss of interest — one of which the DSM algorithm requires before a diagnosis of major depression can be made.

Case-finding.
The use of a screen to produce a shortlist of people who warrant clinical attention, as distinct from diagnosis; the framing later authors preferred for the PHQ's proper role.

Criterion validity.
The degree to which an instrument's scores agree with an external gold standard, here a structured diagnostic interview taken as the truth about who has the disorder.

Cut-off score.
The threshold that partitions a continuous score into a positive and a negative screening verdict; for the PHQ-9 conventionally set at 10.

Diagnostic algorithm.
The rule that reads the PHQ-9 as a symptom count rather than a sum, requiring a set number of symptoms at a given frequency plus a cardinal symptom.

DSM criteria.
The nine symptoms the Diagnostic and Statistical Manual lists for a major depressive episode, each of which the PHQ-9 restates as one item.

Individual-participant-data meta-analysis.
A synthesis that pools the raw records of many studies rather than their published summaries, allowing a single cut-off to be applied uniformly and correcting data-driven inflation.

Internal consistency.
The degree to which the items of a scale measure the same underlying construct, commonly summarised by coefficient alpha.

PHQ-15.
The fifteen-item somatic-symptom module of the family, measuring physical symptom burden rather than depression.

PHQ-2.
The two-item ultra-brief screen made of the PHQ-9's cardinal symptoms, used as a first stage to decide who completes the full nine.

PHQ-9.
The nine-item depression module of the Patient Health Questionnaire, scored 0 to 27, serving as both a severity measure and a screen.

Positive predictive value.
The probability that a person who screens positive truly has the disorder, which depends on the disorder's prevalence in the screened population.

PRIME-MD.
The clinician-administered structured interview for common mental disorders from which the self-report PHQ was derived.

Receiver operating characteristic (ROC) curve.
The plot of sensitivity against the false-positive rate across every cut-off; the area beneath it summarises an instrument's discrimination independent of any single threshold.

Sensitivity.
The proportion of people who truly have the disorder that a screen correctly identifies as positive.

Severity band.
One of the conventional ranges — minimal, mild, moderate, moderately severe, severe — into which a PHQ-9 total is grouped to describe symptom load.

Specificity.
The proportion of people who truly do not have the disorder that a screen correctly clears as negative.

Key Researchers

Simon Gilbody. Professor of psychological medicine and health services research at the University of York; led the diagnostic meta-analyses that quantified the PHQ-9's screening accuracy and interrogated its optimal cut-off, showing that the conventional threshold trades sensitivity against specificity rather than fixing a single true value. ORCID - Google Scholar - Faculty Page

Kurt Kroenke. Professor of medicine at Indiana University and the Regenstrief Institute; principal architect of the PHQ family, who with Spitzer and Williams converted the clinician-administered PRIME-MD into the self-report PHQ and then authored the PHQ-9, PHQ-2, and PHQ-15 validation papers that made the instrument the most widely used depression measure in primary care. ORCID - Faculty Page

Bernd Löwe. Professor of psychosomatic medicine at the University Medical Center Hamburg-Eppendorf; extended the PHQ into treatment monitoring and somatic-symptom assessment and led its adaptation across languages and care settings, work that established the instrument's sensitivity to change over the course of treatment. ORCID - Google Scholar - Faculty Page

Robert L. Spitzer (1932-2015). Psychiatrist at Columbia University; architect of DSM-III and co-developer of PRIME-MD, the structured primary-care interview from which the PHQ was derived, whose criterion-based approach to diagnosis is why the PHQ-9's nine items map one-to-one onto the DSM criteria for major depression. Wikipedia

Brett D. Thombs. Professor of psychiatry at McGill University and the Lady Davis Institute; leads the DEPRESSD individual-participant-data meta-analyses, which pooled primary studies to show that published PHQ-9 accuracy was inflated by data-driven cut-off selection, sharpening the modern case for a fixed threshold. ORCID - Google Scholar - Faculty Page

Janet B. W. Williams. Psychiatric epidemiologist formerly at Columbia University; co-developer of PRIME-MD and the PHQ and author of the Structured Clinical Interview for DSM that served as the diagnostic gold standard against which the PHQ was validated. Wikipedia

Frequently Asked Questions

What does the PHQ-9 actually measure?
The PHQ-9 measures the frequency of the nine symptoms the DSM lists for a major depressive episode over the preceding two weeks, each rated from 0 to 3, giving a total from 0 to 27 that indexes depression severity (Kroenke et al., 2001).

Is a PHQ-9 score a diagnosis of depression?
No. A score above the cut-off flags a person who should receive a diagnostic assessment, but the questionnaire cannot make the diagnosis itself, which requires a clinical interview weighing history and context (Moriarty et al., 2015).

Why is the cut-off usually set at 10?
The threshold of 10 was chosen as a compromise between catching true cases and limiting false alarms, and pooled re-analyses confirmed it as a defensible fixed cut-off, though different services may reasonably choose other values (Manea et al., 2012).

What is the difference between the PHQ-9 and the PHQ-2?
The PHQ-2 keeps only the two cardinal items, depressed mood and loss of interest, as an ultra-brief first-stage screen, while the PHQ-9 adds the remaining seven symptoms to grade severity and support a diagnostic algorithm (Kroenke et al., 2003).

Why did the instrument's reported accuracy change over time?
Early studies chose their cut-off after seeing which value performed best, which inflated accuracy; pooled individual-participant analyses applying one pre-specified cut-off produced more modest and trustworthy figures (Levis et al., 2019).

Does a high sensitivity mean a positive screen is probably correct?
Not by itself. The chance a positive screen is a true case is the positive predictive value, which depends on how common depression is in the screened group and can be low even when sensitivity is high (Gilbody et al., 2007).

Can the PHQ-9 track whether treatment is working?
Yes. Because the summed score is sensitive to change, a falling total across a course of treatment is good evidence of improvement, which is the use for which the instrument is least contested (Löwe et al., 2004).

Is the PHQ-9 reliable across different populations?
The updated pooled analyses were large enough to examine accuracy across subgroups and broadly supported the instrument, though whether a given score means the same thing across every language and culture remains an active measurement question (Negeri et al., 2021).

References

Costantini, L., Pasquarella, C., Odone, A., Colucci, M. E., Costanza, A., Serafini, G., Aguglia, A., Belvederi Murri, M., Brakoulias, V., Amore, M., Ghaemi, S. N., & Amerio, A. (2021). Screening for depression in primary care with Patient Health Questionnaire-9 (PHQ-9): A systematic review. Journal of Affective Disorders, 279, 473-483. https://doi.org/10.1016/j.jad.2020.09.131

Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. https://doi.org/10.1007/BF02310555

Gilbody, S., Richards, D., Brealey, S., & Hewitt, C. (2007). Screening for depression in medical settings with the Patient Health Questionnaire (PHQ): A diagnostic meta-analysis. Journal of General Internal Medicine, 22(11), 1596-1602. https://doi.org/10.1007/s11606-007-0333-y

Hanley, J. A., & McNeil, B. J. (1982). The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology, 143(1), 29-36. https://doi.org/10.1148/radiology.143.1.7063747

Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606-613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x

Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2002). The PHQ-15: Validity of a new measure for evaluating the severity of somatic symptoms. Psychosomatic Medicine, 64(2), 258-266. https://doi.org/10.1097/00006842-200203000-00008

Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2003). The Patient Health Questionnaire-2: Validity of a two-item depression screener. Medical Care, 41(11), 1284-1292. https://doi.org/10.1097/01.MLR.0000093487.78664.3C

Kroenke, K., Spitzer, R. L., Williams, J. B. W., & Löwe, B. (2010). The Patient Health Questionnaire Somatic, Anxiety, and Depressive Symptom Scales: A systematic review. General Hospital Psychiatry, 32(4), 345-359. https://doi.org/10.1016/j.genhosppsych.2010.03.006

Kroenke, K., & Spitzer, R. L. (2002). The PHQ-9: A new depression diagnostic and severity measure. Psychiatric Annals, 32(9), 509-515. https://doi.org/10.3928/0048-5713-20020901-06

Levis, B., Benedetti, A., & Thombs, B. D. (2019). Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: Individual participant data meta-analysis. BMJ, 365, l1476. https://doi.org/10.1136/bmj.l1476

Löwe, B., Unützer, J., Callahan, C. M., Perkins, A. J., & Kroenke, K. (2004). Monitoring depression treatment outcomes with the Patient Health Questionnaire-9. Medical Care, 42(12), 1194-1201. https://doi.org/10.1097/00005650-200412000-00006

Manea, L., Gilbody, S., & McMillan, D. (2012). Optimal cut-off score for diagnosing depression with the Patient Health Questionnaire (PHQ-9): A meta-analysis. CMAJ, 184(3), E191-E196. https://doi.org/10.1503/cmaj.110829

Moriarty, A. S., Gilbody, S., McMillan, D., & Manea, L. (2015). Screening and case finding for major depressive disorder using the Patient Health Questionnaire (PHQ-9): A meta-analysis. General Hospital Psychiatry, 37(6), 567-576. https://doi.org/10.1016/j.genhosppsych.2015.06.012

Negeri, Z. F., Levis, B., Sun, Y., He, C., Krishnan, A., Wu, Y., Bhandari, P. M., Neupane, D., Brehaut, E., Benedetti, A., & Thombs, B. D. (2021). Accuracy of the Patient Health Questionnaire-9 for screening to detect major depression: Updated systematic review and individual participant data meta-analysis. BMJ, 375, n2183. https://doi.org/10.1136/bmj.n2183

Spitzer, R. L., Kroenke, K., & Williams, J. B. W. (1999). Validation and utility of a self-report version of PRIME-MD: The PHQ Primary Care Study. JAMA, 282(18), 1737-1744. https://doi.org/10.1001/jama.282.18.1737