GuidesHow-to

Teaching with the PHQ-9 and GAD-7: Screening, Sensitivity and Specificity, and the Base-Rate Fallacy

A session that uses two 2-minute scales to teach, with numbers, that screening is not diagnosis. Kroenke et al.'s PHQ-9 cut-off of 10 (88% sensitivity and specificity, 2001), Spitzer et al.'s GAD-7 (89%/82%, 2006), Levis et al.'s large meta-analysis (2019), a 1,000-student calculation that yields a 45% positive predictive value, and the safeguards a classroom needs.

Soundary · 9 min read · Updated

Depression and anxiety screening scales are a rare teaching resource that fits a health class, the abnormal-psychology unit of an intro course, and the conditional-probability unit of a statistics course alike. The items are short, the cut-offs are explicit, and sensitivity and specificity are printed in the papers, so students can compute for themselves why a positive screen is not a diagnosis. The session must be designed on the assumption that someone in the room is struggling right now. This guide covers what the two scales measure, how to run one session with safeguards, and which numbers to take from where.

What they measure

The PHQ-9 asks how often nine depressive symptoms occurred over the past two weeks, 0–3 each, for a total of 0–27 (Kroenke, Spitzer & Williams, 2001). The GAD-7 does the same for seven anxiety symptoms, 0–21 (Spitzer, Kroenke, Williams & Löwe, 2006). Both ask about 'the past two weeks', so they measure a state rather than a trait — which is why a short-interval pre/post comparison, such as before and after exams, is meaningful.

Scale0–45–910–1415–1920+
PHQ-9 (0–27)MinimalMildModerateModerately severeSevere
GAD-7 (0–21)MinimalMildModerateSevere (15–21)

Both papers proposed a score of 10 or more as a positive screen. At 10, the PHQ-9 had 88% sensitivity and 88% specificity for major depression (Kroenke et al., 2001); the GAD-7 had 89% sensitivity and 82% specificity for generalized anxiety disorder (Spitzer et al., 2006). Later, Levis, Benedetti & Thombs (2019) pooled individual participant data from many countries and confirmed that 10 is the cut-off that maximizes the sum of sensitivity and specificity. These four numbers are today's raw material.

One session, with safeguards

  1. Before class: create a group from the test page ('Create a group link'; test = PHQ-9 or GAD-7, retest = 2 weeks). Prepare one slide with the school's mental-health pathway (counseling office, responsible teacher, outside services), and if the class is under 18, check the guardian-consent rules first.
  2. Briefing, 5 min: make it explicit that participation is voluntary and can be skipped, that each score is visible only to its owner — not to the teacher — and that the group sees only an anonymous distribution. Say in advance that the last PHQ-9 item asks about thoughts of self-harm, and show the help-pathway slide now.
  3. Test, 2 min: quietly, on their own devices. Those who finish keep the result screen, including its professional-help note, open while they wait.
  4. Distribution, 5 min: the group hub shows anonymous counts by score band. Count only how many are at 10 or above, and never ask who they are. That count is the 'positive screens' for the next step.
  5. Lecture, 12 min: screening vs diagnosis (a screen is a sieve for who to look at more closely), sensitivity (the share of true cases caught) and specificity (the share of non-cases correctly cleared), and how the share of positives that are real (positive predictive value) depends on prevalence.
  6. Calculation, 12 min: work through the '1,000-student calculation' below on the board. Start with 1,000 students, 10% prevalence, 88% sensitivity and specificity, and arrive at 88 true cases among 196 positives = 45% positive predictive value. Then have them redo it with 3% prevalence.
  7. Wrap-up, 4 min: show the help-pathway slide once more and announce the retest in two weeks (ideally straddling an exam period). Leave time after class — a student may come to you individually.

The 1,000-student calculation — a 45% positive predictive value

Screen ↓ / Truth →Actually has it (100)Does not (900)Total
Screen positive (10+)88 (88% sensitivity)108 (false positives)196
Screen negative12 (missed)792 (88% specificity)804

Of 196 positives, 88 actually have a depressive disorder — 45%. Even with a fairly good test at 88% sensitivity and 88% specificity, more than half of positive screens do not meet diagnostic criteria. Conversely, only 12 of the 804 negatives (1.5%) are missed, so the test is strong at reassurance and weak at confirmation. At 3% prevalence the positive predictive value drops much further — which is why a positive screen must lead to a fuller assessment.

Discussion questions

  • If a test with 88% sensitivity has only a 45% positive predictive value, does a school-wide screening program help or harm? What happens to the 108 false positives?
  • If the cut-off rises from 10 to 15, which way do sensitivity and specificity each move? When is missing a case the bigger cost, and when is a false alarm?
  • If a 'past two weeks' scale runs high during exam season, is that a flaw in the test or the test working as designed?
  • Our class group shows N people at 10 or above and nothing else. How should a school act on that? How do you increase help without identifying individuals?

FAQ

Can it be used with high-school or middle-school students?

The original papers studied adult primary-care patients. For adolescents, follow the school's mental-health policy and guardian-consent rules first, never compel participation, and have a crisis pathway ready. If the goal is to teach the statistics, the '1,000-student calculation' alone makes a lesson without anyone taking the test.

How can a teacher know which students scored high?

You cannot, by design. That is why the right approach is to repeat the help pathway to everyone and keep time open for students to come individually. If identifying individuals is required, that is the job of the school's counseling system, not a class.

Why is a positive screen not a diagnosis?

A diagnosis comes from a clinical interview that covers duration, functional impairment, and the exclusion of other causes (medical illness, substances, bereavement). A screening scale self-reports symptom frequency only. As the calculation shows, even a good scale yields positives of which about half do not meet criteria — a positive screen is a ticket to a fuller assessment, not a conclusion.

Related tests

References

  1. Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606–613.
  2. Spitzer, R. L., Kroenke, K., Williams, J. B. W., & Löwe, B. (2006). A brief measure for assessing generalized anxiety disorder: The GAD-7. Archives of Internal Medicine, 166(10), 1092–1097.
  3. Levis, B., Benedetti, A., Thombs, B. D., & the DEPRESsion Screening Data (DEPRESSD) Collaboration. (2019). Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: Individual participant data meta-analysis. BMJ, 365, l1476.

This article is general information written by Soundary from published literature and diagnostic criteria. It is not a medical diagnosis or treatment recommendation for any individual. If you are concerned about symptoms, please consult a mental health professional.