Abstract
Psychiatric status rating scales, which the Medical Subject Headings vocabulary classifies under neuropsychological tests, are standardized instruments that convert the severity of psychiatric symptoms into a quantitative score. Where a diagnostic interview asks whether a disorder is present, a rating scale asks how severe it is now, so that change can be tracked over time and compared across patients and trials. This article treats the rating scale as a measurement instrument: the split between clinician-administered and self-report formats, the psychometric properties that decide whether a score can be trusted, and the response-and-remission thresholds that turn a number into a clinical decision. It covers the major depression, mania, and psychosis scales, the content heterogeneity that makes two scales disagree, and the measurement-based-care movement. Three interactive demonstrations model severity banding, treatment response and remission, and cross-scale symptom overlap.
Keywords: psychiatric rating scale, symptom severity, measurement-based care, psychometrics, treatment response
A psychiatric status rating scale is a structured instrument that assigns a number to the current severity of a defined set of psychiatric symptoms. It is not a diagnosis. A diagnostic instrument sorts patients into categories — depressed or not, psychotic or not — whereas a status rating scale assumes the condition and measures its intensity, so that the same patient can be scored repeatedly and a rising or falling total read as deterioration or improvement. The oldest of the modern scales, the Brief Psychiatric Rating Scale, was introduced to give the clinical trials of early psychopharmacology a common quantitative yardstick (Overall & Gorham, 1962), and the design problem it faced — how to turn a clinician's global impression of a patient into a reproducible number — defines the field still. The value of a scale rests entirely on its measurement properties: a total score is only useful if two raters watching the same patient would assign nearly the same number, if that number rises and falls with genuine clinical change, and if a stated cutoff means the same thing in one clinic as in another.
- A psychiatric status rating scale measures the severity of a symptom state, not the presence of a diagnosis; its purpose is to track change over time.
- Scales divide into clinician-administered instruments, which depend on a trained rater's judgment, and self-report instruments, which depend on the patient's insight and reading of each item.
- A total score is trustworthy only to the extent the scale is reliable between raters, valid against an external criterion, and sensitive to real clinical change.
- Treatment outcome is defined by score thresholds: response is conventionally a 50% reduction from baseline, and remission is a score below a scale-specific cutoff.
- Different scales for the same disorder sample different symptoms, so their scores are not interchangeable and a change of instrument can change the apparent result.
What Psychiatric Status Rating Scales Are
A rating scale decomposes a clinical construct into a fixed set of items, each scored on an ordered set of anchors, and sums the items into a total. The Hamilton Rating Scale for Depression, the instrument that made this design dominant, breaks depression into seventeen items — depressed mood, guilt, suicide, insomnia, agitation, somatic symptoms — each rated on a three- or five-point scale from a semi-structured interview (Hamilton, 1960). The total is treated as a single dimension of severity even though the items sample distinct symptom domains, and this compression of many symptoms into one number is both the source of the scale's utility and the root of its measurement problems.
Two features separate a status rating scale from the broader family of psychological tests. First, it is state-sensitive by design — built to move when the patient's condition moves, which is the opposite of a trait measure built to be stable. Montgomery and Åsberg constructed their depression scale specifically by selecting, from a larger item pool, the items most responsive to change during treatment, so that the scale would register drug effects that a less sensitive instrument would miss (Montgomery & Åsberg, 1979). Second, its output is ordinal and anchored: each score point is tied to a described behavioural or subjective threshold, so that a rating of 3 on an item is meant to denote the same severity for every rater and every patient. Whether that intention is met is an empirical question, and much of the psychometric literature exists to test it.
Types of Psychiatric Status Rating Scales
Within the Medical Subject Headings classification, Psychiatric Status Rating Scales is a node under Neuropsychological Tests, and it carries two narrower descriptors as direct subtypes. These subtypes are not mutually exclusive categories of a single taxonomy, nor a theory of how psychiatric measurement is organized; MeSH is an indexing vocabulary, so the list below reflects how the assessment literature is catalogued for retrieval, not a claim about the natural structure of the field. Both subtypes are indexed terms that do not yet have their own articles on this site.
| Subtype | In brief |
|---|---|
| Brief Psychiatric Rating Scale | A short clinician-rated scale of psychotic and general psychopathology symptoms, introduced to give psychopharmacology trials a rapid, reproducible severity measure (Overall & Gorham, 1962). |
| Mental Status Schedule | A structured interview schedule that standardizes the recording of current mental-status signs and symptoms, so that findings from the psychiatric examination can be scored and compared. |
Clinician-Administered and Self-Report Scales
The single most consequential design choice is who does the rating. In a clinician-administered scale, a trained rater conducts an interview, weighs the patient's report against observed behaviour, and assigns each score; in a self-report scale, the patient reads each item and rates it directly. The two formats measure the same construct through different apertures. A clinician can discount a patient's denial of a symptom that is plainly present, and can probe an ambiguous answer, but introduces the rater's own judgment as a source of error. A self-report scale removes rater variance and captures the patient's internal experience directly, but depends on the patient's insight, candour, and reading ability, and cannot see what the patient will not report.
The clinician-administered tradition runs from the Brief Psychiatric Rating Scale (Overall & Gorham, 1962) through the Hamilton depression and anxiety scales (Hamilton, 1960; Hamilton, 1959) to disorder-specific instruments: the Positive and Negative Syndrome Scale for schizophrenia, which scores thirty items across positive, negative, and general psychopathology subscales (Kay et al., 1987); the Young Mania Rating Scale for the manic pole of bipolar disorder (Young et al., 1978); and the Montgomery–Åsberg Depression Rating Scale (Montgomery & Åsberg, 1979). Global instruments sit alongside the symptom scales: the Global Assessment Scale reduces overall psychiatric disturbance to a single 1-to-100 rating of functioning (Endicott et al., 1976), and the Clinical Global Impressions scale asks the clinician for a single judgment of severity and of change, a deliberately coarse measure whose simplicity is its clinical appeal (Busner & Targum, 2007).
The self-report tradition begins with the Beck Depression Inventory, which asked patients to select, for each of twenty-one symptom groups, the statement that best described them (Beck et al., 1961). Its modern successor in primary care is the nine-item Patient Health Questionnaire, whose items map directly onto the diagnostic criteria for a major depressive episode and whose brevity made routine depression measurement feasible outside psychiatry (Kroenke et al., 2001). Some instruments are published in parallel clinician and self-report versions of the same items — the Quick Inventory of Depressive Symptomatology among them — precisely so that the two vantage points can be compared on a common scale (Rush et al., 2003).
Figure 1
From Symptom State to Clinical Decision
Reading the Score: Severity Bands
A raw total means little until it is placed on a severity scale. Instrument developers publish cutoffs that partition the score range into bands — none, mild, moderate, severe — validated against clinical judgment or an external criterion. The Patient Health Questionnaire, for example, is conventionally read at cutoffs of 5, 10, 15, and 20, dividing its 0-to-27 range into minimal, mild, moderate, moderately severe, and severe depression, with a score of 10 the usual threshold for probable major depression (Kroenke et al., 2001). These bands are conveniences, not natural boundaries: the underlying severity is continuous, and a patient one point either side of a cutoff differs trivially in symptom burden while being labelled differently. The first demonstration makes the arbitrariness of banding visible by letting the reader move a score across the published thresholds of several scales.
Demonstration 1 — Reading a score as a severity band
Pick a depression scale and slide a raw total through its published cutoffs. The coloured bands are the labels developers assign to ranges of the score; the marker shows where the current total falls.
The bands differ between scales — a score of 10 is moderate on the PHQ-9 but mild on the MADRS — because each range was calibrated separately. When the marker sits one point from a cutoff, a single item rated one step higher moves the patient into the next label, though their symptom burden has barely changed.
Psychometric Properties
Three properties decide whether a score can be believed. Inter-rater reliability asks whether two clinicians rating the same patient produce the same total; without it, a change between visits could reflect a change of rater rather than of patient. Validity asks whether the scale measures what it claims — whether its total correlates with an accepted external criterion and rises in the conditions where the construct should be present. Sensitivity to change, or responsiveness, asks whether the score moves when the patient's condition moves, the property Montgomery and Åsberg engineered directly into their scale (Montgomery & Åsberg, 1979). A scale can be reliable yet invalid, or valid yet insensitive, and a critical review of the Hamilton depression scale's many versions found its clinimetric properties uneven across forms, with the original structure performing better on some criteria than later revisions (Carrozzino et al., 2020).
A separate problem is what a given number means. A total of 75 on the Positive and Negative Syndrome Scale is not self-interpreting; only by linking scale scores to the global clinical impressions of experienced raters could the field establish that particular PANSS totals correspond to mildly, moderately, or markedly ill states (Leucht et al., 2005). This anchoring of an arbitrary metric to clinical meaning is what allows a trial's reported score change to be read as a clinically meaningful improvement rather than a statistical artefact.
Response and Remission
The clinical and regulatory use of these scales rests on two derived definitions. Response is a proportional improvement from baseline — conventionally a reduction of at least 50% in the total score — and marks a patient who has improved substantially. Remission is an absolute state — a score falling below a scale-specific threshold — and marks a patient who is now, by the instrument's standard, well. The two are not the same: a patient can respond without remitting, halving a very high score yet remaining symptomatic, and the distinction drives treatment decisions about whether to continue, augment, or switch. Because these definitions are computed from the total score, the choice of scale and cutoff directly determines who counts as recovered. The second demonstration lets the reader set a baseline and a follow-up score and watch the response and remission criteria resolve.
Demonstration 2 — Response and remission are different verdicts
Set a baseline Hamilton depression score and a follow-up score after treatment. Response is a reduction of at least 50% from baseline; remission is a follow-up total of 7 or below. A patient can meet one without the other.
At the default 24 → 11 the reduction is 54.2%, so the patient is a responder, but 11 is above the remission threshold of 7 — improved yet still symptomatic. Lower the follow-up to 7 and remission is also met; raise the baseline and the same follow-up score suddenly clears the 50% bar, because response is relative to where treatment began.
Routine use of rating scales to guide treatment — measuring symptoms at each visit and adjusting care when the score fails to fall — is the core of measurement-based care, an approach that improves outcomes over treatment as usual but which has been adopted slowly in practice despite decades of evidence (Aboraya et al., 2018). The obstacles are practical rather than theoretical: the time each scale takes, the training raters require, and the difficulty of integrating repeated measurement into a clinical workflow.
| Scale | Construct | Format | Rater |
|---|---|---|---|
| Brief Psychiatric Rating Scale (BPRS) | General psychopathology, psychosis | 18-24 items | Clinician |
| Hamilton Depression Rating Scale (HAM-D) | Depression severity | 17-21 items | Clinician |
| Montgomery–Åsberg (MADRS) | Depression, change-sensitive | 10 items | Clinician |
| Positive and Negative Syndrome Scale (PANSS) | Schizophrenia symptoms | 30 items | Clinician |
| Young Mania Rating Scale (YMRS) | Manic severity | 11 items | Clinician |
| Beck Depression Inventory (BDI) | Depression severity | 21 items | Self-report |
| Patient Health Questionnaire (PHQ-9) | Depression, primary care | 9 items | Self-report |
| Clinical Global Impressions (CGI) | Global severity and change | 2 items | Clinician |
The Content-Overlap Problem
Because each scale samples a different set of symptoms, two instruments for the same disorder are not measuring quite the same thing. An analysis of seven common depression scales found they collectively assessed fifty-two distinct symptoms, with little overlap in their specific content; some symptoms appeared on nearly every scale while most appeared on only one or two (Fried, 2017). The consequence is that a patient's depression score depends materially on which scale was used, and a trial result obtained with one instrument need not replicate with another. This is a measurement instance of the jangle fallacy — treating two differently named measures as different constructs, or, in reverse, assuming that identically labelled scales measure an identical thing. The third demonstration shows the overlap directly, comparing the symptom content of two selected depression scales.
Demonstration 3 — Two scales, different symptoms
Choose two depression scales and see which symptoms they share and which each measures alone. The Jaccard overlap is the fraction of all symptoms named by either scale that both scales include — a direct measure of how interchangeable their totals are.
HAM-D only
Shared
MADRS only
Even two long-established depression scales overlap only partially: much of each instrument's content is unique to it. This is the measurement face of Fried's finding that seven common depression scales together sample 52 distinct symptoms — two totals labelled "depression" are not scoring the same thing. Content is illustrative of representative scale items.
Worked Example
Consider a patient entering treatment for a major depressive episode with a baseline Hamilton depression score of 24 — in the severe range. After eight weeks of treatment the score is 11.
Response is a reduction of at least 50% from baseline. The absolute reduction is 24 − 11 = 13 points, and the proportional reduction is 13 ÷ 24 = 0.542, or 54.2%. Because 54.2% exceeds the 50% threshold, the patient is a responder.
Remission on the Hamilton scale is conventionally a total of 7 or below. The follow-up score of 11 is above 7, so the patient has not remitted despite responding. Clinically this is the common and important case: the patient is substantially better but still symptomatic, and residual symptoms at this level predict relapse, so treatment would typically be intensified rather than considered complete.
Had the same 13-point improvement started from a baseline of 20 rather than 24, the proportional reduction would be 13 ÷ 20 = 0.65, or 65% — still a response, and the follow-up score of 7 would now also meet remission. The identical absolute change yields a different clinical verdict depending on where it started, which is exactly why response (relative) and remission (absolute) are reported as separate endpoints.
Discussion
The psychiatric status rating scale solved a real problem — it gave psychiatry a reproducible way to quantify severity and to demonstrate that a treatment changes it — and the entire evidence base of psychopharmacology is built on the scores these instruments produce. But the solution carries permanent costs that no refinement removes. Compressing a multidimensional symptom state into a single ordinal total discards information about which symptoms are present; treating that total as an interval quantity, so that a drop from 24 to 11 is called a 54% improvement, imposes arithmetic on anchors that were never equal-interval; and setting a cutoff to define remission draws a categorical line through a continuous distribution. Each of these is a defensible simplification, and each can mislead when the number is read as if it were the patient.
The methodological literature has responded by tightening the standards for how scales are used rather than abandoning them. Recommendations for the conduct of clinical trials now specify how outcome measures should be chosen, administered, and reported, so that a reported score change can be interpreted and compared across studies (Guidi et al., 2018). The pressure runs in two directions at once: toward more rigorous use of the established scales in research, and toward simpler, briefer instruments that a busy clinic can actually deploy at every visit. The reconciliation of those pressures — rigorous measurement that is also feasible — is the practical frontier of the field.
Current Directions
Three lines of work are active. The first is the critical re-examination of the legacy scales themselves. Detailed clinimetric review has shown that the most-used instruments are not psychometrically uniform — the several versions of the Hamilton depression scale differ in how well they perform, and the widely assumed unidimensionality of a total score is often not supported (Carrozzino et al., 2020). This has prompted calls to report subscale or individual-symptom information rather than a single sum. The second is the content-heterogeneity problem: the demonstration that common depression scales share little symptom content (Fried, 2017) has fed a broader movement toward analysing symptoms individually — as a network of interacting complaints rather than as interchangeable indicators of one latent severity — which changes what a rating scale is taken to measure. The third is implementation: the slow translation of measurement-based care from evidence into routine practice remains an active target, with work focused on the workflow, training, and technology that would let repeated measurement fit into ordinary clinical time (Aboraya et al., 2018). Digital administration and ecological momentary methods promise to move measurement out of the clinic visit entirely, sampling symptoms in daily life rather than reconstructing them retrospectively in an interview.
Common Misconceptions
- A rating scale diagnoses the disorder.
- A status rating scale measures severity, not category. It presupposes that the clinician is assessing a given condition and quantifies how intense it is; a high score is not a diagnosis and a scale was never validated to make one. The Patient Health Questionnaire is often used as a screen, but even there a positive result indicates probable depression to be confirmed by interview, not a diagnosis conferred by the number (Kroenke et al., 2001).
- Scores from different scales for the same disorder are interchangeable.
- They are not. Seven common depression scales together sample fifty-two different symptoms with little shared content, so a patient's score depends on which instrument was used and results obtained with one scale need not hold with another (Fried, 2017). A change of scale can change the apparent finding.
- A point on the scale is the same amount of illness everywhere on the range.
- Total scores are ordinal, built from anchored item ratings that were never guaranteed to be equal-interval, so a one-point move near the floor need not equal a one-point move near the ceiling. Establishing what a given total actually corresponds to clinically requires explicit linkage to global impressions, as was done to make PANSS totals interpretable (Leucht et al., 2005).
Glossary
- Anchor.
- The described threshold attached to each score point of an item, intended to make a given rating denote the same severity for every rater.
- Ceiling effect.
- Clustering of scores at the top of a scale's range, so that the instrument cannot register further worsening in the most severely ill.
- Clinician-administered scale.
- An instrument scored by a trained rater from an interview and observation, rather than completed by the patient.
- Cutoff.
- A score value chosen to divide the range into bands or to define a state such as remission; a convenience drawn through a continuous distribution.
- Inter-rater reliability.
- The degree to which two raters scoring the same patient produce the same total; a precondition for reading change over time as change in the patient.
- Jangle fallacy.
- The error of assuming that two measures with different names assess different constructs, or that identically named scales measure an identical thing.
- Measurement-based care.
- The routine use of rating-scale scores at each visit to guide treatment decisions, adjusting care when symptoms fail to improve.
- Ordinal scale.
- A scale whose values are ordered but not guaranteed equal-interval, so that differences between scores are not strictly additive.
- Remission.
- An absolute outcome state defined by the total score falling below a scale-specific threshold, denoting a patient now well by the instrument's standard.
- Response.
- A proportional outcome, conventionally a reduction of at least 50% in total score from baseline, marking substantial improvement short of full recovery.
- Self-report scale.
- An instrument the patient completes directly, capturing internal experience without rater variance but depending on insight and candour.
- Sensitivity to change.
- Responsiveness; the capacity of a scale's score to move when the patient's clinical condition moves, essential for detecting treatment effects.
- Severity band.
- A labelled range of the score continuum — such as mild, moderate, severe — bounded by published cutoffs.
- Status rating scale.
- An instrument that quantifies the current severity of a presumed psychiatric condition, designed to be re-administered so that change can be tracked.
- Validity.
- The degree to which a scale measures the construct it claims to, judged against external criteria and expected patterns of association.
Key Researchers
Aaron T. Beck (1921-2021). American psychiatrist and founder of cognitive therapy; created the Beck Depression Inventory, one of the first widely used self-report severity measures. Wikipedia
Eiko Fried (contemporary). Clinical psychologist at Leiden University; showed that common depression scales share little symptom content and advanced the network approach to psychopathology. Faculty Page - ORCID
Max Hamilton (1912-1988). British psychiatrist; created the Hamilton Rating Scale for Depression and the Hamilton Anxiety Scale, which set the template for clinician-administered symptom rating. Wikipedia
Kurt Kroenke (contemporary). Chancellor's Professor of Medicine at Indiana University and the Regenstrief Institute; co-developed the PHQ-9, which brought routine depression measurement into primary care. Faculty Page - ORCID
Stefan Leucht (contemporary). Professor of psychiatry at the Technical University of Munich; established what PANSS scores mean clinically by linking them to global clinical impressions. Faculty Page - ORCID
Stuart A. Montgomery (contemporary). Emeritus Professor of psychiatry at Imperial College London; with Marie Åsberg he built a depression scale designed specifically to be sensitive to treatment change. Faculty Page
John E. Overall (1929-2016). Professor of psychiatry at the University of Texas; with Donald Gorham he created the Brief Psychiatric Rating Scale, the first widely adopted psychiatric severity measure. Memorial
A. John Rush (b. 1942). American psychiatrist who led the STAR*D trial; developed the Quick Inventory of Depressive Symptomatology in parallel clinician and self-report forms. Faculty Page - ORCID
Marie Åsberg (b. 1938). Professor of psychiatry at the Karolinska Institutet; with Stuart Montgomery she developed the Montgomery–Åsberg Depression Rating Scale, selecting its items for sensitivity to change. Faculty Page - ORCID
Frequently Asked Questions
What is the difference between a psychiatric rating scale and a diagnosis?
A rating scale measures the current severity of a symptom state, assuming the condition is present, whereas a diagnosis is a categorical judgment about whether a disorder exists. A scale total tracks how the condition changes over time; it is not designed to establish the diagnosis itself (Overall & Gorham, 1962).
What is the difference between clinician-rated and self-report scales?
A clinician-rated scale is scored by a trained rater from an interview and observation, which allows judgment about a patient's report but adds rater variance; a self-report scale is completed by the patient, removing rater variance but depending on the patient's insight and candour. The Beck Depression Inventory is a classic self-report instrument (Beck et al., 1961).
What do response and remission mean?
Response is conventionally a reduction of at least 50% in the total score from baseline, marking substantial improvement; remission is a total falling below a scale-specific threshold, marking a patient now well by the instrument's standard. A patient can respond without remitting (Rush et al., 2003).
Why do two depression scales give different results?
Because they sample different symptoms. An analysis of seven common depression scales found fifty-two distinct symptoms assessed across them with little overlap, so a patient's score depends materially on which scale was used (Fried, 2017).
What makes a rating scale sensitive to change?
Its items must be responsive, chosen so that their ratings move when the patient's condition moves. Montgomery and Åsberg built their depression scale by selecting, from a larger pool, the items most responsive to change during treatment (Montgomery & Åsberg, 1979).
What is measurement-based care?
It is the routine use of rating-scale scores at each visit to guide treatment, adjusting care when symptoms fail to improve. It improves outcomes over usual care but has been adopted slowly because of the time and training that repeated measurement demands (Aboraya et al., 2018).
Does a high score on a rating scale mean the total is an exact amount of illness?
No. Total scores are ordinal sums of anchored item ratings that are not guaranteed equal-interval, so a given total must be linked to clinical impressions before it can be interpreted, as was done for the Positive and Negative Syndrome Scale (Leucht et al., 2005).
Are the classic scales still considered psychometrically sound?
They remain in wide use, but critical review has found their properties uneven; the several versions of the Hamilton depression scale differ in performance and their assumed unidimensionality is often unsupported, prompting calls to report symptom-level detail rather than a single sum (Carrozzino et al., 2020).
References
Aboraya, A., Nasrallah, H. A., Elswick, D. E., Ahmed, E., Estephan, N., Aboraya, D., Berzingi, S., Chumbers, J., Berzingi, S., Justice, J., Zafar, J., & Dohar, S. (2018). Measurement-based care in psychiatry: Past, present, and future. Innovations in Clinical Neuroscience, 15(11-12), 13-26. https://pmc.ncbi.nlm.nih.gov/articles/PMC6380611/
Beck, A. T., Ward, C. H., Mendelson, M., Mock, J., & Erbaugh, J. (1961). An inventory for measuring depression. Archives of General Psychiatry, 4(6), 561-571. https://doi.org/10.1001/archpsyc.1961.01710120031004
Busner, J., & Targum, S. D. (2007). The Clinical Global Impressions Scale: Applying a research tool in clinical practice. Psychiatry (Edgmont), 4(7), 28-37. https://pmc.ncbi.nlm.nih.gov/articles/PMC2880930/
Carrozzino, D., Patierno, C., Fava, G. A., & Guidi, J. (2020). The Hamilton Rating Scales for Depression: A critical review of clinimetric properties of different versions. Psychotherapy and Psychosomatics, 89(3), 133-150. https://doi.org/10.1159/000506879
Endicott, J., Spitzer, R. L., Fleiss, J. L., & Cohen, J. (1976). The Global Assessment Scale: A procedure for measuring overall severity of psychiatric disturbance. Archives of General Psychiatry, 33(6), 766-771. https://doi.org/10.1001/archpsyc.1976.01770060086012
Fried, E. I. (2017). The 52 symptoms of major depression: Lack of content overlap among seven common depression scales. Journal of Affective Disorders, 208, 191-197. https://doi.org/10.1016/j.jad.2016.10.019
Guidi, J., Brakemeier, E. L., Bockting, C. L. H., Cosci, F., Cuijpers, P., Jarrett, R. B., Linden, M., Marks, I., Peretti, C. S., Rafanelli, C., Rief, W., Schneider, S., Schnyder, U., Sensky, T., Tomba, E., Vazquez, C., Vieta, E., Zipfel, S., Wright, J. H., & Fava, G. A. (2018). Methodological recommendations for trials of psychological interventions. Psychotherapy and Psychosomatics, 87(5), 276-284. https://doi.org/10.1159/000490574
Hamilton, M. (1959). The assessment of anxiety states by rating. British Journal of Medical Psychology, 32(1), 50-55. https://doi.org/10.1111/j.2044-8341.1959.tb00467.x
Hamilton, M. (1960). A rating scale for depression. Journal of Neurology, Neurosurgery, and Psychiatry, 23(1), 56-62. https://doi.org/10.1136/jnnp.23.1.56
Kay, S. R., Fiszbein, A., & Opler, L. A. (1987). The Positive and Negative Syndrome Scale (PANSS) for schizophrenia. Schizophrenia Bulletin, 13(2), 261-276. https://doi.org/10.1093/schbul/13.2.261
Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606-613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x
Leucht, S., Kane, J. M., Kissling, W., Hamann, J., Etschel, E., & Engel, R. R. (2005). What does the PANSS mean? Schizophrenia Research, 79(2-3), 231-238. https://doi.org/10.1016/j.schres.2005.04.008
Montgomery, S. A., & Åsberg, M. (1979). A new depression scale designed to be sensitive to change. British Journal of Psychiatry, 134(4), 382-389. https://doi.org/10.1192/bjp.134.4.382
Overall, J. E., & Gorham, D. R. (1962). The Brief Psychiatric Rating Scale. Psychological Reports, 10(3), 799-812. https://doi.org/10.2466/pr0.1962.10.3.799
Rush, A. J., Trivedi, M. H., Ibrahim, H. M., Carmody, T. J., Arnow, B., Klein, D. N., Markowitz, J. C., Ninan, P. T., Kornstein, S., Manber, R., Thase, M. E., Kocsis, J. H., & Keller, M. B. (2003). The 16-item Quick Inventory of Depressive Symptomatology (QIDS), clinician rating (QIDS-C), and self-report (QIDS-SR): A psychometric evaluation in patients with chronic major depression. Biological Psychiatry, 54(5), 573-583. https://doi.org/10.1016/S0006-3223(02)01866-8
Young, R. C., Biggs, J. T., Ziegler, V. E., & Meyer, D. A. (1978). A rating scale for mania: Reliability, validity and sensitivity. British Journal of Psychiatry, 133(5), 429-435. https://doi.org/10.1192/bjp.133.5.429