Abstract
A behavior rating scale is a type of psychological test: a standardized questionnaire on which an informant who knows a person well — a parent, a teacher, or the individual themselves — rates the frequency or severity of specified behaviors, which are then scored against norms. This article treats the scale as a measurement instrument: how raw item ratings become norm-referenced T-scores read against clinical cut-points, why the same child is deliberately rated by several informants, what the low agreement among those informants means, and how divergent reports are combined into a single case decision. Three interactive demonstrations model the raw-to-T-score conversion that anchors interpretation, the cross-informant discrepancy that every multi-rater profile contains, and the decision rules that turn several informants' ratings into one identification.
Keywords: norm-referenced score, informant discrepancy, T-score
The behavior rating scale is the workhorse of child and adolescent clinical assessment: cheaper than direct observation, broader than an interview, and standardized in a way a clinician's impression is not. It is a type of psychological test, sharing that class's machinery of fixed items, a representative standardization sample, and scores expressed relative to that sample. What distinguishes it is who supplies the data. The scale does not measure behavior directly; it measures an informant's report of behavior, and because a child behaves differently with a parent than with a teacher, the modern instrument is built to be completed by several informants at once. Following the scale from a raw checklist to a multi-informant profile is a compact account of how behavioral assessment learned to treat disagreement among raters as information rather than error.
- A behavior rating scale is a standardized, norm-referenced questionnaire on which an informant rates specified behaviors; the score reflects the informant's report, not a direct observation of the behavior.
- Raw item ratings are converted to T-scores (mean 50, standard deviation 10) against a normative sample, and clinical cut-points — typically a T of 65 or 70 — mark the borderline and clinical ranges.
- The same person is deliberately rated by multiple informants — parent, teacher, self — because behavior is context-specific and no single vantage point is complete.
- Agreement among informants is low, with a mean cross-informant correlation near .28, so a multi-rater profile almost always contains discrepancies that must be interpreted rather than averaged away.
- Combining informants requires an explicit decision rule — an or-rule, an and-rule, or averaging — and the rule chosen can change whether a case is identified.
What a Behavior Rating Scale Is
A behavior rating scale is a standardized questionnaire on which a knowledgeable informant judges how often, or how severely, a person shows each of a fixed set of behaviors, usually on a short ordinal scale such as not true / somewhat true / very true. The responses are summed into scales — attention problems, aggression, anxiety, and so on — and each scale total is converted into a norm-referenced score read against a representative standardization sample (Achenbach, Ivanova, & Rescorla, 2017). MeSH files the behavior rating scale as a narrower descriptor directly under psychological tests, the general class of standardized behavioral measurement of which it is one instance. Its defining move is indirection: where an intelligence test measures a person's performance directly, a behavior rating scale measures an informant's report of that person's typical behavior, trading the precision of direct observation for the breadth and low cost of a questionnaire (Hunsley & Mash, 2007).
Two features mark the instrument. The first is norm-referencing: a raw score is meaningless until it is placed against the distribution of scores in a standardization sample, so that a high score means high relative to same-age, often same-sex, peers rather than high in the abstract (Achenbach et al., 1987). The second is the reliance on multiple informants. Because a child's behavior varies across settings, the empirically based tradition of scale construction collects ratings from parents, teachers, and — for older children — the youth themselves, and treats the resulting profile, discrepancies included, as the object of interpretation (Achenbach, Ivanova, & Rescorla, 2017). As with any psychological measure, the value of a rating-scale score is judged by its reliability and validity, the standards codified for evidence-based assessment (Hunsley & Mash, 2007).
From the Conners Scales to the Multi-Informant Paradigm
The modern behavior rating scale begins with a single construct. In the late 1960s C. Keith Conners published brief parent and teacher rating scales for the symptoms of what would become attention-deficit/hyperactivity disorder, giving pediatric psychopharmacology a standardized, repeatable outcome measure where clinical impression had reigned (Conners et al., 1998). The Conners scales made hyperactivity the archetypal target of a rating scale and established the template: a short list of behaviors, an ordinal frequency judgment, and a norm-referenced total.
The paradigm widened with Thomas Achenbach's empirically based approach. Rather than starting from a diagnostic checklist, Achenbach derived syndrome scales by factor-analyzing large samples of problem-behavior ratings, and built the Child Behavior Checklist family (ASEBA) to be completed in parallel by parents, teachers, and youth, its scales and profiles set out in the school-age forms manual (Achenbach & Rescorla, 2001; Achenbach, Ivanova, & Rescorla, 2017). His 1987 meta-analysis of cross-informant agreement made low inter-rater correlation an inescapable fact of the field (Achenbach et al., 1987). Robert Goodman later showed that a brief, freely available 25-item scale, the Strengths and Difficulties Questionnaire, could match longer proprietary instruments for screening (Goodman, 1997; Goodman, 2001). By the mid-2000s Andres De Los Reyes had reframed the disagreement Achenbach documented, arguing in the Attribution Bias Context model that discrepant informant reports carry valid, context-specific signal rather than noise (De Los Reyes & Kazdin, 2005). Figure 1 sets out these milestones.
Figure 1
The behavior rating scale, 1969-2005
Scoring: From Raw Ratings to Clinical Cut-Points
The output that a clinician reads is not the sum of the item ratings but its normed transformation. A scale total is converted to a T-score, a standardized score with a mean of 50 and a standard deviation of 10 in the normative sample, computed as T = 50 + 10 × z, where z is the raw score expressed in standard-deviation units above or below the normative mean (Achenbach, Ivanova, & Rescorla, 2017). The transformation is what makes scores from different scales, and different informants, comparable on one metric: a T of 70 means the same distance above the mean whether it comes from an aggression scale or an anxiety scale.
Interpretation then rests on clinical cut-points. By convention a T-score in the mid-60s marks the borderline range and a T of 70 — two standard deviations above the mean, roughly the 98th percentile — marks the clinical range on broad-band instruments such as the CBCL, while narrower screening scales often set the flag lower (Goodman, 2001). Table 1 sets out the major behavior rating scales and what each is built to measure.
| Scale | Principal author | What it targets |
|---|---|---|
| Child Behavior Checklist (ASEBA) | Achenbach | Broad-band internalizing and externalizing syndromes, rated in parallel by parents, teachers, and youth. |
| Conners Rating Scales | Conners | Attention and hyperactivity symptoms, the archetypal narrow-band target and a standard ADHD outcome measure. |
| Strengths and Difficulties Questionnaire | Goodman | A brief 25-item screen of five subscales; a freely available scale translated into more than 80 languages. |
| Behavior Assessment System for Children | Reynolds & Kamphaus | Broad-band clinical and adaptive scales paired with validity indexes across multiple raters. |
| Behavior Rating Inventory of Executive Function | Gioia | Everyday executive function — inhibition, working memory, planning, emotional control — from informant report (Gioia et al., 2000). |
The cut-point is where the arithmetic meets the decision, so it is worth making the conversion manipulable. The first demonstration builds a T-score from a raw rating. The reader sets a raw scale total, the normative mean, and the normative standard deviation, and watches the z-score, the T-score, and the resulting classification — normal, borderline, or clinical — respond, the T-score being the quantity a clinician actually reads.
From a raw rating to a T-score and a clinical range
The raw total of 27 is 2.00 standard deviations from the normative mean, which the T-metric renders as a score of 70 — the clinical range. The raw number alone says nothing; only its place in the standardization sample makes it a clinical quantity, and the same raw total shifts range as the norms change.
Informants and the Problem of Disagreement
The central empirical fact of behavior rating scales is that informants disagree. In Achenbach's meta-analysis the mean correlation between two informants rating the same child's problems was about .28 — low by any measurement standard, and lowest of all between informants who see the child in different settings, such as a parent and a teacher (Achenbach et al., 1987). A generation of later work confirmed the pattern across instruments, ages, and problem types (De Los Reyes et al., 2015).
The interpretation of that disagreement is what changed. The older reading treated a low correlation as measurement error to be reduced or averaged out. The Attribution Bias Context model reframed it: a parent and a teacher genuinely observe different samples of behavior in different contexts and bring different attributional frames, so their divergence is partly valid variance about where and with whom a behavior occurs (De Los Reyes & Kazdin, 2005; De Los Reyes et al., 2015). In a school context the parent-versus-teacher discrepancy is not a nuisance but a datum about setting-specificity (De Los Reyes et al., 2019). The practical consequence is that a multi-rater profile is read for its pattern, not collapsed to a mean. The second demonstration makes the discrepancy manipulable: the reader sets a parent's T-score and a teacher's T-score and watches the cross-informant discrepancy — the gap between them — respond, the gap being the quantity the context-specific reading treats as signal.
Two informants, one child: the discrepancy that carries context
The two informants differ by 10 points, a gap wide enough that the older reading would have called one of them wrong. Because agreement across raters averages only about .28, some discrepancy is the rule, not the exception. The context-specific reading treats this gap as a datum about where the behavior occurs — more at home, more at school — rather than error to be smoothed into a single average.
Integrating Informants: The Decision Rule
If informants disagree and their disagreement is not simply averaged away, a clinician still has to reach a single decision: is this case identified or not? That requires an explicit combination rule, and the choice of rule is consequential (Kraemer et al., 2003; Makol et al., 2020). Three rules dominate practice. The or-rule identifies a case if any informant's score exceeds the cut-point — maximizing sensitivity, at the cost of more false positives. The and-rule requires every informant to exceed the cut-point — maximizing specificity, at the cost of missed cases. Averaging identifies a case if the mean of the informants' scores exceeds the cut-point — a middle path that can wash out a real, setting-specific problem seen by only one rater.
None of the three is correct in the abstract; each encodes a different tolerance for the two kinds of error, and the measurement models that formalize them make that trade-off explicit rather than hidden (Kraemer et al., 2003; Makol et al., 2020). The third demonstration makes the choice manipulable: the reader sets a parent's T-score, a teacher's T-score, and a cut-point, and watches each of the three rules deliver its own identification decision, so that the same two ratings can yield a case or no case depending only on the rule applied.
One case, three rules: how the decision depends on the combination
The three rules disagree on this very case: the or-rule flags the case, the and-rule clears it, and averaging the two scores to 65.0 flags it. The or-rule buys sensitivity with more false positives, the and-rule buys specificity at the cost of missed cases, and averaging can wash out a real problem seen by only one informant. No rule is correct in the abstract, so it must be chosen deliberately and stated.
Reliability, Validity, and Clinical Use
Behavior rating scales earn their place in assessment on evidence, and the standard is the same as for any psychological measure: internal consistency and test-retest reliability within an informant, and validity for the interpretations the scores support (Hunsley & Mash, 2007). Within a single informant the better scales are highly reliable, and their syndrome structure replicates across large samples (Achenbach, Ivanova, & Rescorla, 2017). The brief SDQ demonstrates that even a 25-item screen can reach acceptable psychometric properties and clinically useful cut-points (Goodman, 2001).
The validity question that behavior rating scales raise most sharply is the one their multi-informant design creates. Low cross-informant agreement means that no single informant's report is a sufficient criterion, and a scale validated against one informant may not generalize to another (Achenbach et al., 1987; De Los Reyes et al., 2015). Evidence-based assessment therefore treats the rating scale as one source among several — to be combined with history, observation, and, increasingly, dimensional and latent-variable models that place the scale on a continuous spectrum rather than a single cut (Hunsley & Mash, 2007; Kaat et al., 2019). Read this way, a rating-scale score is informative and consequential, and it is a report of behavior under standardized conditions rather than the behavior itself.
Worked Example
Follow the three demonstrations through one coherent case, checking that the arithmetic on the page matches the arithmetic in the demos.
Start with the T-score conversion. A parent completes a hyperactivity scale and the child's raw total is 27. In the normative sample this scale has a mean of 15 and a standard deviation of 6. The standardized score is z = (27 − 15) / 6 = 12 / 6 = 2.0, so the T-score is T = 50 + 10 × 2.0 = 70. Against a clinical cut-point of 65 a T of 70 falls in the clinical range. Now the teacher completes the same scale for the same child, and the raw total is 21. With the same norms, z = (21 − 15) / 6 = 6 / 6 = 1.0, giving a teacher T-score of T = 50 + 10 × 1.0 = 60, which against the same cut-point is below the clinical range.
Now the cross-informant discrepancy. The parent's T-score is 70 and the teacher's is 60, so the discrepancy is 70 − 60 = 10 points. Under the context-specific reading this ten-point gap is not error to be discarded but a datum: the hyperactive behavior is reported more strongly at home than at school, exactly the kind of setting-specificity the parent-teacher contrast is meant to surface.
Finally the combination rule, with the cut-point set at T = 65. The or-rule asks whether either informant exceeds 65: the parent's 70 does, so the case is identified. The and-rule asks whether both exceed 65: the teacher's 60 does not, so the case is not identified. Averaging takes the mean, (70 + 60) / 2 = 65, which meets the cut-point of 65, so the case is identified. The same two ratings yield identified, not identified, and identified under the three rules — a single child whose classification depends entirely on how the informants are combined, which is precisely why the rule must be chosen deliberately and stated.
Discussion
The behavior rating scale is to child clinical assessment what the intelligence test is to ability measurement: the instrument the field standardized around. Conners's decision to put hyperactivity on a norm-referenced scale, and Achenbach's decision to derive syndromes empirically and to collect them from several informants at once, gave the field a cheap, broad, repeatable measure that a clinical impression could not match. The T-score metric and its cut-points made scores from different scales and different raters comparable on one axis.
The tensions that remain are the ones the multi-informant design creates. Agreement among informants is low and will stay low, because it reflects a real fact about behavior rather than a flaw in the scales (Achenbach et al., 1987). That fact forces two choices the scale cannot make for the clinician: how to read a discrepancy — as error or as context-specific signal (De Los Reyes & Kazdin, 2005) — and how to combine informants into a decision, where the or-rule, the and-rule, and averaging encode different tolerances for false positives and false negatives (Kraemer et al., 2003; Makol et al., 2020). The measured score is informative and consequential, and it is a report of behavior under standardized conditions rather than the whole of a person's conduct, a distinction the careful clinician keeps in view.
Current Directions
Two active lines of work bear directly on how a rating-scale profile should be read. The first is the formalization of informant integration. Rather than defaulting to an or-rule or an average, researchers are building explicit conceptual and measurement models — latent-variable and dimensional approaches — that treat each informant's report as an indicator of a partly shared, partly context-specific latent construct, so that the combination rule is derived from a model rather than chosen by convention (Makol et al., 2020). Linking studies that place a broad-band scale such as the CBCL onto a continuous, item-response-calibrated dimension of disruptive behavior are part of the same move away from a single categorical cut (Kaat et al., 2019).
The second is the extension of the discrepancy framework into the settings where rating scales are most used. The parent-versus-teacher divergence that Achenbach documented is now studied as a substantive variable in its own right in school-based services, where it carries information about where a problem manifests and which intervention setting to target (De Los Reyes et al., 2019). Across both lines the reframing is the same: informant disagreement, once treated as the measurement problem to be minimized, has become a source of clinically useful information about context, and the research frontier is the methodology for extracting it (De Los Reyes et al., 2015).
Key Researchers
Thomas M. Achenbach (1940-2023). Psychologist at the University of Vermont; built the empirically based, multi-informant assessment paradigm and the Child Behavior Checklist family (ASEBA), and his 1987 meta-analysis established the low cross-informant agreement every rating-scale interpretation must contend with. ORCID - Wikipedia
C. Keith Conners (1933-2017). Psychologist at Duke University; created the Conners Rating Scales, the parent- and teacher-report instruments that made attention and hyperactivity the archetypal target of a standardized behavior rating scale and anchored decades of ADHD assessment research. Wikipedia
Gerard A. Gioia. Pediatric neuropsychologist at Children's National Hospital and George Washington University; lead author of the Behavior Rating Inventory of Executive Function (BRIEF), the instrument that made everyday executive function measurable from parent and teacher report rather than only from performance tests. Faculty Page
Robert Goodman (1953-2025). Child psychiatrist at the Institute of Psychiatry, King's College London; designed the Strengths and Difficulties Questionnaire, a brief 25-item scale now translated into more than 80 languages, and showed that a short, free instrument could match longer proprietary scales for screening. Wikipedia
Randy W. Kamphaus. Psychologist at the University of Oregon; co-author of the Behavior Assessment System for Children (BASC) and a researcher on the classification and dimensional structure of child behavior problems, advancing the use of rating scales for population screening as well as individual assessment. Wikipedia
Andres De Los Reyes. Professor of psychology at the University of Maryland, College Park; the leading contemporary theorist of informant discrepancy, whose Attribution Bias Context model reframed disagreement among raters as valid, context-specific signal rather than error to be reconciled. ORCID - Wikipedia
Cecil R. Reynolds. Emeritus professor at Texas A&M University; co-author of the Behavior Assessment System for Children (BASC), the broad-band multi-rater system that pairs clinical and adaptive scales with validity indexes, and an authority on test bias and psychometric fairness in behavioral assessment. Wikipedia
Glossary
- And-rule.
- A combination rule that identifies a case only when every informant's score exceeds the clinical cut-point; it maximizes specificity at the cost of missing setting-specific problems seen by a single rater.
- Attribution Bias Context (ABC) model.
- De Los Reyes and Kazdin's framework holding that informant discrepancies carry valid, context-specific information — different observers see behavior in different settings and interpret it through different frames — rather than being mere measurement error.
- Behavior rating scale.
- A standardized, norm-referenced questionnaire on which an informant judges the frequency or severity of specified behaviors; a type of psychological test that measures a report of behavior rather than a direct observation.
- Broad-band scale.
- A rating scale covering a wide span of problem domains — such as the internalizing and externalizing syndromes of the CBCL — as opposed to a narrow-band scale focused on a single construct.
- Clinical cut-point.
- The T-score threshold above which a scale flags a problem as clinically significant; commonly a T of 65 for the borderline range and 70 for the clinical range on broad-band instruments.
- Cross-informant agreement.
- The degree to which two informants rating the same person converge; Achenbach's meta-analysis put the mean correlation near .28, lowest between informants who observe the person in different settings.
- Externalizing behavior.
- The broad syndrome grouping outwardly directed conduct — aggression, rule-breaking, hyperactivity — that empirically based scales separate from internalizing problems.
- Informant discrepancy.
- The difference between two informants' ratings of the same person; under the ABC model a partly valid signal about the context-specificity of behavior rather than error to be averaged away.
- Internalizing behavior.
- The broad syndrome grouping inwardly directed distress — anxiety, depression, withdrawal, somatic complaints — that empirically based scales separate from externalizing problems.
- Multi-informant assessment.
- The deliberate collection of ratings from several informants — parent, teacher, self — for the same person, adopted because behavior is context-specific and no single vantage point is complete.
- Narrow-band scale.
- A rating scale focused on a single construct or disorder, such as the Conners scales for attention and hyperactivity, as opposed to a broad-band instrument spanning many domains.
- Norm-referenced score.
- A score interpreted by comparison with the distribution in a representative standardization sample rather than against an absolute standard; the form of scoring intrinsic to a behavior rating scale.
- Or-rule.
- A combination rule that identifies a case when any single informant's score exceeds the cut-point; it maximizes sensitivity at the cost of more false positives.
- Rater (informant).
- The person who completes a behavior rating scale about a target individual — a parent, teacher, or the individual as a self-reporter — whose vantage point and interpretive frame shape the resulting score.
- Standardization sample.
- The representative reference group against which raw scores are normed, so that a T-score expresses standing relative to same-age, often same-sex, peers.
- T-score.
- A standardized score with a mean of 50 and a standard deviation of 10, computed as T = 50 + 10z; the common metric that makes different rating scales and different informants comparable.
Frequently Asked Questions
What is a behavior rating scale?
It is a standardized questionnaire on which a knowledgeable informant, such as a parent, teacher, or the person themselves, rates the frequency or severity of specified behaviors, which are then scored against norms. It is a type of psychological test that measures a report of behavior rather than a direct observation (Achenbach, Ivanova, & Rescorla, 2017; Hunsley & Mash, 2007).
How is a raw rating turned into a score a clinician reads?
The scale total is converted to a T-score, a standardized score with a mean of 50 and a standard deviation of 10, using T = 50 + 10z. Clinical cut-points, often a T of 65 for the borderline range and 70 for the clinical range, then mark where a scale flags a problem (Achenbach, Ivanova, & Rescorla, 2017; Goodman, 2001).
Why are several informants asked to rate the same person?
Because behavior is context-specific: a child behaves differently with a parent than with a teacher, so no single vantage point is complete. The empirically based tradition collects parent, teacher, and youth reports in parallel and interprets the whole profile (Achenbach, Ivanova, & Rescorla, 2017).
Why do informants disagree so much?
Achenbach's meta-analysis found a mean cross-informant correlation near .28, lowest between informants who see the person in different settings. The disagreement reflects a real fact about where behavior occurs, not merely a flaw in the scales (Achenbach et al., 1987).
Is informant disagreement just measurement error?
Not entirely. The Attribution Bias Context model holds that a parent and a teacher genuinely observe different samples of behavior in different contexts, so their divergence carries valid, context-specific information rather than only noise (De Los Reyes & Kazdin, 2005; De Los Reyes et al., 2015).
How are several informants' scores combined into one decision?
By an explicit rule. An or-rule identifies a case if any informant exceeds the cut-point; an and-rule requires all of them to; averaging uses the mean. Each encodes a different tolerance for false positives versus false negatives, and the rule can change the outcome (Kraemer et al., 2003; Makol et al., 2020).
What is the difference between a broad-band and a narrow-band scale?
A broad-band scale such as the CBCL or BASC spans many problem domains at once, separating internalizing from externalizing syndromes, whereas a narrow-band scale such as the Conners scales focuses on a single construct like attention and hyperactivity (Achenbach, Ivanova, & Rescorla, 2017; Conners et al., 1998).
Can a short, free rating scale be as good as a long proprietary one?
For screening, yes. Goodman's 25-item Strengths and Difficulties Questionnaire reaches acceptable psychometric properties and clinically useful cut-points, and is freely available and widely translated, showing that brevity need not sacrifice screening accuracy (Goodman, 1997; Goodman, 2001).
References
Achenbach, T. M., McConaughy, S. H., & Howell, C. T. (1987). Child/adolescent behavioral and emotional problems: Implications of cross-informant correlations for situational specificity. Psychological Bulletin, 101(2), 213-232. https://doi.org/10.1037/0033-2909.101.2.213
Achenbach, T. M., & Rescorla, L. A. (2001). Manual for the ASEBA school-age forms & profiles. University of Vermont, Research Center for Children, Youth, & Families.
Achenbach, T. M., Ivanova, M. Y., & Rescorla, L. A. (2017). Empirically based assessment and taxonomy of psychopathology for ages 1.5-90+ years: Developmental, multi-informant, and multicultural findings. Comprehensive Psychiatry, 79, 4-18. https://doi.org/10.1016/j.comppsych.2017.03.006
Conners, C. K., Sitarenios, G., Parker, J. D. A., & Epstein, J. N. (1998). The revised Conners' Parent Rating Scale (CPRS-R): Factor structure, reliability, and criterion validity. Journal of Abnormal Child Psychology, 26(4), 257-268. https://doi.org/10.1023/A:1022602400621
De Los Reyes, A., & Kazdin, A. E. (2005). Informant discrepancies in the assessment of childhood psychopathology: A critical review, theoretical framework, and recommendations for further study. Psychological Bulletin, 131(4), 483-509. https://doi.org/10.1037/0033-2909.131.4.483
De Los Reyes, A., Augenstein, T. M., Wang, M., Thomas, S. A., Drabick, D. A. G., Burgers, D. E., & Rabinowitz, J. (2015). The validity of the multi-informant approach to assessing child and adolescent mental health. Psychological Bulletin, 141(4), 858-900. https://doi.org/10.1037/a0038498
De Los Reyes, A., Cook, C. R., Gresham, F. M., Makol, B. A., & Wang, M. (2019). Informant discrepancies in assessments of psychosocial functioning in school-based services and research: Review and directions for future research. Journal of School Psychology, 74, 74-89. https://doi.org/10.1016/j.jsp.2019.05.005
Gioia, G. A., Isquith, P. K., Guy, S. C., & Kenworthy, L. (2000). Behavior Rating Inventory of Executive Function. Child Neuropsychology, 6(3), 235-238. https://doi.org/10.1076/chin.6.3.235.3152
Goodman, R. (1997). The Strengths and Difficulties Questionnaire: A research note. Journal of Child Psychology and Psychiatry, 38(5), 581-586. https://doi.org/10.1111/j.1469-7610.1997.tb01545.x
Goodman, R. (2001). Psychometric properties of the Strengths and Difficulties Questionnaire. Journal of the American Academy of Child & Adolescent Psychiatry, 40(11), 1337-1345. https://doi.org/10.1097/00004583-200111000-00015
Hunsley, J., & Mash, E. J. (2007). Evidence-based assessment. Annual Review of Clinical Psychology, 3, 29-51. https://doi.org/10.1146/annurev.clinpsy.3.022806.091419
Kaat, A. J., Blackwell, C. K., Estabrook, R., Burns, J. L., Petitclerc, A., Briggs-Gowan, M. J., Gershon, R. C., Cella, D., Perlman, S. B., & Wakschlag, L. S. (2019). Linking the Child Behavior Checklist (CBCL) with the Multidimensional Assessment Profile of Disruptive Behavior (MAP-DB): Advancing a dimensional spectrum approach to disruptive behavior. Journal of Child and Family Studies, 28(2), 343-353. https://doi.org/10.1007/s10826-018-1272-4
Kraemer, H. C., Measelle, J. R., Ablow, J. C., Essex, M. J., Boyce, W. T., & Kupfer, D. J. (2003). A new approach to integrating data from multiple informants in psychiatric assessment and research: Mixing and matching contexts and perspectives. American Journal of Psychiatry, 160(9), 1566-1577. https://doi.org/10.1176/appi.ajp.160.9.1566
Makol, B. A., Youngstrom, E. A., Racz, S. J., Qasmieh, N., Glenn, L. E., & De Los Reyes, A. (2020). Integrating multiple informants' reports: How conceptual and measurement models may address long-standing problems in clinical decision-making. Clinical Psychological Science, 8(6), 953-970. https://doi.org/10.1177/2167702620924439