Child development

Reading a GARS-3 Score in Context

Save

An Autism Index is a comparison, not a quantity. It says where someone's ratings fall relative to the group the test was built on — which makes who was in that group part of what the number means. That is the thing a family most needs to understand before a GARS-3 result changes what they believe about their child.

Last updated: July 2026

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

Continue in Claude

Open a chat with this article’s link already in the message, and keep asking questions there. Claude reads the article and its sources; nothing about you is included.

Continue in Claude →

The button opens the Claude desktop app and fills in the message for you to review before sending. No desktop app, or reading on a phone? Copy the prompt and paste it into any AI.

What the GARS-3 Is and What It Produces

The Gilliam Autism Rating Scale, Third Edition is a norm-referenced instrument of 58 items organized into six subscales, intended for individuals aged 3 through 22, and used either to screen for autism spectrum disorder or to contribute to a fuller autism evaluation. Results come back as scaled scores on each subscale plus an overall Autism Index, and the scoring runs in the symptomatic direction: higher means greater likelihood of autism 1.

Its publisher is a commercial test company and the test manual is not a peer-reviewed document, so the citable source of record for what this instrument is and how it behaves is an independent test review published in the Journal of Psychoeducational Assessment 1. That sounds like bibliographic housekeeping. It is not. It is the reason several numbers attached to this test online have no independent backing at all.

The subscale profile carries more usable information than the single Index number. Six subscales exist because the behaviors they cover do not move in lockstep, and one person can land in very different places across them. The pattern is what an evaluator reads. The Index is what gets quoted.

What an Autism Index Actually Compares You To

Norm-referenced means a score is not a count of anything. It is a position — where this person's ratings fall relative to the sample of people the test was standardized on. Change the reference sample and an identical set of ratings produces a different number. That is true of every norm-referenced test ever built, and it is the most useful single thing to understand about an Index score.

So the makeup of the norming sample is not a footnote. It is part of the measurement. The peer-reviewed review of the GARS-3 states that its norming sample consisted only of individuals who already carried autism diagnoses, and that it had limited racial and ethnic diversity 1. A score therefore answers a narrower question than it appears to be answering.

The practical translation is this. An Index tells you how someone's rated behavior compares with a particular reference group under particular conditions. It does not tell you how autistic someone is, because that is not a quantity that exists to be measured. And no rating instrument is the gold-standard autism assessment — that role belongs to a full diagnostic evaluation by clinicians who have met the person.

The Cautions Its Peer Review Attaches

The review of record recommends caution in two specific situations: assessing individuals aged 20 to 22, and assessing people from minority populations 1. Both cautions trace back to the norming sample rather than to anything about the person being rated. An instrument generalizes only as far as the group it was built on reaches, and the review says plainly where this one thins out.

This page prints no GARS-3 cutoff, and no sensitivity or specificity figure. The independent review establishes the instrument's structure and its psychometric cautions 1; the accuracy figures that circulate for this test come from publisher materials rather than from that review, and the gap between a vendor's number and an independently reported one is invisible to a reader once the number has been copied onto a third-party page.

If a report in front of you carries an Index score, the questions worth asking are who completed the ratings, what the subscale profile looked like underneath the composite, and what the evaluator concluded after setting it beside everything else they had. Being unable to interpret a standardized score without the manual is not a gap in your understanding. It is how these instruments are designed.

Where the GARS-3 Sits Among the Other Instruments

Autism instruments differ in a way that matters more than their names: who supplies the information, and whether anyone interacts with the child. The GARS-3 gathers ratings about a person from someone who knows them 1. A free parent-report screen works differently again — the M-CHAT-R/F asks caregivers 20 questions about toddlers aged 16 to 30 months, with a structured follow-up interview when a screen comes back positive, and it is a screen rather than a diagnostic test 2.

Some instruments are done with the child directly. The STAT is an interactive, play-based screener for children aged 24 to 36 months that works through play, communication, and imitation skills. Its cutoff was derived through signal detection analysis, its sensitivity, specificity, and predictive values were confirmed in both a development sample and an independent validation sample, and its risk categories were checked against ADOS-G classification 3.

Others are rated by an observer who watched. The CARS-2 rating scale is that kind: fourteen areas of behavior plus a general-impressions rating, each scored 1 to 4 in half-point steps, giving a total between 15 and 60 in which higher means greater impairment 4.

None of these is a smaller or larger version of another. They collect different evidence, from different people, at different ages. A full evaluation is where several of them get reconciled against each other.

Why Who Filled It In Matters as Much as What It Says

A rating instrument records one person's account of another person's behavior, which means the account arrives carrying the rater's vantage point. An adult who sees a child only at home sees a different child from one who sees them only in a classroom of twenty-four. Neither is wrong. Both are partial, and one set of ratings from one rater is one vantage point.

That is why teacher input is requested so often during autism evaluations, and why evaluators put weight on cross-setting observation. Agreement between raters in different settings is informative; disagreement is informative too. A child who reads as unremarkable at school and comes apart at home every afternoon is telling you something specific about where their effort is being spent.

Diagnosis is not a matter of accumulating rating forms. It rests on developmental history and directly observed behavior, and a comprehensive evaluation may draw on developmental pediatricians, child psychologists or psychiatrists, and neurologists, because no laboratory test settles the question 5. Rating scales feed that process rather than shortcut it.

What Happens at the Top of the Age Range

The GARS-3 stops at 22, and its review recommends particular caution in the 20-to-22 band 1 — which is precisely the age at which many people first go looking for an answer about themselves. Adults never assessed as children are a large group, and instruments built and normed on children do not simply extend upward into adult life.

Adult assessment uses adult instruments. The RAADS-R is one: an 80-question scale developed to assist the diagnosis of autism spectrum disorder in adults, organized across social relatedness, circumscribed interests, language, and sensory-motor domains. Its international validation ran across 779 people — 201 with autism and 578 comparison subjects — and reported sensitivity of 97%, specificity of 100%, test-retest reliability of .987, and concurrent validity of 95.59% against the Social Responsiveness Scale-Adult 6.

Those are strong figures, and they come from a study in which the autism group had already been diagnosed against established criteria before the scale was applied 6. That is the normal shape of an instrument validation, and it describes how cleanly the scale separated groups that were already separated. It is a different claim from how any questionnaire performs on an unsorted person walking into a clinic for the first time.

What to Do With a GARS-3 Result

Treat it as one document in a file rather than as an answer. A GARS-3 result contributes to an evaluation; the diagnosis comes from a clinician integrating developmental history, direct observation, and testing, sometimes across several specialties 5. If a rating result is the only thing a family has, what they have is a reason to pursue an evaluation, not the outcome of one.

The useful requests are specific. Ask what the subscale profile showed rather than only the Index. Ask who completed the ratings and how much of the person's life that rater actually sees. Ask what the evaluator would expect to find if this particular result turned out to be misleading — clinicians who work with these instruments have a ready answer, and the answer tells you how much weight they are putting on it themselves.

When cost or waiting is what stands between a rating result and a real evaluation, two routes sometimes open things up: a university training clinic, where supervised trainees evaluate at reduced fees, and a telehealth evaluation, which widens the geography a family can choose from. Both carry limits worth asking about directly. Both are better than another year holding a questionnaire score nobody has interpreted.

Common questions

It is a rating instrument, so the items are answered about a person by someone who knows them and has observed them, rather than administered to the person directly. A report should state who provided the ratings and in what setting. If it does not, that is worth asking, because the answer changes how much weight the result deserves.

No. It means the ratings someone provided placed this child in a particular position relative to the sample the test was normed on. A diagnosis requires a clinician who has taken a developmental history, observed the child directly, and weighed testing and school information. High Index scores occur in children who are not diagnosed, and the reverse happens too.

No, and the difference is structural. The GARS-3 is a 58-item norm-referenced questionnaire producing subscale scores and an Autism Index for ages 3 to 22. The CARS-2 descends from a fifteen-item observer-rated scale with a fixed 15-to-60 total. They ask related questions of different people using different reference points.

It is a commercial clinical instrument sold to qualified professionals, and its items are not distributed to families. Beyond the licensing question, a self-administered version would not be norm-referenced against anything, because the scoring depends on comparison with a standardization sample and on training in how the items are meant to be rated.

The caution is about the test, not about the person being assessed. The instrument's norming sample carried limited racial and ethnic diversity, and a norm-referenced score means comparison with that sample. Where the reference group does not represent the person well, the comparison is less trustworthy, which is why the reviewer flagged it.

Its stated range reaches 22, but the peer-reviewed review specifically recommends caution between 20 and 22. Adults seeking assessment are usually better served by a clinician who does adult diagnostic work and uses instruments developed and validated on adults, alongside a history that reaches back into childhood.

Related

Say it back

How would you explain this to someone you love?

Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

When an evaluation should not be the next step

  • Language, gestures, or play skills that were already established and have since been lost, at any age.
  • Self-injury that is new, escalating, or causing damage, or aggression that is putting the person or people around them at risk.
  • A marked loss of everyday abilities over weeks in a teenager or adult — no longer managing tasks they were doing without help a month earlier.
  • Hopelessness, withdrawal from everything, or any statement about not wanting to be alive in a teenager or adult awaiting assessment.

Talk of suicide or self-harm means calling or texting 988, the Suicide and Crisis Lifeline, now rather than at the next appointment. Injuries needing treatment, or aggression that is not safe to contain at home, are 911 or emergency-room situations.

This article describes what a norm-referenced rating instrument measures and how its results are interpreted. It contains no items, no scoring method, and no way to assess anyone. Any conclusion about a specific person belongs to a qualified clinician who has evaluated them directly.

References

  1. 1.Karren BC (2017). A Test Review: Gilliam, J. E. (2014). Gilliam Autism Rating Scale–Third Edition (GARS-3). Journal of Psychoeducational Assessment, 35(3), 342–346. doi:10.1177/0734282916635465That the GARS-3 comprises 58 items across six subscales, is norm-referenced for ages 3 to 22, is used to screen for or contribute to an autism evaluation, and reports subscale scaled scores plus an Autism Index in which higher indicates greater likelihood of ASD; that this peer-reviewed review is the citable source of record standing in for the commercial manual; and the review's stated cautions regarding ages 20 to 22, minority populations, and a norming sample composed only of individuals with ASD diagnoses with limited racial and ethnic diversity. Deliberately not cited for cutoffs or accuracy figures.
  2. 2.Robins DL, Fein D, Barton M (2009). M-CHAT-R/F (Modified Checklist for Autism in Toddlers, Revised, with Follow-Up) — official screening instrument. mchatscreen.com (Robins, Fein & Barton, copyright holders). linkThat the M-CHAT-R/F is a free 20-item parent-report screen for toddlers aged 16 to 30 months with a structured two-stage follow-up interview for positive screens, and that it is a screening instrument rather than a diagnostic test. No items are reproduced or paraphrased here.
  3. 3.Stone WL, Coonrod EE, Turner LM, et al. (2004). Psychometric Properties of the STAT for Early Autism Screening. Journal of Autism and Developmental Disorders 2004;34(6):691-701. doi:10.1007/s10803-004-5289-8That the STAT is an interactive, play-based Level 2 autism screener for children aged 24 to 36 months assessing play, communication, and imitation; that its cutoff was derived through signal detection analysis with sensitivity, specificity, and predictive values confirmed in a development and an independent validation sample; and that its risk categories were compared against ADOS-G classification. No numeric cutoff, item count, or score range is asserted.
  4. 4.Schopler E, Reichler RJ, DeVellis RF, Daly K (1980). Toward objective classification of childhood autism: Childhood Autism Rating Scale (CARS). Journal of Autism and Developmental Disorders. doi:10.1007/BF02408436The observer-rated structure the CARS-2 descends from: fourteen behavioral areas plus a general-impressions rating, each scored 1 to 4 in half-point increments, giving a 15-to-60 total in which higher indicates greater impairment. Used here only for the contrast in instrument type; no cutoff bands are cited.
  5. 5.Centers for Disease Control and Prevention (2024). Clinical Testing and Diagnosis for Autism Spectrum Disorder. CDC — Autism Spectrum Disorder (ASD), Healthcare Providers. linkThat autism diagnosis rests on developmental history and observed behavior rather than a laboratory test, and that a comprehensive evaluation may involve developmental pediatricians, child psychologists or psychiatrists, and neurologists.
  6. 6.Ritvo RA, Ritvo ER, Guthrie D, et al. (2011). The Ritvo Autism Asperger Diagnostic Scale-Revised (RAADS-R): a scale to assist the diagnosis of Autism Spectrum Disorder in adults: an international validation study. Journal of Autism and Developmental Disorders, 41(8), 1076-89. doi:10.1007/s10803-010-1133-5That the RAADS-R is an 80-question scale assisting adult autism diagnosis across social relatedness, circumscribed interests, language, and sensory-motor domains; its international validation across 779 subjects (201 ASD, 578 comparisons) whose ASD participants already met established diagnostic criteria; and its reported sensitivity of 97%, specificity of 100%, test-retest reliability of .987, and concurrent validity of 95.59% with the Social Responsiveness Scale-Adult. No score range or cutoff is cited.

6 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — citations link their sources. Editorial policy