Let me tell you about the most expensive coin flip in healthcare.
You're preparing for your kid's autism assessment like it's a Supreme Court case. You've researched, organized binders, prepped them. Because you know that in your world — competitive school districts, anxious pediatricians, waitlists longer than your kid's attention span — the right score from the right test is the key that unlocks services, accommodations, and answers.
The entire system has told you the ADOS-2 is the "gold standard." An objective, unimpeachable truth-telling machine.
That "gold standard" flags kids as autistic who turn out not to be, about a third of the time. Its false-positive rate, in one referred clinical sample, was 34%.
The Science: A Takedown in Three Acts
A 2025 review in CNS Spectrums (Gupta and colleagues) pulls together the data that exposes the cracks in the ADOS-2's "gold standard" reputation — drawing on empirical studies so damning, I genuinely wonder how this test is still being used uncritically. The hardest number in it comes from Greene and colleagues (2022, The Clinical Neuropsychologist). Let me walk you through it.
Act I: The 34% False Positive Rate
When Greene and colleagues gave the ADOS-2 (Module 3) to 214 children and adolescents (ages 5–16) referred for autism evaluation, 101 scored in the test's "clinically elevated" range — yet the study found a 34% false-positive rate (against just a 1% false-negative rate). The kids wrongly flagged as autistic were disproportionately those carrying anxiety and trauma histories, not autism. As discussed in Gupta's 2025 review, that one finding alone should give every clinician pause.
Let me put that in perspective. Among the kids who scored in the autism range, nearly half — 45 of 101 — turned out not to be autistic, disproportionately kids carrying anxiety or trauma whose distress the test misread. (Greene reports the test's overall false-positive rate as 34%.) That's not science. That's not even a good guess. In a test that costs families thousands of dollars and shapes the entire trajectory of their child's life, that's closer to a coin flip.
In any other high-stakes field — aviation, medicine, engineering — a 34% failure rate would be grounds for immediate grounding. In psychological testing, it's the "gold standard."
(That should make your blood boil. It makes mine boil, and I do this for a living.)
A separate study backs this up: in a specialist adult autism service (BMC Psychiatry, 2021, n=88), the ADOS-2 (Module 4) showed about 92% sensitivity but only 57% specificity. Translation: it's great at catching autism when it's there, but it's terrible at ruling it out when it isn't. A "negative" result was actually more informative than a "positive" one.
Act II: The Systemic Bias Built Into the Test
The study reinforces what people in the neurodivergent community have been saying for years: the ADOS-2 was built for a specific, narrow prototype — young, white, cisgender males.
A 2022 study in JAMA Network Open (Kalb et al.) confirmed it: 11% of ADOS-2 diagnostic items show significant differential item functioning by race or sex. Meaning: for that slice of items, the test may be picking up race or sex differences rather than autism itself. The study's authors read the overall bias as modest — but reasonable people read 11% as a lot when a single point can change a kid's life.
It also struggles to catch the nuanced, internally-focused presentation of autism in girls and women: a 2023 analysis in the Journal of Autism and Developmental Disorders (Rea and colleagues) found females showed fewer atypicalities on most social-communication items and lower overall scores, and concluded that "gold-standard assessments may be less sensitive to female presentations of ASD." The standard "gold standard" diagnostic tools for ASD — including the ADOS-2 — are fundamentally biased against women and girls →. If your daughter has learned to mask, make eye contact, and mirror social cues (as most autistic girls do by elementary school), the ADOS-2 will likely miss her completely.
And the discrepancy isn't limited to sex. When the ADOS and ADI-R (another "gold standard" tool) are both administered to the same child, they frequently disagree — de Bildt and colleagues (2004) found agreement between the two was only fair, and weakest in older children. In many of those disagreements, the ADOS said "no autism" while the ADI-R said "autism" — and the clinician sided with the ADOS.
Ultimately, it is up to the clinician to make the call on the final diagnosis — not the test.
That principle runs all through the research literature, and it should be tattooed on every assessment report. The test doesn't diagnose. The clinician diagnoses. The test is just one data point.
Act III: The Abdication of Clinical Judgment
Perhaps the most damning point is how the test is used in practice. The Gupta review describes a disturbing trend: clinicians have been pressured — by schools, insurance companies, and a liability-averse system — to defer entirely to test scores. They've become technicians who administer the ADOS, tally the score, and deliver the verdict, even when their own clinical judgment screams that the test is wrong.
I'll say it plainly: when a clinician trusts a test score over their own brain, they've stopped being a clinician. They've become a scoring machine with a PhD.
In my practice, I've sat across from kids where every observation, every interaction, every piece of the puzzle says this is autism — and a different clinician dismissed it because the ADOS score came in one point below the cutoff. That one point is not science. It's an arbitrary line on a flawed instrument. And that kid doesn't get services, doesn't get understanding, doesn't get the language to describe their own experience — because of one point.
What to Demand Instead
So what does a good assessment actually look like? Here's my criteria:
- Multiple sources of data. The ADOS-2, if used at all, should be a single data point among dozens. A comprehensive, collaborative assessment → integrates clinical interview, behavioral observation across settings, developmental history, school data, parent/caregiver interview, and — critically — the person's own description of their inner experience.
- A clinician who uses their brain. You want someone who looks at the whole, complex, brilliant human in front of them — not someone who tallies a score and reads from a decision tree.
- Awareness of masking and bias. Your clinician should be asking: Is this presentation being missed because of who this person is? If they're not asking that question, they're administering the bias, not detecting it.
- A diagnosis that makes sense to you. If the result doesn't match your lived experience, that's a data point too. You are the expert on your own brain. The clinician is a translator, not a judge.
Treating assessment as a therapeutic intervention — something that helps you, not just labels you — is a genuine paradigm shift. The research backs it: a meta-analysis of 17 studies (Poston & Hanson, 2010) found that collaborative, feedback-rich assessment produces real, measurable benefit for the people sitting in the chair.
This is the paradigm shift I built my assessment practice → around. An assessment isn't something that happens to you. It's something we build together. Your experience, your history, your observations all carry weight — not just the ADOS score.
Your Kid Is Not a Score
Your child is not a number. Your child is a story — a complex, contradictory, brilliant story that cannot be reduced to a single test administered in a single hour by a single person.
The ADOS-2 became the "gold standard" not because it was the best tool, but because it was the most convenient. It gave the system a score. It gave insurance companies a checkbox. It gave school districts a number to put in an IEP.
But your kid deserves more than a checkbox. They deserve someone who knows how to listen.
Stop preparing your kid for a flawed test. Start preparing to challenge the system that relies on it.
When you're ready to build your team: Start here →
For the Clinicians.
A note for the clinicians, because the numbers say most of you are.
The searches that bring people here are scoring interpretation and module charts. Those come from somebody who administers the thing, or is being taught to.
I'm an LPC in St. Louis. Getting a practice found online turned out to be a whole second job I had to learn on my own, and everything I worked out sits on a second site, collab.enlitens.com. The guides are free.
It's about running a practice rather than assessing anyone: how people actually look for a therapist, what your own site needs to say, which agency pitches are selling weather. A strategy session costs $250 and the build afterwards is free. You can read the lot and stop there.
Disagreeing with the standard tool makes you harder to find. That is what it is for.
Marketing help for therapists →
Part of: The Science Library → | Related: The Myth of the Normal Brain · When Your Diagnosis Gets Invalidated