Muscle, joint & pain

Reading Your Shoulder SPADI Score

Save

Most people are handed the number without the manual. Three questions unlock it: which direction does this scale run, what is it being compared against, and how big does a change have to be before it means anything? The SPADI answers the first plainly, the second surprisingly, and the third not at all.

Last updated: July 2026

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

What the number on the form actually is

A reported SPADI score is a proportion of the worst possible answer, rescaled to run from 0 to 100. Because the questionnaire has two unequal halves — a pain section and a longer disability section — each half is converted to that same 0-to-100 footing before the two are averaged into the total, which is why a form of thirteen questions produces a percentage rather than a count 1.

That conversion is the source of most confusion about the score. Nothing was counted out of a hundred. A 62 does not mean sixty-two of anything; it means the answers landed 62 per cent of the way toward the worst response the form allows.

Higher is worse. A SPADI of 20 describes a better shoulder than a SPADI of 70. The instrument runs in the direction of disease rather than health, so improvement shows up as the number falling.

The three numbers usually reported are the pain subscale, the disability subscale and the total. They are worth reading in that order rather than skipping to the total, because a shoulder can improve on one and stall on the other, and the average conceals which.

There are no official mild, moderate and severe bands

This is the part people most want and the part the evidence does not supply. The paper that developed the SPADI defines no severity categories, no cutoff separating a manageable shoulder from a bad one, and neither a minimal clinically important difference nor a minimal detectable change 1. If a chart or an app shows your score inside a coloured band labelled moderate, that band came from somewhere other than the instrument's own foundation.

This absence is normal rather than a scandal, and it recurs across the shoulder questionnaires. The Oxford Shoulder Score, developed and validated in 1996 on patients before and six months after shoulder surgery, likewise reports no minimal clinically important difference and no minimal detectable change 2.

A score with no severity label attached is not a score that has failed you. Most function questionnaires are built to track change in one person, not to grade people against a threshold.

The Oxford score also supplies the cautionary example about direction. As originally published, it ran from 12 to 60 — with 12, the bottom of the range, representing the best possible outcome 2. Anyone who assumes shoulder questionnaires always run one way will misread half the forms they encounter, which is the practical reason a number should never travel without the name of the form that produced it.

Direction is not a convention you can assume

Related instruments run in opposite directions on purpose, and the choice usually reflects what the developers wanted the score to feel like. The HOOS, JR, a six-item hip questionnaire built from a longer parent measure, converts its raw answers through a crosswalk table onto a 0-to-100 interval scale in which 0 represents total hip disability and 100 represents perfect hip health 3. That is the exact inverse of the SPADI's arrangement.

So a 90 is close to the best result a person can have on the hip measure and close to the worst on the shoulder one. There is no way to infer which from the number.

Whatever form you were handed — the womac score, the dash score, the lefs score, the constant-murley score — the same three questions apply before the result means anything: which way does this scale run, what is it being compared against, and how much movement counts as movement.

The first is usually printed on the form or in its scoring instructions, and it is a fair thing to ask a clinician outright. Patient-reported outcome measure is the general name for this family of forms: instruments whose score comes from the patient's answers rather than from a test. Their conventions are not standardised, which is a defect of the field and not something you are failing to understand.

Compared against what?

This is the question that separates two genuinely different kinds of number, and it changes what a score can be used for. The SPADI is referenced against itself: the useful comparison is your own answers from six weeks ago, because nothing in the instrument tells you where the general population sits.

Other measures are referenced against a population instead. The PROMIS Physical Function measure reports scores on a standardised T-score metric normed so that the average of a United States general-population sample is 50, with a standard deviation of 10, and higher scores indicating better physical function 4. On that scale a 40 is immediately interpretable — one standard deviation below the general population average — without knowing anything about that person's previous score.

PROMIS Physical Function was calibrated in a sample of 16,065 adults, which is what buys it a population reference point 4. The SPADI's derivation sample was three orders of magnitude smaller, which is one reason it was never positioned as a norm-referenced instrument.

Neither design is better. A norm-referenced score answers the question where do I stand; a change-referenced score answers am I getting better, which is usually the question in a course of shoulder rehabilitation. Trouble starts only when a change-referenced score is read as though it carried a population meaning it does not have.

When the scale runs out at the top

A questionnaire can only measure inside its own range, and a person whose shoulder has recovered to near-normal will pile up at the good end where the form has nothing left to detect. Statisticians call it a ceiling effect, and it is often substantial in outcome measures. The HOOS, JR reported ceiling effects between 37 and 46 per cent in its validation cohorts, against floor effects between 0.6 and 1.9 per cent 3.

Read that plainly: in those hip-replacement cohorts, roughly four in ten people ended up at the top of the scale. Ceiling effects of 37 to 46 per cent, against floor effects of 0.6 to 1.9 per cent, in the HOOS, JR validation cohorts 3.

What that means for anyone tracking their own recovery is that the form loses resolution exactly where the good news is. Once someone is near the best possible score, further genuine improvement — being able to do a demanding overhead job again, sleeping through the night reliably — may not register at all, because there is no room left in the questions. The same happens at the bottom of the range, more rarely. In the middle of the scale is where the instrument does its actual work.

How much change counts as real change?

The honest answer is that the SPADI's own development paper does not say, and its measurement stability sets a floor under how confidently any small gap can be read. Repeat administrations in the development sample agreed at 0.64 to 0.66, while the internal consistency of the items ran 0.86 to 0.95 1. The first of those figures is the relevant one here: some of the difference between two administrations is the instrument, not the shoulder.

So a three-point drop between two visits is not evidence of anything on its own. A sustained downward direction across four measurements, with both subscales moving, is a different kind of evidence — not because any single gap became meaningful, but because noise does not usually trend.

Other instruments in this family have had thresholds worked out in later research, and the useful move is to ask which threshold your clinician is using and where it came from. The oswestry score is the familiar comparison from the low back: the Oswestry Disability Index is a validated ten-section measure reported as a percentage, and it comes with an interpretation literature the SPADI's founding paper does not contain 5.

Ask what change your clinic treats as meaningful on this form, and what study that number came from. It is a reasonable question and it has an answer.

Your score moves for reasons that are not treatment

Some shoulder conditions have a natural course that will move a questionnaire score whether or not anything is done, and reading improvement as proof that a treatment worked is the most common interpretive error in this whole area. Frozen shoulder is the clearest case: it progresses through freezing, frozen and thawing stages and usually resolves over one to three years, with range-of-motion-focused physical therapy as the primary treatment 6.

Frozen shoulder typically runs its course over one to three years, through freezing, frozen and thawing stages 6. A SPADI score falling steadily during the thawing stage is describing a condition doing what that condition does.

That is not an argument against treatment, and it is not a reason to discount your own improvement. It is a reason to be careful about causal stories built from a single trajectory, particularly when someone is selling the treatment that happened to coincide with the thaw.

Shorter-run wobbles have humbler explanations. A form filled in after a bad night, or during a week of heavy lifting, is a snapshot of that state. Making sense of shoulder pain over months means reading a line rather than a point, which is what the instrument was built for.

Common questions

There is no published band that answers that. Fifty means the answers landed halfway toward the worst response the form allows, which describes a shoulder causing substantial pain and difficulty — but the instrument defines no mild, moderate or severe categories. What a 50 means for you depends mostly on what your previous scores were and what you need the shoulder to do.

Down. The SPADI runs in the direction of the problem, so improvement shows as a falling number and zero is the best possible result. This is worth confirming on any form, because closely related questionnaires run the opposite way — on some hip and shoulder measures, one hundred is the ideal score.

The pain subscale, the disability subscale and the total. Pain and function often improve at different rates, so the two subscales carry information the average hides. A shoulder whose pain has settled while daily tasks remain difficult points toward stiffness and confidence rather than an angry tendon, and that shows only in the split.

The instrument's development paper sets no threshold, and its repeat-testing figures mean single small gaps carry measurement noise. Clinicians who quote a specific number of points are drawing on later research, which is worth asking about. In practice a consistent direction across several visits is more convincing evidence than any single difference between two.

Not usefully. There is no population reference for the SPADI, so the score carries no information about where you sit relative to other people your age. Some measures are built for exactly that comparison and are scored against population norms; this one is built to track one shoulder over time.

Two common reasons. A shoulder near the top of the scale can improve in ways the questions no longer capture, because the form runs out of room. And the questionnaire asks about a defined set of everyday tasks, so gains in something it does not ask about — sleep, confidence, sport — may not appear in the number at all.

Related

Say it back

How would you explain this to someone you love?

Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

Shoulder symptoms that no score should be used to reason about

  • Sudden inability to raise the arm after an injury, especially with visible deformity or the arm held fixed against the body
  • Fever, chills, or a red, hot and swollen shoulder, particularly after a joint injection or a recent infection elsewhere
  • Shoulder or arm ache that comes on with exertion and eases with rest, or that arrives with chest tightness, breathlessness or sweating
  • New numbness, pins and needles, or weakness spreading down the arm and into the hand

Shoulder or arm pain occurring with chest tightness, breathlessness, sweating or nausea can be a heart attack. Call 911 immediately rather than waiting to see whether it settles.

This page explains how a shoulder questionnaire score is constructed and what can and cannot be read from it. It is general education, not medical advice, and no questionnaire result can identify the cause of shoulder pain or substitute for assessment by a clinician who can examine you.

References

  1. 1.Roach KE, Budiman-Mak E, Songsiridej N, Lertratanakul Y (1991). Development of a shoulder pain and disability index. Arthritis Care Res. 1991;4(4):143-9. doi:10.1002/art.1790040403That the SPADI is a self-administered index of thirteen items split into a pain subscale and a longer disability subscale; the original psychometrics of internal consistency 0.8604-0.9507 and test-retest reliability 0.6377-0.6552 in the derivation cohort; and the absence from this defining paper of any minimal clinically important difference, minimal detectable change, or severity cutoff.
  2. 2.Dawson J, Fitzpatrick R, Carr A (1996). Questionnaire on the perceptions of patients about shoulder surgery. J Bone Joint Surg Br 1996;78-B(4):593-600. doi:10.1302/0301-620X.78B4.0780593That the Oxford Shoulder Score was developed and validated prospectively in patients assessed before shoulder surgery and again at six months; that it reports no minimal clinically important difference and no minimal detectable change; and that its original 1996 scoring ran from 12 to 60 with 12 representing the best outcome — used here as the illustration that scoring direction cannot be assumed.
  3. 3.Lyman S, Lee YY, Franklin PD, et al. (2016). Validation of the HOOS, JR: A Short-form Hip Replacement Survey. Clinical Orthopaedics and Related Research, 474(6):1472-1482. doi:10.1007/s11999-016-4718-2That the HOOS, JR is a six-item short form whose raw sum is converted by a crosswalk table to a 0-100 interval measure in which 0 represents total hip disability and 100 perfect hip health (higher = better), and that its validation cohorts showed ceiling effects of 37% to 46% and floor effects of 0.6% to 1.9%.
  4. 4.Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014). The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency. Journal of Clinical Epidemiology, 67(5):516-526. doi:10.1016/j.jclinepi.2013.10.024That PROMIS Physical Function scores are reported on a standardised T-score metric normed to a mean of 50 and standard deviation of 10 in a US general population sample, with higher scores indicating better physical function, and that the item bank was calibrated in 16,065 adults — the contrasting case of a norm-referenced rather than change-referenced score.
  5. 5.Fairbank JCT, Pynsent PB (2000). The Oswestry Disability Index. Spine. doi:10.1097/00007632-200011150-00017That the Oswestry Disability Index is a validated ten-section patient-reported measure of low-back-related disability scored as a percentage, cited here as the familiar comparison case of a regional disability index with an established score-interpretation literature.
  6. 6.American Academy of Orthopaedic Surgeons (OrthoInfo) (2024). Frozen Shoulder (Adhesive Capsulitis). OrthoInfo — AAOS. linkThat frozen shoulder progresses through freezing, frozen and thawing stages and usually resolves over one to three years, with range-of-motion-focused physical therapy as the primary treatment — the natural-history caution against reading a falling questionnaire score as proof that a treatment caused the improvement.

6 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — every citation independently verified. Editorial policy