How Much Score Change Actually Counts as Better
SaveEvery outcome questionnaire produces a number, and every number invites the question of how far it has to move before the change is real. Researchers answer it with a threshold called the minimal clinically important difference. It is the reason a study can be statistically positive and clinically empty at the same time, and the reason your own score sheet gets easier to read once you know roughly where that line sits.
Last updated: July 2026
What does minimal clinically important difference mean?
The minimal clinically important difference is the smallest change in a score that the person filling in the form would recognise as a genuine improvement in their life. It sits above the noise of the questionnaire and below the dramatic. A change smaller than it can be perfectly real in the arithmetic and still be invisible from inside the body the form describes.
MCID — the smallest score change a patient would recognise as a real improvement, rather than the smallest change a statistician can detect.
The distinction matters because questionnaires are not thermometers. A knee form does not read out a quantity of knee. It converts answers about limping, stairs and swelling into a number that carries the imprecision of every step in the conversion. Ask one person twice in a week and you may get two different answers. The threshold is the field's attempt to draw a line under all that wobble and say: below here, we should not claim anything happened.
Why a statistically significant result can still be meaningless
Statistical significance answers one narrow question: is this difference likely to be a fluke? It says nothing whatsoever about size. Recruit enough participants and a difference far too small for any human being to feel will clear the significance bar comfortably. This is not a hypothetical failure mode. It is the ordinary, published result of several of the largest trials in musculoskeletal care.
The cleanest example sits in low back pain. Early referral to physical therapy for recent-onset low back pain produced a small but statistically significant improvement in disability at three months compared with usual care — and by one year, the difference between the two groups was no longer clinically important 1Ref 1Fritz JM, Magel JS, McFadden M, et al. (2015).Early Physical Therapy vs Usual Care in Patients With Recent-Onset Low Back Pain: A Randomized Clinical Trial.That early physical therapy for recent-onset low back pain produced a small statistically significant improvement in disability at 3 months versus usual care, and that the between-group difference was not clinically important at 1 year — the article's central worked example of statistical significance diverging from clinical importance..
A statistically significant three-month gain from early physical therapy was no longer clinically important by one year 1Ref 1Fritz JM, Magel JS, McFadden M, et al. (2015).Early Physical Therapy vs Usual Care in Patients With Recent-Onset Low Back Pain: A Randomized Clinical Trial.That early physical therapy for recent-onset low back pain produced a small statistically significant improvement in disability at 3 months versus usual care, and that the between-group difference was not clinically important at 1 year — the article's central worked example of statistical significance diverging from clinical importance..
The same shape recurs in medication trials, where it arguably matters more because the drugs are so widely taken:
- Anti-inflammatories for chronic low back pain. Slightly more effective than placebo for short-term pain and disability — but the effect is small and may not be clinically important 2Ref 2Enthoven WTM, Roelofs PDDM, Deyo RA, van Tulder MW, Koes BW (2016).Non-steroidal anti-inflammatory drugs for chronic low back pain.That NSAIDs are slightly more effective than placebo for short-term pain and disability in chronic low back pain, but that the effect is small and may not be clinically important — an example of a real effect falling below a meaningful threshold..
- Paracetamol for spinal pain and arthritis. Ineffective for low back pain outright, and in hip and knee osteoarthritis it produces only a small effect on pain and disability that does not reach clinical importance 3Ref 3Machado GC, Maher CG, Ferreira PH, et al. (2015).Efficacy and safety of paracetamol for spinal pain and osteoarthritis: systematic review and meta-analysis of randomised placebo controlled trials.That paracetamol is ineffective for low back pain and produces only a small, not clinically important effect on pain and disability in hip and knee osteoarthritis — a second example of a measurable but sub-threshold effect..
Each of those studies found something. What it found was too small to matter — and only a threshold like this one makes that sentence sayable.
Where the line comes from
The threshold is not a property of the questionnaire, and it is not calculated from the form itself. It is established by research, and the underlying move is disarmingly simple: ask people. Everyone completes the questionnaire, undergoes some period of care, completes it again, and is separately asked a plain human question — compared with when you started, are you better, about the same, or worse?
That plain question is the reference point. The score change reported by people who said they were a little better — the smallest improvement anyone bothers to name — is what the threshold gets built around. The definition of important therefore comes from patients rather than statisticians or surgeons, which is the quietly radical thing about the idea.
- It is a convention, not a constant. Reasonable researchers make different choices about how to ask, whom to ask, and where to draw the line. Different studies produce different thresholds for the same form.
- It is a judgement about people, not a fact about tissue. Nothing in the biology of a hip determines the number. It is a summary of what people said mattered.
A threshold quoted without naming its questionnaire and its population is an unfinished sentence.
Why the threshold is different for every questionnaire
There is no universal MCID. Each questionnaire carries its own, because each is built on a different scale, asks about a different joint, and reacts differently to the same underlying change in the same person. A threshold borrowed from one instrument and applied to another means nothing, however confidently it gets quoted.
There is no universal MCID. It belongs to one specific questionnaire in one specific population, and it does not travel between forms.
This is easiest to see when two forms cover the same joint. The hoos questionnaire and the womac index both measure hip pain and function, and they do not move in step: the HOOS is more responsive than the WOMAC on its pain and symptom subscales 4Ref 4Nilsdotter AK, Lohmander LS, Klässbo M, Roos EM (2003).Hip disability and osteoarthritis outcome score (HOOS)—validity and responsiveness in total hip replacement.That the HOOS is a validated hip patient-reported outcome measure and is more responsive than the WOMAC on its pain and symptom subscales — used to show that two instruments covering the same joint do not move by the same amount for the same underlying change, so a threshold cannot be shared between them.. Same hip, same patient, same real improvement — different amounts of score movement. If one improvement produces two different numbers depending on which form was handed to you, the line for "this counted" has to be drawn separately on each form.
Responsiveness is the property underneath this. A questionnaire that barely twitches when someone genuinely improves needs a small threshold and will still miss real gains. A form scored across several subscales has, in effect, several thresholds — a person can improve on pain while nothing at all moves on sport and recreation.
What the threshold looks like inside a real surgical trial
The idea does its hardest work in surgery, because surgery carries a large and entirely legitimate placebo component: the anaesthetic, the incision, the recovery, the attention, the months of enforced rest. Reading surgical trials well means separating the effect of the procedure from the effect of having been operated on. The only clean way to do that is to operate on everyone and perform the actual procedure in just one group.
That is what a placebo-controlled shoulder trial did. Arthroscopic subacromial decompression for subacromial shoulder pain provided no clinically important benefit over placebo surgery — arthroscopy alone, without the decompression — or over no treatment at all 5Ref 5Beard DJ, Rees JL, Cook JA, et al. (CSAW) (2018).Arthroscopic subacromial decompression for subacromial shoulder pain (CSAW): a multicentre, pragmatic, parallel group, placebo-controlled, three-group, randomised surgical trial.That arthroscopic subacromial decompression provided no clinically important benefit over placebo surgery or over no treatment for subacromial shoulder pain — the article's illustration of a clinical-importance threshold applied to a between-group difference in a placebo-controlled surgical trial..
Notice what the threshold is doing there. The trial is not claiming nobody got better. It is claiming that the gap between the groups did not clear the line. Sham surgery methodology is what makes that a meaningful sentence rather than a rhetorical one: without a placebo arm, every improvement in the operated group would have been credited to the operation, and the score would have looked like a triumph. The threshold is the second filter. It asks whether the difference that survived is one a patient would have noticed.
When the difference does clear the bar
It would be a serious misreading of this literature to conclude that surgery does not work. Some of these trials find the opposite, using exactly the same threshold logic. For femoroacetabular impingement syndrome, hip arthroscopy produced modestly better patient-reported hip function at twelve months than personalised physiotherapist-led conservative care — at substantially higher cost 6Ref 6Griffin DR, Dickenson EJ, Wall PDH, et al. (UK FASHIoN) (2018).Hip arthroscopy versus best conservative care for the treatment of femoroacetabular impingement syndrome (UK FASHIoN): a multicentre randomised controlled trial.That for femoroacetabular impingement syndrome, hip arthroscopy led to modestly better patient-reported hip function at 12 months than personalised physiotherapist-led conservative care, at substantially higher cost — the article's counterweight example of a difference that does clear a meaningful threshold..
The question is never "surgery or not." It is which operation, for which problem, in which order, and at what cost.
That result is the honest counterweight. The benefit was real, it was modest, and it came with a price a score sheet does not show. The threshold did not make the decision. It made the decision legible.
The other guard rail is the population. None of these trials studied a broken bone, a joint that will not stay located, a tendon torn through by a sudden injury, or nerve compression worsening week by week. They studied painful, largely degenerative conditions in stable joints. A result does not travel outside the population it enrolled, and treating one as a verdict on surgery as a category is the same error as quoting a threshold without naming its questionnaire. The frame is a sequence of care.
What it means for a score of your own
A trial reports an average, and the threshold gets compared against that average. But an average is not a description of anyone inside it. The same average can come from a group where everybody changed alike, or from a group where a third improved dramatically, a third did nothing, and a third got worse. It is not a forecast for one person, in either direction.
- "It missed the threshold, so this cannot help me." It means the average benefit was too small to matter. It does not identify who, if anyone, did well.
- "It cleared the threshold, so this will help me." It means the average benefit was noticeable. That is still not a prediction about you.
- Your own judgement counts. A clinically meaningful difference on paper and a meaningful difference to you are not obliged to coincide. If the change you got is the one that lets you sleep on that side again, it counts.
- The trend beats any single reading. One score is a snapshot taken on one day, in one mood. Three or four fills describe a direction, and a direction is what is worth discussing.
- Cost and burden are not on the form. The score sheet has no box for months of recovery, time off work, or what the treatment costs. A benefit that clears the threshold can still be a poor trade.
A score that moved less than the threshold is not a wasted three months. It is one instrument's opinion, formed on one day, from eight or ten questions — and it is built to be less sensitive than you are.
Common questions
Related
Muscle, joint & pain
The Difference Between Statistically Better and Actually BetterMuscle, joint & pain
The Global Rating of Change, a One-Question CheckMuscle, joint & pain
The Patient-Specific Functional Scale: Define Your Own Goals
Say it back
How would you explain this to someone you love?
Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.
What a score sheet has no box for
- —New weakness in an arm or leg that is worsening over days — a grip that keeps failing, or a foot that catches on the floor when you walk
- —Loss of bladder or bowel control, or numbness across the area that would contact a saddle, occurring alongside back pain
- —A single joint that is hot, red and swollen with fever, or a joint you cannot bear any weight through after an injury
- —Unexplained weight loss, or night pain that wakes you and does not settle when you change position
Loss of bladder or bowel control with back pain, or weakness that is visibly worsening day by day, is an emergency department visit the same day rather than a questionnaire question. Call 911 if you cannot get there safely.
This article explains how outcome scores are interpreted in research and in clinical practice. It is general education about a statistical idea, not medical advice, and it cannot assess your condition or tell you what any score of yours means. Decisions about your care belong with a clinician who can examine you.
References
- 1.Fritz JM, Magel JS, McFadden M, et al. (2015). Early Physical Therapy vs Usual Care in Patients With Recent-Onset Low Back Pain: A Randomized Clinical Trial. JAMA. doi:10.1001/jama.2015.11648 ✓That early physical therapy for recent-onset low back pain produced a small statistically significant improvement in disability at 3 months versus usual care, and that the between-group difference was not clinically important at 1 year — the article's central worked example of statistical significance diverging from clinical importance.
- 2.Enthoven WTM, Roelofs PDDM, Deyo RA, van Tulder MW, Koes BW (2016). Non-steroidal anti-inflammatory drugs for chronic low back pain. Cochrane Database of Systematic Reviews. doi:10.1002/14651858.CD012087 ✓That NSAIDs are slightly more effective than placebo for short-term pain and disability in chronic low back pain, but that the effect is small and may not be clinically important — an example of a real effect falling below a meaningful threshold.
- 3.Machado GC, Maher CG, Ferreira PH, et al. (2015). Efficacy and safety of paracetamol for spinal pain and osteoarthritis: systematic review and meta-analysis of randomised placebo controlled trials. BMJ. doi:10.1136/bmj.h1225 ✓That paracetamol is ineffective for low back pain and produces only a small, not clinically important effect on pain and disability in hip and knee osteoarthritis — a second example of a measurable but sub-threshold effect.
- 4.Nilsdotter AK, Lohmander LS, Klässbo M, Roos EM (2003). Hip disability and osteoarthritis outcome score (HOOS)—validity and responsiveness in total hip replacement. BMC Musculoskeletal Disorders. PMID 12777182 ✓That the HOOS is a validated hip patient-reported outcome measure and is more responsive than the WOMAC on its pain and symptom subscales — used to show that two instruments covering the same joint do not move by the same amount for the same underlying change, so a threshold cannot be shared between them.
- 5.Beard DJ, Rees JL, Cook JA, et al. (CSAW) (2018). Arthroscopic subacromial decompression for subacromial shoulder pain (CSAW): a multicentre, pragmatic, parallel group, placebo-controlled, three-group, randomised surgical trial. The Lancet. doi:10.1016/S0140-6736(17)32457-1That arthroscopic subacromial decompression provided no clinically important benefit over placebo surgery or over no treatment for subacromial shoulder pain — the article's illustration of a clinical-importance threshold applied to a between-group difference in a placebo-controlled surgical trial.
- 6.Griffin DR, Dickenson EJ, Wall PDH, et al. (UK FASHIoN) (2018). Hip arthroscopy versus best conservative care for the treatment of femoroacetabular impingement syndrome (UK FASHIoN): a multicentre randomised controlled trial. The Lancet. doi:10.1016/S0140-6736(18)31202-9That for femoroacetabular impingement syndrome, hip arthroscopy led to modestly better patient-reported hip function at 12 months than personalised physiotherapist-led conservative care, at substantially higher cost — the article's counterweight example of a difference that does clear a meaningful threshold.
6 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — citations link their sources. Editorial policy