Muscle, joint & pain

The Difference Between Statistically Better and Actually Better

Save

Trials report two different things and readers routinely merge them. Statistical significance says a difference probably isn't noise. Clinical meaningfulness says the difference is big enough to change a life. Most of the modern musculoskeletal evidence turns on the gap between the two — treatments that are real, measurable, and still too small for the person receiving them to notice.

Last updated: July 2026

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

What does clinically meaningful improvement mean?

Clinically meaningful improvement is a change big enough to register in a person's life: less pain climbing stairs, a return to sleeping through the night, a shoulder that reaches the top shelf. Researchers turn that idea into a number — a threshold on whatever questionnaire the trial used — and treatments that move people less than the threshold are counted as not clinically important, however certain the finding.

The threshold has a formal name: the minimal clinically important difference, or MCID — the smallest change a patient would call worthwhile. It is deliberately a patient's judgement rather than a statistician's. Researchers ask people at the end of a study whether they feel better, then read off how far their scores actually moved. The change separating "somewhat better" from "no real change" becomes the threshold.

A near neighbour does a different job. The minimal detectable change asks not whether a shift matters but whether it is measurable at all — how far a score must move to clear the instrument's own noise. A change can pass one bar and fail the other.

Why isn't statistical significance the same thing?

Statistical significance answers one narrow question: is this difference likely to be more than chance? It says nothing about size. Feed enough people into a trial and a difference so small no one could feel it will still come back significant, because significance is partly a function of sample size. Meaningfulness asks the question significance cannot: is the difference big enough to matter?

The two get merged constantly. A press release says a treatment "significantly improved" pain — the word doing statistical work while being heard as ordinary English, where significant means substantial. In the paper it means nearer to detectable.

A p-value tells you a difference is probably there. It never tells you the difference is worth having.

This bites hardest in exactly the trials built to settle things. Large, expensive studies are powered to detect small differences — that is why they are made large. So "statistically significant" in a big trial should open the question of how big, not close it.

Where does the threshold come from, and what can't it tell you?

It comes from patients, by one of two routes: the anchor-based method asks whether people feel better and reads off how far their scores moved, while the distribution-based method works from the instrument's measurement error. Neither yields a universal number. Thresholds differ by questionnaire, by condition, and by how badly affected people were at the start — and none of them can tell you what will happen to you.

That last point surprises people. Someone in severe pain generally needs a larger absolute change before calling it worthwhile than someone mildly affected. And a threshold is applied to a group average, which hides its parts: inside a trial reporting no clinically important difference are people who improved enormously and people who got worse.

A trial finding no meaningful benefit on average is not a statement about whether your pain is real. It is a statement about which levers reliably move it.

That cuts both ways, which is why it is rarely quoted honestly. A disappointing average does not prove nobody was helped. But "it worked for some people" is not evidence either — after any treatment for a fluctuating condition, some people are better next month and would have been regardless. Telling those apart needs a comparison group, the same arithmetic that decides how to read prp evidence.

What does the gap look like in real trials?

It looks like a study that succeeds and fails at the same time. Several of the largest musculoskeletal trials of the last decade found differences that were statistically real and, measured against the accepted threshold, too small to matter to a patient. It is the finding, not a technicality, and an abstract's opening sentence can easily obscure it.

TreatmentWhat the trial found
Early physical therapy, recent-onset low back painA small, statistically significant disability improvement at three months versus usual care; by one year the between-group difference was not clinically important 1
Paracetamol for spinal pain and osteoarthritisIneffective for low back pain; for hip and knee osteoarthritis a small effect on pain and disability, not clinically important 2
Oral NSAIDs for chronic low back painSlightly more effective than placebo short-term, with an effect that is small and may not be clinically important 3
Arthroscopic subacromial decompressionHigh-certainty evidence of no clinically important benefit over placebo surgery or non-surgical care 4

The shoulder is the clearest case, because the operation was not visibly failing anyone. Patients had it and got better. The placebo-controlled trial showed decompression produced no clinically important benefit over arthroscopy alone or over no treatment at all 5 — people improved either way, and the improvement had been credited to the procedure for years. That is the machinery behind sham surgery methodology, and it is why reading surgical trials with the threshold in hand changes what a person would choose.

When the difference is real but modest, who decides if it's enough?

You do, with help — because "modest" is not a verdict, it is a measurement waiting for a judgement. Some trials find a benefit that clears statistical significance, sits at or near the threshold, and arrives with real costs in money, recovery time, and risk. Nothing in the statistics resolves that trade. It resolves where the numbers meet a particular life.

Hip arthroscopy for femoroacetabular impingement syndrome is the instructive counter-case. Against personalised physiotherapist-led conservative care, it produced modestly better patient-reported hip function at twelve months — at substantially higher cost 6. That is a genuinely different result from the shoulder one and deserves to be read differently. A benefit that is real and modest is not a benefit that is absent.

It is also where the rest of the arithmetic starts: what you would owe once you know your surgery cost after deductible, how long before you could return to work after surgery, and what the recovery asks of the people you live with. None of that is in the trial; all of it is in the decision. The same applies to any claim that minimally invasive surgery is better — a smaller incision is a real difference; a clinically important difference a year later is a separate question.

When "clinically meaningful" is the wrong question entirely

Some operations are not judged by a questionnaire at all, and applying the threshold to them is a category error. When the problem is mechanical or dangerous — a fracture that will not hold its alignment, a joint physically blocked by a displaced fragment, an infected joint, a nerve losing power week by week — the operation is judged on whether it fixes the structure. Clinicians generally treat those as clear indications, and the trials above have nothing to say about them.

The studies that found no clinically important benefit were asking about one specific kind of surgery: elective, offered for pain, in people who could reasonably wait, with a real non-surgical alternative available. That is the setting the threshold was built for, and where most orthopaedic decisions live. But the boundary is real, and blurring it turns a useful idea into a harmful one.

The threshold governs elective surgery offered for pain. It does not govern surgery that fixes a mechanical or dangerous problem.

Where the arithmetic does not decide:

  • A joint that is genuinely locked, not merely painful — something is physically in the way.
  • Neurological loss that is progressing — weakness measurably worse than last month, a foot that catches on stairs.
  • Fractures that are displaced or unstable, and joints that dislocate rather than ache.
  • Infection in or around a joint, a surgical problem from the first hour.

None of this argues against operating. It argues for knowing which question you are in before you answer it.

How to use this in the room

The question that does the work is short: how much better, on what scale, compared with what? It can be asked of any number offered — a surgeon's, a physical therapist's, an advertisement's. Someone who can answer it, including naming the threshold on the scale they are quoting, is telling you something real. Someone who can only say "it's proven" is telling you something else.

What tends to open a number up:

  • Compared with what? Placebo, no treatment, or the other thing I could do instead — three very different answers.
  • How many points, out of how many? A benefit described only as "significant" has been rounded into a word.
  • What counts as a meaningful change on that scale, and was the number set before the results were known?
  • How long did it last? Several treatments here work at three months and have stopped separating from the comparison by a year 1.

A clinically meaningful difference is the only kind worth reorganising a life around. Everything smaller is a number that is true and does not help.

Common questions

No. Statistical significance means a difference probably isn't chance. It says nothing about whether the difference is large enough for anyone to notice. Large trials are specifically designed to detect small differences, so significance in a big study should prompt the next question — how big? — rather than settle it.

The minimal clinically important difference is the smallest change a patient would call worthwhile. The minimal detectable change is the smallest change that exceeds the instrument's own measurement error. One is about mattering, the other about measuring. A shift can clear one and not the other, and papers sometimes blur them.

Often they don't disagree about the effect size — they disagree about the yardstick. Different questionnaires carry different thresholds, and thresholds shift with how severely affected people were at baseline. Two papers can report the same underlying benefit and reach opposite verdicts on whether it was clinically important.

Not quite. It means that across the group studied, the benefit over the comparison was too small to count on. Individuals inside that average did improve, and some improved a great deal — but people in the comparison group improved too. Without the comparison, individual improvement proves nothing.

No, and the framework is precise about where it applies. It governs elective operations offered for pain, where waiting is reasonable and a non-surgical option exists. It does not govern surgery for a mechanically blocked joint, progressive nerve weakness, an unstable fracture, or an infection — those are judged on whether the structure gets fixed.

It is usually a sentence in the methods, near the description of the primary outcome, naming the instrument and the number that counts as important. Papers that pre-specify it and cite its source are the ones most worth trusting. If no threshold appears anywhere, the paper is reporting size without context.

Related

Say it back

How would you explain this to someone you love?

Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

Numbers are for elective decisions — these symptoms are not

  • New or worsening weakness in a foot, leg, or hand — catching a toe on stairs, a foot that slaps the ground, dropping objects you could hold last month
  • Numbness in the groin or inner thighs, or new loss of bladder or bowel control, alongside back pain
  • Fever with new spinal or joint pain, or a single joint that becomes hot, swollen, and too painful to move
  • Pain that wakes you every night and is not eased by any position change, especially with unexplained weight loss

Numbness in the saddle area, new loss of bladder or bowel control, or rapidly progressing leg weakness with back pain needs emergency assessment the same day — go to an emergency department or call 911. A hot, swollen, immovable joint with fever needs urgent evaluation, not a wait-and-see.

This explains how researchers and clinicians read trial results. It is general education, not medical advice, and it cannot tell you what a specific number means for your body or your decision. Bring the questions here to the clinician who knows your case.

References

  1. 1.Fritz JM, Magel JS, McFadden M, et al. (2015). Early Physical Therapy vs Usual Care in Patients With Recent-Onset Low Back Pain: A Randomized Clinical Trial. JAMA. doi:10.1001/jama.2015.11648Early physical therapy for recent-onset low back pain produced a small statistically significant disability improvement at 3 months versus usual care, with between-group differences that were not clinically important at 1 year — used as the worked example of a statistically real, clinically unimportant difference and of benefits that narrow over time.
  2. 2.Machado GC, Maher CG, Ferreira PH, et al. (2015). Efficacy and safety of paracetamol for spinal pain and osteoarthritis: systematic review and meta-analysis of randomised placebo controlled trials. BMJ. doi:10.1136/bmj.h1225Paracetamol is ineffective for low back pain and produces only a small, not clinically important effect on pain and disability in hip and knee osteoarthritis — used as an example of significance without clinically important size.
  3. 3.Enthoven WTM, Roelofs PDDM, Deyo RA, van Tulder MW, Koes BW (2016). Non-steroidal anti-inflammatory drugs for chronic low back pain. Cochrane Database of Systematic Reviews. doi:10.1002/14651858.CD012087NSAIDs are slightly more effective than placebo for short-term pain and disability in chronic low back pain, but the effect is small and may not be clinically important — used as an example of an honestly reported modest benefit.
  4. 4.Karjalainen TV, Jain NB, Page CM, et al. (2019). Subacromial decompression surgery for rotator cuff disease. Cochrane Database of Systematic Reviews. doi:10.1002/14651858.CD005619.pub3High-certainty evidence that subacromial decompression surgery provides no clinically important benefit over placebo or non-surgical care for rotator cuff disease — used as the surgical example where the clinical-importance threshold, not statistical significance, carries the result.
  5. 5.Beard DJ, Rees JL, Cook JA, et al. (CSAW) (2018). Arthroscopic subacromial decompression for subacromial shoulder pain (CSAW): a multicentre, pragmatic, parallel group, placebo-controlled, three-group, randomised surgical trial. The Lancet. doi:10.1016/S0140-6736(17)32457-1Arthroscopic subacromial decompression provided no clinically important benefit over placebo (arthroscopy only) or over no treatment for subacromial shoulder pain — used to show that improvement after a procedure had been credited to the procedure without a comparison group.
  6. 6.Griffin DR, Dickenson EJ, Wall PDH, et al. (UK FASHIoN) (2018). Hip arthroscopy versus best conservative care for the treatment of femoroacetabular impingement syndrome (UK FASHIoN): a multicentre randomised controlled trial. The Lancet. doi:10.1016/S0140-6736(18)31202-9Hip arthroscopy for femoroacetabular impingement syndrome produced modestly better patient-reported hip function at 12 months than personalised physiotherapist-led conservative care, at substantially higher cost — used as the counter-case where a benefit is real and modest and the trade-off passes to the patient.

6 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — citations link their sources. Editorial policy