Muscle, joint & pain

How to Read a Surgical Trial Before You Trust It

Save

A surgery study can be technically positive and still tell you almost nothing. This is how to weigh one — the comparator, who crossed between groups, whether anyone was blinded, how long patients were followed, and the distance between a statistically significant result and a difference a patient would actually feel.

Last updated: July 2026

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

Why is a surgery study harder to trust than a pill study?

A surgery study is harder to trust because it is harder to blind. In a good drug trial, a placebo pill looks identical to the real one, so belief and expectation cancel out on both sides. Surgery carries a large placebo effect of its own — the anaesthetic, the recovery ritual, the conviction that something has been fixed — so the only way to isolate what the operation does is to compare it against a sham operation that reproduces everything except the definitive step.

A sham operation is a placebo procedure — the incisions, the theatre, and the recovery, but without the surgical act being tested. When surgeons have run these trials, the results have been humbling. In one, arthroscopic partial meniscectomy for a worn meniscus produced the same symptom relief a year later as a sham knee operation in which no tissue was removed 1. In a placebo-controlled shoulder trial, arthroscopic subacromial decompression gave no clinically important benefit over a placebo arthroscopy — or over no surgery at all 2.

That is what sham surgery, sometimes called placebo surgery, reveals: some operations work mainly through the ceremony around them. A trial that compares surgery only against a waiting list, or against no treatment, cannot separate the operation from that ceremony. The first question to ask of any surgical trial is not whether surgery beat nothing, but what, exactly, surgery was measured against.

Statistical significance is not the same as feeling better

A result can be statistically significant — unlikely to be due to chance — and still be far too small for any patient to notice. Those are two different questions. Statistical significance asks whether an effect is real; clinical significance asks whether it is big enough to matter to the person living in the body. A large trial can detect a difference so tiny that no one would trade a scar for it.

The minimal clinically important difference is the smallest change on a pain or function scale that a patient can actually feel. It is the bar a real effect has to clear. Early physical therapy for recent-onset low back pain, for example, produced a small, statistically significant improvement in disability at three months — but by one year the gap over usual care was no longer clinically important 3. The finding was true and the effect was trivial. Both things can be so at once.

When you read a trial, look past the p-value to the effect size, and compare it against that clinically meaningful difference. A useful habit: ask what the number means in plain terms — a few points on a hundred-point scale, a week less of pain, one fewer flare a year. If a study reports only that a result was significant without telling you how large it was, it has answered the easy question and dodged the one that decides whether an operation is worth it.

Who actually got the operation? Crossover and intent-to-treat

Trials randomly assign people to surgery or to non-surgical care, but people do not always stay put. Some assigned to wait get worse and choose surgery; some assigned to surgery improve while waiting and cancel. This switching is called crossover, and when a lot of it happens, the tidy comparison of two groups blurs into two mixed populations.

A trial's answer depends on whether it counts patients by the group they were assigned to or by the treatment they actually received. Counting by assignment — intent-to-treat — is the honest, conservative method, because it preserves the randomisation that makes the groups comparable in the first place. But when crossover is high, intent-to-treat can wash a real difference out to nothing. In SPORT, the large trial of surgery versus non-operative care for a herniated lumbar disc, so many patients crossed between groups that the head-to-head comparison was inconclusive — even though both groups improved substantially 4.

That is not a failed trial; it is a trial telling you something real. High crossover usually means the condition improves either way and the decision is rarely urgent — the same pattern that shows up when people read about microdiscectomy for sciatica. When you meet a study whose headline is "no significant difference," check the crossover rate before concluding surgery does nothing. Sometimes it means the two paths lead to a similar place, and the choice is about how fast you want to get there, not whether you arrive.

One trial or many? What a systematic review adds

A single trial is one data point, run in one set of clinics on one kind of patient. A systematic review gathers every trial that has asked the same question, pools them, and grades how much confidence the total picture deserves. That grade — often described as high, moderate, low, or very low certainty — is doing quiet work: it tells you whether a finding is likely to hold up or likely to change with the next study.

The shoulder decompression story shows why this matters. A Cochrane review pooling the placebo-controlled trials concluded, with high-certainty evidence, that subacromial decompression does not deliver clinically important benefits over placebo or non-surgical care 5. High certainty means further trials are unlikely to overturn it — a far stronger statement than any one study could make.

When you read about a surgery, look for the systematic review or the clinical practice guideline that sits above the individual trials, and read its certainty rating before you read its conclusion. A guideline that recommends against a procedure on high-certainty evidence is a different animal from a single small trial with a surprising result. The individual trials are the raw material; the graded review is the finished judgement. Both matter, but they answer at different levels of confidence, and confusing the two is how a single provocative study gets treated as settled fact.

When the trial favors surgery, read it just as carefully

Sometimes surgery does win, and the same discipline applies. When a trial favors an operation, the question is not over — it is beginning. How much did surgery win by, measured against the difference a patient would feel? For how long? At what cost and risk? And for which patients, since a trial's average can hide people who did much better and much worse.

Surgery winning a trial is the start of the question — how much, for whom, at what cost — not the end of it. In the UK FASHIoN trial, hip arthroscopy for femoroacetabular impingement produced modestly better hip function at one year than a physiotherapist-led program — a real edge, but a small one, achieved at substantially higher cost 6. That is a defensible reason to operate for some people and a poor one for others; the trial gives you the size of the trade, not the decision.

None of this is an argument against surgery. It is an argument for sequence — trying the lower-risk, reversible option first when the evidence says outcomes are similar, and reserving the operation for when it earns its place. And some situations do not belong in this weighing at all. Surgery is clearly the right call, and often urgent, for progressive nerve damage or muscle weakness, loss of bowel or bladder control, an unstable fracture or dislocation, a joint destroyed by advanced arthritis that no longer responds to anything, a deep infection, or a tumor. Those are not elective decisions to be paced against a rehab program, and knowing the difference between an elective and an urgent operation is part of reading the evidence honestly.

A reader's checklist for any surgery study

You can size up most surgical trials with a short list of questions, none of which requires a medical degree. Run through them before you accept a headline, and the difference between a strong study and a weak one usually becomes visible within a paragraph or two of the results.

  • What was the comparator? Surgery versus a sham operation is the gold standard. Surgery versus no treatment or a waiting list cannot separate the operation from its placebo effect.
  • Was anyone blinded? In the best surgical trials, the patient and the people measuring outcomes do not know who got the real procedure.
  • How much crossover was there? High switching between groups blurs the comparison and often signals that both paths improve.
  • How big was the effect, in plain terms? Compare it against the smallest change a patient could actually feel, not just the p-value.
  • How long were patients followed? An advantage at six weeks that has vanished by one year is a different finding from one that lasts.
  • Who was in the trial? Age, severity, and whether the problem was degenerative or from a fresh injury decide whether the result applies to you.
  • What were the harms and the cost? Every operation carries risk; a small benefit and a real complication rate change the math.

A study that answers these questions well deserves weight. One that leaves several blank deserves caution, however confident its abstract sounds.

Turning a study into a decision with your surgeon

A trial describes an average patient across a population; you are one person with your own goals, and the study is an input to a conversation, not a verdict handed down. The point of reading it well is to walk into the room able to ask better questions and to take part in the decision rather than receive it. This is what shared decision making and informed consent are supposed to mean in practice.

Good questions before surgery grow directly out of the checklist: What does the best evidence say this operation adds over not having it, and how large is that difference? What happens if I wait, and is waiting safe in my case? What has to be true about my problem for surgery to be the right call, and is it true? What are the odds and the stakes of a complication? The answers should line up with the trials, and where they do not, that gap is worth naming out loud.

The same reading applies across orthopedics — whether the choice is acl surgery vs rehab, carpal tunnel surgery vs a splint, or a spinal decompression against a course of physical therapy. The evidence rarely says "never operate." More often it says the sequence matters: start with what is reversible and low-risk when outcomes are similar, escalate deliberately, and keep the operation for the situations where a trial — read carefully — actually shows it pulling its weight.

Common questions

Sham surgery is a placebo operation: the surgeon reproduces the incisions and the procedure but omits the step being tested, so patients cannot tell which they received. It is used only under strict ethical review, with informed consent, when the operation's real benefit is genuinely uncertain. These trials have repeatedly shown that some common procedures work mainly through their placebo effect.

It means the study did not detect a reliable gap between the two groups. That can happen because the treatments truly work about equally, or because high crossover between groups blurred the comparison, or because the trial was too small. It does not automatically mean surgery is useless — check the crossover rate, the effect size, and how outcomes moved in both groups before concluding.

Statistical significance only tells you an effect is unlikely to be chance, not that it is large. A big trial can detect a difference too small for any patient to feel. The relevant bar is the minimal clinically important difference — the smallest change a person actually notices. Always read the size of the effect in plain terms, not just whether it was significant.

A clinical practice guideline or systematic review that pools all the trials and grades the certainty of the evidence generally carries more weight than any single study, which reflects one group of patients in one set of clinics. A guideline recommending for or against a procedure on high-certainty evidence is a stronger signal than a lone trial with a surprising headline.

No. The evidence supports a sequence of care, not avoidance. When trials show non-surgical and surgical outcomes are similar, trying the lower-risk option first is reasonable. But surgery is clearly indicated — often urgently — for progressive nerve or muscle weakness, loss of bowel or bladder control, unstable fractures, deep infection, tumors, and advanced joint destruction. Those decisions are not paced against a rehab program.

Related

Say it back

How would you explain this to someone you love?

Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

When a wait-and-read approach is the wrong frame

  • Progressive weakness or numbness, a foot that drags, or a hand losing its grip over days
  • New loss of bowel or bladder control, or numbness in the saddle area between the legs
  • A joint that is hot, swollen, and red with fever — a possible joint infection
  • Severe pain or deformity after a fall or accident, suggesting a fracture or dislocation

Loss of bowel or bladder control with back or leg symptoms, or a suspected fracture or joint infection, is an emergency — go to the emergency department or call 911; do not wait to weigh the evidence.

This article explains how to read surgical research and is educational only. It is not medical advice and cannot tell you whether a specific operation is right for you. Decisions about surgery should be made with a qualified clinician who knows your history and can examine you.

References

  1. 1.Sihvonen R, Paavola M, Malmivaara A, et al. (FIDELITY) (2013). Arthroscopic Partial Meniscectomy versus Sham Surgery for a Degenerative Meniscal Tear. New England Journal of Medicine. doi:10.1056/NEJMoa1305189Used as a worked example of a sham-controlled surgical trial: arthroscopic partial meniscectomy was no better than sham surgery for a degenerative meniscal tear.
  2. 2.Beard DJ, Rees JL, Cook JA, et al. (CSAW) (2018). Arthroscopic subacromial decompression for subacromial shoulder pain (CSAW): a multicentre, pragmatic, parallel group, placebo-controlled, three-group, randomised surgical trial. The Lancet. doi:10.1016/S0140-6736(17)32457-1Cited as a placebo-controlled surgical trial in which subacromial decompression gave no clinically important benefit over placebo arthroscopy or no treatment.
  3. 3.Fritz JM, Magel JS, McFadden M, et al. (2015). Early Physical Therapy vs Usual Care in Patients With Recent-Onset Low Back Pain: A Randomized Clinical Trial. JAMA. doi:10.1001/jama.2015.11648Illustrates statistical versus clinical significance: a small statistically significant disability improvement at 3 months that was no longer clinically important at 1 year.
  4. 4.Weinstein JN, Tosteson TD, Lurie JD, et al. (SPORT) (2006). Surgical vs Nonoperative Treatment for Lumbar Disk Herniation: The Spine Patient Outcomes Research Trial (SPORT): A Randomized Trial. JAMA. linkUsed to explain crossover and intent-to-treat: high crossover made the head-to-head comparison inconclusive while both surgical and nonoperative groups improved substantially.
  5. 5.Karjalainen TV, Jain NB, Page CM, et al. (2019). Subacromial decompression surgery for rotator cuff disease. Cochrane Database of Systematic Reviews. doi:10.1002/14651858.CD005619.pub3Used to explain what a graded systematic review adds: high-certainty evidence that subacromial decompression provides no clinically important benefit over placebo or non-surgical care.
  6. 6.Griffin DR, Dickenson EJ, Wall PDH, et al. (UK FASHIoN) (2018). Hip arthroscopy versus best conservative care for the treatment of femoroacetabular impingement syndrome (UK FASHIoN): a multicentre randomised controlled trial. The Lancet. doi:10.1016/S0140-6736(18)31202-9Used as an example of a trial that modestly favors surgery: hip arthroscopy gave slightly better hip function at 12 months than conservative care, at substantially higher cost.

6 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — citations link their sources. Editorial policy