Muscle, joint & pain

Why the Same Questionnaire Comes Back After Surgery

Save

It feels like paperwork and it is not. Outcome questionnaires were designed to be answered twice: the Oxford Knee Score was built by asking 117 people the same twelve questions before a knee replacement and again six months later. Without that first set of answers, the second set is a number floating free of anything to compare it against.

Last updated: July 2026

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

What the questionnaire before surgery is actually for

It builds the baseline, and the baseline is the whole apparatus. Nobody can interpret a post-operative score without knowing what it replaced. This is not a habit that grew up around the instruments — it is how they were constructed. The Oxford Knee Score, a 12-item questionnaire completed by the patient, was developed and validated prospectively in 117 people assessed before total knee replacement and again six months afterwards 1.

The Oxford Shoulder Score was built the same way: 12 items, 111 patients answering before shoulder surgery and again at six months, alongside the SF-36, the Stanford health assessment questionnaire, and the surgeon-rated Constant score 2. The paired design is not how these questionnaires happen to be used. It is what they are.

Which is why a clinic that misses the pre-operative form has lost something it cannot get back. There is no reconstructing how bad the knee was in March once it is September and the knee is better. Memory quietly revises the past to match the present.

Why the wording never changes

Because a change score is a subtraction, and subtraction needs the same units at both ends. If the questions shift between visits, the difference between the two totals stops being about you and starts being about the form. That is why the questionnaire can feel rigid, repetitive, and occasionally beside the point: the items were fixed during development and validation, and the instrument's value lives in that fixity.

The Oxford Knee Score's construct validity was established by showing that its scores moved with American Knee Society clinical scores and with the relevant sections of the SF-36 and the Stanford health assessment questionnaire 1. Those relationships were demonstrated for that exact set of items in that exact wording. Rewrite one question to sound friendlier and you have a new instrument with nothing behind it.

It also explains why a form may ask about something that has nothing to do with your life. It is not asking because a clinician wondered. It is asking because the item earns its place in the total.

Why the second questionnaire comes at six months, not six days

Because the early weeks measure the operation rather than the outcome. Both Oxford instruments were validated against a six-month follow-up point 12, and that choice reflects what the first months after orthopedic surgery are actually like. The first two weeks after surgery are dominated by swelling, wound pain, and a body still clearing an anesthetic. Sleep after surgery is broken for reasons that have nothing to do with whether the procedure worked.

Scored at three weeks, almost any joint replacement looks like a disaster. The follow-up window is placed to land after that noise has cleared.

The window also depends on the operation. A carpal tunnel release recovery runs on a different clock from a rotator cuff repair recovery, and an acl surgery recovery timeline runs longer than either, with return-to-sport milestones counted in months rather than weeks. Bunion surgery recovery has its own arc again. The date on the follow-up questionnaire is chosen to sit where the real answer lives, which is rarely convenient and never arbitrary.

The trap: two versions of one score run in opposite directions

This is the most confusing thing about these forms, and it is worth knowing before you compare your number to anything you read elsewhere. The paper that defined the Oxford Knee Score scored each of its 12 items from 1 to 5 and summed them to a total between 12 and 60 — in which lower was better, and 60 meant the worst symptoms 1.

The founding paper for the Oxford Shoulder Score did the same thing: 12 to 60, with 12 as the best possible outcome 2. The versions most clinics print today run from 0 to 48, and on those, higher is better. That rescoring arrived later, in a separate publication, and it deliberately reversed the direction. Two people can quote the same named instrument and mean precisely opposite things.

Before you compare your score to anything, ask which direction is good on this version. The question is not naive — the literature itself changed its mind.

What counts as a real improvement?

This is less settled than you would expect. The Oxford Knee Score paper demonstrated sensitivity to change — effect sizes comparable with the relevant SF-36 dimensions, and significantly greater change scores in patients who reported substantial improvement 1 — but it derived no threshold from that. No minimal clinically important difference appears in it, and no minimal detectable change either 1.

Neither does the shoulder paper, which showed a standardised effect size comparing favourably with the SF-36 and the health assessment questionnaire, and the sharpest discrimination of patients who said their shoulder was much better 2.

So when a surgeon quotes you a number of points that counts as a meaningful gain, that figure comes from later research. It is not printed in the paper that built the score.

The property itself is real and measurable. The quebec back pain disability scale, a 20-item self-administered instrument, was shown to detect significant change in disability over time and to separate groups expected to change in opposite directions — which is why its authors recommended it for monitoring individual patients and not only for clinical trials 3.

Where your answers go after the six-month form

Into the evidence that decides what the next person gets offered. The trials that shape surgical decisions are assembled out of exactly these questionnaires. When researchers compared early surgery against prolonged conservative treatment for sciatica caused by a lumbar disc herniation, early surgery gave faster relief of leg pain — but at one year, outcomes were similar between the two strategies 4.

That is a finding about timing rather than about winners, and it exists only because people filled in the same form again and again. The answer is not the same for every problem, which is the entire reason for asking. In patients with lumbar spinal stenosis without spondylolisthesis, decompressive surgery produced greater improvement than nonsurgical care over two years, while the nonsurgical patients also improved modestly and rarely got worse 5.

Two conditions, two different answers — faster but equal at a year in one, clearly better at two years in the other. Neither conclusion could have been reached by asking people, in general, how they were getting on.

When the after score is flat, or worse

It is information, not a verdict, and it belongs in the conversation rather than in your stomach at 2am. A score that has not moved says that the thing being measured has not changed. It does not say the operation failed, that nothing further can be done, or that you are imagining the trouble. It says the next appointment has a real subject, which is better than a next appointment spent guessing.

Sometimes the explanation is that the form is measuring what you are not doing rather than what you cannot do. Fear of movement after an operation is common and it suppresses a function score exactly as a mechanical problem would. A questionnaire cannot separate those two. A clinician who examines you can.

And sometimes the honest reading is that the procedure was not the right answer to that problem. Low-value care for low back pain — unnecessary imaging, opioids, injections, and surgery — is widespread globally and needs reducing 6. The only reason anyone knows that is that somebody measured the after.

Common questions

They have your answers from before. The point of the exercise is the difference between the two, and a difference needs both ends. The first form describes the problem you brought in; the second describes what is left of it. Neither number is interesting on its own, and neither can be worked out from the other.

These instruments were designed and validated to track individual patients, not only to fill research databases. Whether a particular clinic reviews yours before your appointment varies, and it is a fair thing to ask. If the answer is that nobody looks, that is worth knowing too — the form takes your time and should buy you something.

You are not expected to. That is precisely why the questionnaire is administered before the operation rather than reconstructed afterwards. Recall of past pain and disability drifts toward whatever is true now, in both directions, which is one of the better-known problems in outcome measurement and the reason baselines are collected in advance.

Read the wording at the top of the form, because it tells you: most instruments specify a time frame, such as the past week, and mean it. Answering about your worst-ever day when the form asks about the past week produces a baseline that nothing afterwards can match, which makes a genuinely good result look like a failure.

No. It is one input among the examination, the imaging where relevant, and what you say in the room. A questionnaire measures one construct through a fixed set of questions; it has no way of knowing about the thing it did not ask. Success is a judgment made with you, not a number crossing a line.

Related

Say it back

How would you explain this to someone you love?

Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

Symptoms that outrank any questionnaire

  • A surgical wound that grows redder, hotter, or starts draining fluid, especially alongside a fever
  • Calf pain and swelling on one side, or new shortness of breath or chest pain, after lower-limb surgery
  • New numbness, or weakness that makes a limb give way, that was not present when you left the hospital
  • Pain that escalates sharply after a stretch of steady improvement, rather than easing week by week

New shortness of breath or chest pain after an operation, or a hot and swollen calf, needs an emergency department the same day. Call 911 if breathing is difficult.

This article explains why outcome questionnaires are repeated and how they are read. It is general education, not medical advice, and no score replaces an assessment by the clinician looking after your recovery.

References

  1. 1.Dawson J, Fitzpatrick R, Murray D, Carr A. (1998). Questionnaire on the perceptions of patients about total knee replacement. J Bone Joint Surg Br. 1998;80-B(1):63-69. doi:10.1302/0301-620X.80B1.0800063The Oxford Knee Score's development and validation in 117 patients assessed pre-operatively and at six months after total knee replacement; its 12-item patient-completed structure; its construct validity against American Knee Society clinical scores and relevant SF-36 and Stanford HAQ sections; its sensitivity to change, with significantly greater change scores in patients self-reporting substantial improvement; its original 1-to-5-per-item scoring summing to 12-60 with lower indicating better outcome; and the fact that this paper derives no MCID and no MDC.
  2. 2.Dawson J, Fitzpatrick R, Carr A (1996). Questionnaire on the perceptions of patients about shoulder surgery. J Bone Joint Surg Br 1996;78-B(4):593-600. doi:10.1302/0301-620X.78B4.0780593The Oxford Shoulder Score's development and prospective validation in 111 patients assessed before shoulder surgery and again at six months alongside the SF-36, Stanford HAQ, and the surgeon-rated Constant score; its 12-item structure; its responsiveness, including a standardised effect size comparing favourably with the SF-36 and HAQ and the best discrimination of patients reporting their shoulder was much better; its original 12-60 scoring with 12 as the best outcome; and that it reports no MCID and no MDC.
  3. 3.Kopec JA, Esdaile JM, Abrahamowicz M, et al. (1995). The Quebec Back Pain Disability Scale. Measurement properties. Spine (Phila Pa 1976). 1995;20(3):341-52. doi:10.1097/00007632-199502000-00016The Quebec Back Pain Disability Scale as a 20-item self-administered instrument; its responsiveness — detecting significant change in disability over time and distinguishing change scores between groups expected to differ in direction of change; and the authors' recommendation of it as an outcome measure both in clinical trials and for monitoring individual patients.
  4. 4.Peul WC, van Houwelingen HC, van den Hout WB, et al. (2007). Surgery versus Prolonged Conservative Treatment for Sciatica. New England Journal of Medicine. doi:10.1056/NEJMoa064039That for sciatica caused by lumbar disc herniation, early surgery gave faster relief of leg pain than prolonged conservative care, while one-year outcomes were similar between the two strategies.
  5. 5.Weinstein JN, Tosteson TD, Lurie JD, et al. (SPORT) (2008). Surgical versus Nonsurgical Therapy for Lumbar Spinal Stenosis. New England Journal of Medicine. doi:10.1056/NEJMoa0707136That in SPORT, patients with lumbar spinal stenosis without spondylolisthesis improved more with decompressive surgery than with nonsurgical care over two years, while nonsurgical patients also improved modestly and rarely worsened.
  6. 6.Buchbinder R, van Tulder M, Öberg B, et al. (2018). Low back pain: a call for action. The Lancet. doi:10.1016/S0140-6736(18)30488-4That low-value care for low back pain — unnecessary imaging, opioids, injections, and surgery — is widespread globally and should be reduced.

6 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — citations link their sources. Editorial policy