Why the Same Questionnaire Comes Back After Surgery
SaveIt feels like paperwork and it is not. Outcome questionnaires were designed to be answered twice: the Oxford Knee Score was built by asking 117 people the same twelve questions before a knee replacement and again six months later. Without that first set of answers, the second set is a number floating free of anything to compare it against.
Last updated: July 2026
What the questionnaire before surgery is actually for
It builds the baseline, and the baseline is the whole apparatus. Nobody can interpret a post-operative score without knowing what it replaced. This is not a habit that grew up around the instruments — it is how they were constructed. The Oxford Knee Score, a 12-item questionnaire completed by the patient, was developed and validated prospectively in 117 people assessed before total knee replacement and again six months afterwards 1Ref 1Dawson J, Fitzpatrick R, Murray D, Carr A. (1998).Questionnaire on the perceptions of patients about total knee replacement.The Oxford Knee Score's development and validation in 117 patients assessed pre-operatively and at six months after total knee replacement; its 12-item patient-completed structure; its construct validity against American Knee Society clinical scores and relevant SF-36 and Stanford HAQ sections; its sensitivity to change, with significantly greater change scores in patients self-reporting substantial improvement; its original 1-to-5-per-item scoring summing to 12-60 with lower indicating better outcome; and the fact that this paper derives no MCID and no MDC..
The Oxford Shoulder Score was built the same way: 12 items, 111 patients answering before shoulder surgery and again at six months, alongside the SF-36, the Stanford health assessment questionnaire, and the surgeon-rated Constant score 2Ref 2Dawson J, Fitzpatrick R, Carr A (1996).Questionnaire on the perceptions of patients about shoulder surgery.The Oxford Shoulder Score's development and prospective validation in 111 patients assessed before shoulder surgery and again at six months alongside the SF-36, Stanford HAQ, and the surgeon-rated Constant score; its 12-item structure; its responsiveness, including a standardised effect size comparing favourably with the SF-36 and HAQ and the best discrimination of patients reporting their shoulder was much better; its original 12-60 scoring with 12 as the best outcome; and that it reports no MCID and no MDC.. The paired design is not how these questionnaires happen to be used. It is what they are.
Which is why a clinic that misses the pre-operative form has lost something it cannot get back. There is no reconstructing how bad the knee was in March once it is September and the knee is better. Memory quietly revises the past to match the present.
Why the wording never changes
Because a change score is a subtraction, and subtraction needs the same units at both ends. If the questions shift between visits, the difference between the two totals stops being about you and starts being about the form. That is why the questionnaire can feel rigid, repetitive, and occasionally beside the point: the items were fixed during development and validation, and the instrument's value lives in that fixity.
The Oxford Knee Score's construct validity was established by showing that its scores moved with American Knee Society clinical scores and with the relevant sections of the SF-36 and the Stanford health assessment questionnaire 1Ref 1Dawson J, Fitzpatrick R, Murray D, Carr A. (1998).Questionnaire on the perceptions of patients about total knee replacement.The Oxford Knee Score's development and validation in 117 patients assessed pre-operatively and at six months after total knee replacement; its 12-item patient-completed structure; its construct validity against American Knee Society clinical scores and relevant SF-36 and Stanford HAQ sections; its sensitivity to change, with significantly greater change scores in patients self-reporting substantial improvement; its original 1-to-5-per-item scoring summing to 12-60 with lower indicating better outcome; and the fact that this paper derives no MCID and no MDC.. Those relationships were demonstrated for that exact set of items in that exact wording. Rewrite one question to sound friendlier and you have a new instrument with nothing behind it.
It also explains why a form may ask about something that has nothing to do with your life. It is not asking because a clinician wondered. It is asking because the item earns its place in the total.
Why the second questionnaire comes at six months, not six days
Because the early weeks measure the operation rather than the outcome. Both Oxford instruments were validated against a six-month follow-up point 1Ref 1Dawson J, Fitzpatrick R, Murray D, Carr A. (1998).Questionnaire on the perceptions of patients about total knee replacement.The Oxford Knee Score's development and validation in 117 patients assessed pre-operatively and at six months after total knee replacement; its 12-item patient-completed structure; its construct validity against American Knee Society clinical scores and relevant SF-36 and Stanford HAQ sections; its sensitivity to change, with significantly greater change scores in patients self-reporting substantial improvement; its original 1-to-5-per-item scoring summing to 12-60 with lower indicating better outcome; and the fact that this paper derives no MCID and no MDC.2Ref 2Dawson J, Fitzpatrick R, Carr A (1996).Questionnaire on the perceptions of patients about shoulder surgery.The Oxford Shoulder Score's development and prospective validation in 111 patients assessed before shoulder surgery and again at six months alongside the SF-36, Stanford HAQ, and the surgeon-rated Constant score; its 12-item structure; its responsiveness, including a standardised effect size comparing favourably with the SF-36 and HAQ and the best discrimination of patients reporting their shoulder was much better; its original 12-60 scoring with 12 as the best outcome; and that it reports no MCID and no MDC., and that choice reflects what the first months after orthopedic surgery are actually like. The first two weeks after surgery are dominated by swelling, wound pain, and a body still clearing an anesthetic. Sleep after surgery is broken for reasons that have nothing to do with whether the procedure worked.
Scored at three weeks, almost any joint replacement looks like a disaster. The follow-up window is placed to land after that noise has cleared.
The window also depends on the operation. A carpal tunnel release recovery runs on a different clock from a rotator cuff repair recovery, and an acl surgery recovery timeline runs longer than either, with return-to-sport milestones counted in months rather than weeks. Bunion surgery recovery has its own arc again. The date on the follow-up questionnaire is chosen to sit where the real answer lives, which is rarely convenient and never arbitrary.
The trap: two versions of one score run in opposite directions
This is the most confusing thing about these forms, and it is worth knowing before you compare your number to anything you read elsewhere. The paper that defined the Oxford Knee Score scored each of its 12 items from 1 to 5 and summed them to a total between 12 and 60 — in which lower was better, and 60 meant the worst symptoms 1Ref 1Dawson J, Fitzpatrick R, Murray D, Carr A. (1998).Questionnaire on the perceptions of patients about total knee replacement.The Oxford Knee Score's development and validation in 117 patients assessed pre-operatively and at six months after total knee replacement; its 12-item patient-completed structure; its construct validity against American Knee Society clinical scores and relevant SF-36 and Stanford HAQ sections; its sensitivity to change, with significantly greater change scores in patients self-reporting substantial improvement; its original 1-to-5-per-item scoring summing to 12-60 with lower indicating better outcome; and the fact that this paper derives no MCID and no MDC..
The founding paper for the Oxford Shoulder Score did the same thing: 12 to 60, with 12 as the best possible outcome 2Ref 2Dawson J, Fitzpatrick R, Carr A (1996).Questionnaire on the perceptions of patients about shoulder surgery.The Oxford Shoulder Score's development and prospective validation in 111 patients assessed before shoulder surgery and again at six months alongside the SF-36, Stanford HAQ, and the surgeon-rated Constant score; its 12-item structure; its responsiveness, including a standardised effect size comparing favourably with the SF-36 and HAQ and the best discrimination of patients reporting their shoulder was much better; its original 12-60 scoring with 12 as the best outcome; and that it reports no MCID and no MDC.. The versions most clinics print today run from 0 to 48, and on those, higher is better. That rescoring arrived later, in a separate publication, and it deliberately reversed the direction. Two people can quote the same named instrument and mean precisely opposite things.
Before you compare your score to anything, ask which direction is good on this version. The question is not naive — the literature itself changed its mind.
What counts as a real improvement?
This is less settled than you would expect. The Oxford Knee Score paper demonstrated sensitivity to change — effect sizes comparable with the relevant SF-36 dimensions, and significantly greater change scores in patients who reported substantial improvement 1Ref 1Dawson J, Fitzpatrick R, Murray D, Carr A. (1998).Questionnaire on the perceptions of patients about total knee replacement.The Oxford Knee Score's development and validation in 117 patients assessed pre-operatively and at six months after total knee replacement; its 12-item patient-completed structure; its construct validity against American Knee Society clinical scores and relevant SF-36 and Stanford HAQ sections; its sensitivity to change, with significantly greater change scores in patients self-reporting substantial improvement; its original 1-to-5-per-item scoring summing to 12-60 with lower indicating better outcome; and the fact that this paper derives no MCID and no MDC. — but it derived no threshold from that. No minimal clinically important difference appears in it, and no minimal detectable change either 1Ref 1Dawson J, Fitzpatrick R, Murray D, Carr A. (1998).Questionnaire on the perceptions of patients about total knee replacement.The Oxford Knee Score's development and validation in 117 patients assessed pre-operatively and at six months after total knee replacement; its 12-item patient-completed structure; its construct validity against American Knee Society clinical scores and relevant SF-36 and Stanford HAQ sections; its sensitivity to change, with significantly greater change scores in patients self-reporting substantial improvement; its original 1-to-5-per-item scoring summing to 12-60 with lower indicating better outcome; and the fact that this paper derives no MCID and no MDC..
Neither does the shoulder paper, which showed a standardised effect size comparing favourably with the SF-36 and the health assessment questionnaire, and the sharpest discrimination of patients who said their shoulder was much better 2Ref 2Dawson J, Fitzpatrick R, Carr A (1996).Questionnaire on the perceptions of patients about shoulder surgery.The Oxford Shoulder Score's development and prospective validation in 111 patients assessed before shoulder surgery and again at six months alongside the SF-36, Stanford HAQ, and the surgeon-rated Constant score; its 12-item structure; its responsiveness, including a standardised effect size comparing favourably with the SF-36 and HAQ and the best discrimination of patients reporting their shoulder was much better; its original 12-60 scoring with 12 as the best outcome; and that it reports no MCID and no MDC..
So when a surgeon quotes you a number of points that counts as a meaningful gain, that figure comes from later research. It is not printed in the paper that built the score.
The property itself is real and measurable. The quebec back pain disability scale, a 20-item self-administered instrument, was shown to detect significant change in disability over time and to separate groups expected to change in opposite directions — which is why its authors recommended it for monitoring individual patients and not only for clinical trials 3Ref 3Kopec JA, Esdaile JM, Abrahamowicz M, et al. (1995).The Quebec Back Pain Disability Scale. Measurement properties.The Quebec Back Pain Disability Scale as a 20-item self-administered instrument; its responsiveness — detecting significant change in disability over time and distinguishing change scores between groups expected to differ in direction of change; and the authors' recommendation of it as an outcome measure both in clinical trials and for monitoring individual patients..
Where your answers go after the six-month form
Into the evidence that decides what the next person gets offered. The trials that shape surgical decisions are assembled out of exactly these questionnaires. When researchers compared early surgery against prolonged conservative treatment for sciatica caused by a lumbar disc herniation, early surgery gave faster relief of leg pain — but at one year, outcomes were similar between the two strategies 4Ref 4Peul WC, van Houwelingen HC, van den Hout WB, et al. (2007).Surgery versus Prolonged Conservative Treatment for Sciatica.That for sciatica caused by lumbar disc herniation, early surgery gave faster relief of leg pain than prolonged conservative care, while one-year outcomes were similar between the two strategies..
That is a finding about timing rather than about winners, and it exists only because people filled in the same form again and again. The answer is not the same for every problem, which is the entire reason for asking. In patients with lumbar spinal stenosis without spondylolisthesis, decompressive surgery produced greater improvement than nonsurgical care over two years, while the nonsurgical patients also improved modestly and rarely got worse 5Ref 5Weinstein JN, Tosteson TD, Lurie JD, et al. (SPORT) (2008).Surgical versus Nonsurgical Therapy for Lumbar Spinal Stenosis.That in SPORT, patients with lumbar spinal stenosis without spondylolisthesis improved more with decompressive surgery than with nonsurgical care over two years, while nonsurgical patients also improved modestly and rarely worsened..
Two conditions, two different answers — faster but equal at a year in one, clearly better at two years in the other. Neither conclusion could have been reached by asking people, in general, how they were getting on.
When the after score is flat, or worse
It is information, not a verdict, and it belongs in the conversation rather than in your stomach at 2am. A score that has not moved says that the thing being measured has not changed. It does not say the operation failed, that nothing further can be done, or that you are imagining the trouble. It says the next appointment has a real subject, which is better than a next appointment spent guessing.
Sometimes the explanation is that the form is measuring what you are not doing rather than what you cannot do. Fear of movement after an operation is common and it suppresses a function score exactly as a mechanical problem would. A questionnaire cannot separate those two. A clinician who examines you can.
And sometimes the honest reading is that the procedure was not the right answer to that problem. Low-value care for low back pain — unnecessary imaging, opioids, injections, and surgery — is widespread globally and needs reducing 6Ref 6Buchbinder R, van Tulder M, Öberg B, et al. (2018).Low back pain: a call for action.That low-value care for low back pain — unnecessary imaging, opioids, injections, and surgery — is widespread globally and should be reduced.. The only reason anyone knows that is that somebody measured the after.
Common questions
Related
Muscle, joint & pain
How to Fill Out a Function Questionnaire HonestlyMuscle, joint & pain
The Oxford Knee Score Before and After ReplacementMuscle, joint & pain
The Oxford Shoulder Score, Explained
Say it back
How would you explain this to someone you love?
Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.
Symptoms that outrank any questionnaire
- —A surgical wound that grows redder, hotter, or starts draining fluid, especially alongside a fever
- —Calf pain and swelling on one side, or new shortness of breath or chest pain, after lower-limb surgery
- —New numbness, or weakness that makes a limb give way, that was not present when you left the hospital
- —Pain that escalates sharply after a stretch of steady improvement, rather than easing week by week
New shortness of breath or chest pain after an operation, or a hot and swollen calf, needs an emergency department the same day. Call 911 if breathing is difficult.
This article explains why outcome questionnaires are repeated and how they are read. It is general education, not medical advice, and no score replaces an assessment by the clinician looking after your recovery.
References
- 1.Dawson J, Fitzpatrick R, Murray D, Carr A. (1998). Questionnaire on the perceptions of patients about total knee replacement. J Bone Joint Surg Br. 1998;80-B(1):63-69. doi:10.1302/0301-620X.80B1.0800063 ✓The Oxford Knee Score's development and validation in 117 patients assessed pre-operatively and at six months after total knee replacement; its 12-item patient-completed structure; its construct validity against American Knee Society clinical scores and relevant SF-36 and Stanford HAQ sections; its sensitivity to change, with significantly greater change scores in patients self-reporting substantial improvement; its original 1-to-5-per-item scoring summing to 12-60 with lower indicating better outcome; and the fact that this paper derives no MCID and no MDC.
- 2.Dawson J, Fitzpatrick R, Carr A (1996). Questionnaire on the perceptions of patients about shoulder surgery. J Bone Joint Surg Br 1996;78-B(4):593-600. doi:10.1302/0301-620X.78B4.0780593 ✓The Oxford Shoulder Score's development and prospective validation in 111 patients assessed before shoulder surgery and again at six months alongside the SF-36, Stanford HAQ, and the surgeon-rated Constant score; its 12-item structure; its responsiveness, including a standardised effect size comparing favourably with the SF-36 and HAQ and the best discrimination of patients reporting their shoulder was much better; its original 12-60 scoring with 12 as the best outcome; and that it reports no MCID and no MDC.
- 3.Kopec JA, Esdaile JM, Abrahamowicz M, et al. (1995). The Quebec Back Pain Disability Scale. Measurement properties. Spine (Phila Pa 1976). 1995;20(3):341-52. doi:10.1097/00007632-199502000-00016 ✓The Quebec Back Pain Disability Scale as a 20-item self-administered instrument; its responsiveness — detecting significant change in disability over time and distinguishing change scores between groups expected to differ in direction of change; and the authors' recommendation of it as an outcome measure both in clinical trials and for monitoring individual patients.
- 4.Peul WC, van Houwelingen HC, van den Hout WB, et al. (2007). Surgery versus Prolonged Conservative Treatment for Sciatica. New England Journal of Medicine. doi:10.1056/NEJMoa064039That for sciatica caused by lumbar disc herniation, early surgery gave faster relief of leg pain than prolonged conservative care, while one-year outcomes were similar between the two strategies.
- 5.Weinstein JN, Tosteson TD, Lurie JD, et al. (SPORT) (2008). Surgical versus Nonsurgical Therapy for Lumbar Spinal Stenosis. New England Journal of Medicine. doi:10.1056/NEJMoa0707136That in SPORT, patients with lumbar spinal stenosis without spondylolisthesis improved more with decompressive surgery than with nonsurgical care over two years, while nonsurgical patients also improved modestly and rarely worsened.
- 6.Buchbinder R, van Tulder M, Öberg B, et al. (2018). Low back pain: a call for action. The Lancet. doi:10.1016/S0140-6736(18)30488-4That low-value care for low back pain — unnecessary imaging, opioids, injections, and surgery — is widespread globally and should be reduced.
6 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — citations link their sources. Editorial policy