Muscle, joint & pain

The Oxford Hip Score, in Plain Language

Save

Twelve questions about a hip, answered by its owner, adding to one total — and no examination anywhere in it. That was a deliberate break from the older way of grading a hip, in which a clinician measured the joint and assigned the verdict. Here is what the Oxford questionnaire asks, why its two scoring schemes contradict each other, and how it sits next to the Harris Hip Score and the HOOS.

Last updated: July 2026

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

What the Oxford Hip Score asks, and who answers it

Twelve items, covering hip pain and hip function, combined into one composite total — and the person supplying every answer is the patient 1. No clinician examines the joint, measures how far it turns, or reads a film. Dawson, Fitzpatrick, Carr and Murray published it in 1996 with one purpose: recording what total hip replacement does for patients 1.

To show twelve questions could carry that weight, the Oxford group tested them on 220 people booked for a new hip, capturing each person twice — once while still waiting, once half a year after the operation 1. They set the results against three yardsticks already trusted at the time: the Charnley hip score, which a surgeon completes, plus the Arthritis Impact Measurement Scales and the SF-36 1. The questionnaire held together internally, reproduced itself acceptably on retest, tracked those established measures in the directions you would predict, and registered clinically important change with effect sizes that stood up well beside them 1.

The Oxford Hip Score is one person's report on one hip, turned into a figure a clinic can set against the same person's figure later.

Notice what is absent. No imaging, no examination — so it cannot separate osteoarthritis from a labral tear from pain referred down from a back, and it was never built to. It records how much a hip hurts and how much it interferes, then stops.

Two scoring schemes, and they disagree

The same twelve answers can be added two different ways, and the two point in opposite directions. Under the older arithmetic each item earns one to five points, giving a total between 12 and 60, where 12 is the healthiest hip and 60 the most troubled. Under the newer arithmetic each item earns zero to four, the total lands between 0 and 48, and now 48 is the healthiest hip.

Nothing about the questions changed; only the counting did. Most paperwork issued today uses the 0-to-48 arithmetic, but most is not yours, and the 1996 publication does not settle it — that paper establishes the twelve items and their measurement properties, not a canonical range 1.

A hip described identically on both schemes yields two totals that add to 60. The conversion is subtraction. The confusion is not.

So a bare figure is not information yet. It becomes information when two facts travel with it: the largest total the form allows, and which end of that span is the good end. Usually both are printed on the sheet. When they are not, the person who handed it over knows, and asking costs nothing.

  • Maximum 48 → the bigger your total, the better the hip.
  • Maximum 60 → the smaller your total, the better the hip.
  • Neither stated → the figure cannot honestly be read yet.

Your account of the hip, or the surgeon's: the Harris Hip Score

The Harris Hip Score is the other name on a hip clinic's paperwork, and the difference between the two instruments is who holds the pen. Harris introduced his in 1969, inside a study of 39 mold arthroplasties at Massachusetts General Hospital, describing it in the paper's own subtitle as a new method of result evaluation 2. A clinician completes it, and completing it means examining you.

It runs 0 to 100 and rises with better hip status, so a large Harris figure is unambiguously a good one — no second scheme, no inversion 2. The hundred points are parcelled across four domains, and the weighting is the interesting part 2:

DomainPointsWho supplies it
Pain44You describe it
Functional capacity47Largely you describe it
Range of motion5A clinician measures it
Absence of deformity4A clinician assesses it

Ninety-one of the hundred ride on pain and function, which tells you what Harris thought a hip was for 2. Only nine depend on the examination — but those nine are why the harris hip score cannot be completed at your kitchen table and the Oxford can. Both instruments are doing hip function measurement; they simply do it from opposite sides of the examination table.

One caution travels less widely than it should. The excellent/good/fair/poor bands routinely quoted for the Harris are not established by the 1969 paper, and neither is any threshold for meaningful change 2. That paper introduced the instrument inside a clinical series, decades before modern outcome-measure psychometrics — a derivation in context, not a validation study 2.

One total, or five subscales: the Oxford next to the HOOS

The HOOS answers the question the Oxford deliberately declines: which part got better? Where the Oxford compresses hip pain and hip function into a single figure 1, the HOOS hip score keeps five results apart — pain, symptoms, activities of daily living, sport and recreation, and hip-related quality of life 3. Same joint, same patient-completed design, finer resolution.

That resolution buys something concrete. One Oxford total cannot reveal that your pain settled while your sport and recreation stayed exactly where it was; five subscales can. The HOOS validation, carried out in total hip replacement, also found it more responsive than the WOMAC on the pain and symptom subscales — it picked up movement there that the older instrument partly missed 3.

None of which makes the HOOS the better questionnaire. Compression is a feature when a clinic needs a figure it can collect from everyone at every visit without friction, and five subscales are five times the paperwork. These instruments answer different questions, and a clinic picks the one matched to the question it is asking. If both have arrived at different appointments, you have not been handed a duplicate.

How much does the total have to move before it counts?

The paper that created the Oxford Hip Score does not say, and that absence is worth understanding rather than working around. What the 1996 study demonstrates is that the score is sensitive to clinically important change, shown through effect sizes measured across those 220 patients and compared against the Arthritis Impact Measurement Scales and the SF-36 1. That is a statement about how the instrument behaves in a population.

A population statement and a personal threshold are not the same object. The founding study derives no cut-off, reports no minimal clinically important difference, and reports no minimal detectable change 1. Every specific figure quoted as the threshold for this instrument was worked out afterwards, by other researchers, in populations and by methods that vary — which is how two clinics can quote two different numbers and neither be wrong.

A total that has barely shifted is a piece of evidence, not a verdict on your recovery or on how hard you tried.

What survives the caution is the comparison the instrument was built for: your total against your own earlier total, on the same scheme. And the most useful reading of a series is rarely arithmetic. A few points mean little alone. A few points arriving alongside sleeping through the night, or putting a sock on without planning the manoeuvre, mean something the figure cannot express by itself.

Does a low total mean you need a hip replacement?

No, and no total triggers an operation. The instrument carries no threshold that would, and the study behind it never set out to select candidates — it set out to measure results after the fact 1. Hip replacement indications rest on the imaging, on what has been tried and for how long, on how the pain behaves overnight, on your other conditions, and on what you want the hip for.

Orthopaedic guidance treats hip osteoarthritis as a sequence rather than a fork. The American Academy of Orthopaedic Surgeons' hip osteoarthritis guideline covers the non-surgical measures — exercise and physical therapy, anti-inflammatory medication — and it covers the surgical options, because both sit on one pathway at different points along it 4. A repeated Oxford total is one way of seeing where on that pathway a person has arrived, which beats reconstructing it from memory. Memory edits pain, in both directions.

The sequence framing has to cut both ways to stay honest. There are hips for which replacement is plainly the right operation, and postponing it in the name of exhausting every alternative first has a cost, paid in years of a life narrowed around a joint. A total that has drifted steadily the wrong way across three appointments, despite genuine and sustained non-surgical treatment, is not an argument for waiting longer. It is precisely the evidence the conversation needs.

The questionnaire is a measuring instrument, and measuring instruments hold no opinion about surgery — including none about yours.

What the knee and shoulder versions reveal about the scoring

The Oxford Hip Score has two siblings built to the same blueprint, and reading across to them settles something the hip's own founding paper leaves open. The oxford knee score arrived in 1998 for total knee replacement 5. The oxford shoulder score arrived in 1996 for shoulder surgery, though not for shoulder instability, which has its own separate instrument 6. Each is twelve patient-completed items producing one summed total.

The family's original arithmetic is documented, and it is the counterintuitive one. The knee score's founding paper defines its scoring outright: each of the twelve items scores 1 to 5, the items sum to a total between 12 and 60, and on that scale lower is better while 60 is the worst symptoms 5. The shoulder paper does the same, running 12 to 60 with 12 marking the best outcome, each item counting upward from fewest symptoms 6. Two instruments, two founding papers, one convention — and it is not the convention most people assume on seeing a two-digit hip score written down. The 0-to-48 recoding that reversed the direction came afterwards, from the same Oxford group, and it is the one most clinics now use. Neither founding paper tells you which is on your sheet.

The knee paper is also explicit about a limit that gets ignored. It was built for knee arthroplasty outcome measures, validated in 117 patients before a total knee replacement and again six months later against American Knee Society clinical scores and relevant sections of the SF-36 and the Stanford Health Assessment Questionnaire 5. It was not built for knee pain in general, and using it outside that population asks more of it than its evidence supports 5.

That limit applies to the hip version too. The Oxford Hip Score was validated in people having a total hip replacement 1 — handed to you for a hip that has never been near an operating theatre, it still records something real, but it is doing a job its founding evidence does not cover.

Common questions

Whichever the form says. Two schemes are in use. Where the maximum possible total is 48, bigger totals describe a better hip. Where the maximum is 60, smaller totals do, and 12 is the best result available. The twelve questions are identical either way — only the counting differs. If the sheet does not state its maximum, the clinic that issued it can.

There is no official band. The 1996 publication that defined the instrument set no severity categories and no threshold dividing a good hip from a poor one, and the grading schemes quoted around the internet come from later work using different populations and methods. Your own total, against your own earlier total, on the same scheme, is the comparison that actually carries meaning — which is the comparison the score was designed for.

No, and the difference is who fills it in. The Oxford is twelve questions you answer yourself, with nothing measured and nobody examining the joint. The Harris is completed by a clinician and includes range of motion and deformity, which need hands on your hip. The Harris runs 0 to 100 with bigger being better. The two are not convertible, and a figure from one says nothing about a figure from the other.

Resolution. The Oxford folds hip pain and hip function into one total, which is quick to collect and quick to repeat. The HOOS reports five separate subscales — pain, symptoms, daily activities, sport and recreation, and hip-related quality of life — so it can show pain improving while sport stays stuck. One number is easier to track. Five numbers show the shape of the change rather than only its size.

It can be filled in for any hip, and physiotherapists sometimes do. Whether it means as much is a separate question. Its founding study validated it in people having a total hip replacement, which is the population where the evidence for what it measures actually sits. Used outside that group it still records something real about pain and difficulty — it is just standing on thinner ground, and the total is worth holding loosely.

The score's only real job is tracking change over time, which means an honest baseline matters more than a flattering one. Answering optimistically is a common instinct — most people want to be a good patient — and it costs you: a flattered first score makes a genuine improvement look smaller than it was. Many people find it helps to answer as if describing a typical recent stretch to someone who wants an accurate picture rather than a brave one.

Related

Say it back

How would you explain this to someone you love?

Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

When a hip is a same-day problem, not a questionnaire problem

  • Hip or groin pain that begins suddenly after a fall, especially if the leg will not take your weight, looks shorter than the other, or has turned outward
  • A hip that becomes hot, swollen and severely painful along with fever or chills, particularly in the weeks after a hip operation or a joint injection
  • Hip or back pain arriving with new numbness between the legs, new weakness in the leg, or a change in bladder or bowel control
  • Deep hip pain that wakes you every night and eases in no position, alongside unexplained weight loss or a past cancer diagnosis

A hip that cannot bear weight after a fall, or a hot and feverish joint, belongs in an emergency department the same day rather than on a questionnaire. New numbness between the legs, or loss of bladder or bowel control alongside leg weakness, is a 911 or emergency-department situation now rather than a next-appointment one.

The Oxford Hip Score is a measurement instrument used in research and clinics. It is not a diagnostic test and not a substitute for assessment. This page explains what the questionnaire measures and how its scoring works; it does not interpret any individual's score. Gale's health library is educational and does not replace advice from a clinician who knows your hip and your history.

References

  1. 1.Dawson J, Fitzpatrick R, Carr A, Murray D. (1996). Questionnaire on the perceptions of patients about total hip replacement. J Bone Joint Surg Br. 1996;78(2):185-90. doi:10.1302/0301-620X.78B2.0780185The Oxford Hip Score's construction and founding psychometrics: a 12-item patient-completed questionnaire covering hip pain and function in a single composite score, developed and validated prospectively in 220 patients assessed pre-operatively and at six-month follow-up against the SF36, the Arthritis Impact Measurement Scales, and the surgeon-assessed Charnley hip score; high internal consistency; satisfactory test-retest reliability; construct validity via correlations with those measures; and sensitivity to clinically important change established by effect size. Also cited for what this paper does NOT establish: no MCID, no MDC, no severity threshold or cut-off, and no canonical score range or scoring direction.
  2. 2.Harris WH (1969). Traumatic arthritis of the hip after dislocation and acetabular fractures: treatment by mold arthroplasty. An end-result study using a new method of result evaluation. J Bone Joint Surg Am. doi:10.2106/00004623-196951040-00012The Harris Hip Score's origin, structure and direction: introduced in 1969 as the 'new method of result evaluation' named in the paper's subtitle, embedded in an end-result study of 39 mold arthroplasties at Massachusetts General Hospital; clinician-administered; scored 0-100 with higher totals indicating better hip status; the 100 points divided across pain (44), functional capacity (47), range of motion (5) and absence of deformity (4). Also cited for its limits: this is a derivation-in-context paper rather than a validation study, and it establishes neither the commonly quoted excellent/good/fair/poor category thresholds nor any MCID or MDC.
  3. 3.Nilsdotter AK, Lohmander LS, Klässbo M, Roos EM (2003). Hip disability and osteoarthritis outcome score (HOOS)—validity and responsiveness in total hip replacement. BMC Musculoskeletal Disorders. PMID 12777182The HOOS as a validated patient-reported outcome measure for hip osteoarthritis and total hip replacement, structured as five separate subscales — pain, symptoms, activities of daily living, sport and recreation, and hip-related quality of life — and more responsive than the WOMAC on the pain and symptom subscales.
  4. 4.American Academy of Orthopaedic Surgeons (AAOS) (2023). Management of Osteoarthritis of the Hip — Clinical Practice Guideline. AAOS. linkThat orthopaedic-society guidance for the management of hip osteoarthritis covers both non-surgical measures (exercise and physical therapy, anti-inflammatory medication) and surgical options, supporting the framing of hip osteoarthritis care as a sequence along one pathway rather than a binary choice between conservative care and surgery.
  5. 5.Dawson J, Fitzpatrick R, Murray D, Carr A. (1998). Questionnaire on the perceptions of patients about total knee replacement. J Bone Joint Surg Br. 1998;80-B(1):63-69. doi:10.1302/0301-620X.80B1.0800063The Oxford Knee Score's 12-item patient-completed structure yielding a single summed score, and its validation in 117 patients assessed pre-operatively and at six months after total knee replacement against American Knee Society clinical scores and relevant SF-36 and Stanford HAQ sections. Also cited for the original scoring this paper defines — each of the 12 items scored 1-5, summing to 12-60, with lower = better and 60 = worst symptoms — which documents the Oxford family's founding scoring convention; and for the paper's population limit, that the instrument was developed and validated specifically for total knee replacement rather than general knee pain.
  6. 6.Dawson J, Fitzpatrick R, Carr A (1996). Questionnaire on the perceptions of patients about shoulder surgery. J Bone Joint Surg Br 1996;78-B(4):593-600. doi:10.1302/0301-620X.78B4.0780593The Oxford Shoulder Score's 12-item structure and its intended population — outcomes of shoulder surgery, excluding shoulder stabilisation, which is covered by a separate Oxford instrument — established prospectively in 111 patients assessed before surgery and again at six months. Also cited for the original 1996 scoring this paper defines: 12 to 60, with 12 marking the best outcome and each item scored 1-5 from least symptoms upward, documenting the Oxford family's founding scoring convention.

6 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — every citation independently verified. Editorial policy