Muscle, joint & pain

Reading a Lower Extremity Functional Scale Score

Save

Your physical therapist wrote a number at the top of the sheet and moved on. There is no band it falls into, no line marking normal, and no way to read it against anybody else's leg. What it can do is tell you whether you have genuinely improved since the last time you filled it in, and there is a specific size that change has to reach before it counts.

Last updated: July 2026

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

What a LEFS score is a score of

It is a total of your own difficulty ratings, and nothing beyond that. The scale hands you a list of ordinary lower-limb activities — walking, stairs, squatting, getting in and out of a car, standing for an hour — and asks how much difficulty each gives you today. The ratings are summed into one total, where a higher total means more function. The maximum is printed on your clinic's scoring sheet.

So the number is a self-report, and it inherits everything about the day you reported it. It was developed for outpatients with lower-extremity musculoskeletal dysfunction 1, which is a broad population on purpose: the same scale follows an ankle sprain, a knee replacement, and a hip that has hurt for a decade.

The scale measures difficulty. It does not measure damage, and the two come apart all the time.

The nine-point threshold

The scale's development paper reports both of the figures anyone needs here, and they happen to land on the same value. The minimal clinically important difference is 9 scale points at a 90% confidence interval, and the minimal detectable change is also 9 scale points at the same confidence 1. Nine is the number to hold on to. Below it, a change cannot be told apart from noise. At or above it, the change is both real and large enough to be worth something.

Those are two distinct ideas that happen to coincide here, and the distinction is worth keeping straight.

  • Minimal detectable change is a statement about the instrument. It is the smallest change that exceeds the measurement error — the point past which you can say something moved rather than the ruler wobbling.
  • Minimal clinically important difference is a statement about people. It is the smallest change patients themselves notice as meaningful.

When the two coincide, as they do here, the instrument is only just sensitive enough for the job. There is no zone where a change is measurable but too small to care about, and no zone where a change matters but hides inside the error.

A six-point gain on the LEFS is not a small improvement. It is a result you cannot distinguish from noise 1.

Nine is also specific to this scale, which is easy to forget when instruments start piling up. The IKDC Subjective Knee Form reports a minimal detectable change of 9.0 points 2 — an identical-looking number on a different scale, measuring a different thing. The coincidence is meaningless. Thresholds never travel between instruments.

How much error sits in a single reading

The same paper puts a figure on this, and it is the most clarifying number on the page. Point-in-time measurement error is ±5.3 scale points at a 90% confidence interval 1. Which means a score of 52 is not really 52. It is a statement that the true value sits somewhere in a band roughly five points either side, and the band belongs to the instrument rather than to you.

This is not a flaw. Every self-report measure has one, and most never tell you its size. The LEFS does, and knowing it changes how you read your own sheet: two scores four points apart are the same score.

Reliability, meanwhile, is high. Test-retest reliability came out at R = .94, with a 95% lower-limit confidence interval of .89 1. A scale can be both highly reliable and carry a five-point band, and holding both facts at once is the whole skill of reading these numbers.

A score that dipped a few points since last visit has not necessarily gone anywhere. That is what the band is for.

Is there a good LEFS score?

Not in the sense people are asking. The scale carries no severity bands, no normal range, and no line above which a leg counts as recovered. That absence is deliberate rather than an oversight — it is what a change-detecting instrument is. Built to be sensitive to change, it outperformed the SF-36 physical function subscale on exactly that in the population it was developed in 1.

Plenty of respected instruments are built this way. PROMIS Physical Function reports on a standardised T-score metric normed to a mean of 50 and standard deviation of 10 in a US general population, higher meaning better function, and it is explicitly a norm-referenced continuous metric rather than a cutoff instrument with severity bands 3. Different design, same refusal to hand out grades.

Others do carry an interpretive format. The Oswestry Disability Index, for the low back, is a ten-section measure reported as a percentage from 0 to 100% 4, and a percentage invites banding in a way a raw total does not. If you have met an oswestry score, that is why it came with a label attached and your LEFS did not.

What you want to knowDoes the LEFS answer it?
Am I better than last month?Yes — if the gap clears nine points
Is my leg normal for my age?No
Am I worse than the person next to me?No, and the comparison is meaningless
Has my recovery stalled?Yes, over a run of scores

The fourth row is the one worth taking seriously, and it needs more than one visit's number to answer.

What the scale never asks

It asks how much difficulty an activity gives you. It never asks why. Whether stairs are hard because the knee will not bend, or hard because you are afraid of what bending it might do, the scale records the same answer — and those two legs need completely different treatment.

That second possibility is measurable, just not here. The Pain Catastrophizing Scale is a 13-item self-report measure, each item rated from 0 to 4, giving a total from 0 to 52 where higher means more catastrophizing, and it resolves into three factors: rumination, magnification, and helplessness 5. People scoring high on it report more negative pain-related thoughts, greater emotional distress, and greater pain intensity when pain is induced experimentally 5. It is worth knowing that its original paper set no clinical cutoff 5, which puts it in the same family as the LEFS — a measure, not a verdict.

The practical version: a LEFS that will not move, in a leg that examines well, is a reason to ask a wider question rather than to repeat the same exercises harder. Fear of movement is common after an injury, it responds to being addressed directly, and no functional score will surface it on its own.

Read the run, not the reading

One score is a dot. Three scores are a direction, and a direction is the thing anybody treating you actually acts on. A leg gaining four points a fortnight is not stalling merely because no single fortnight clears nine — over six weeks the run clears it comfortably. A leg that gained nine points early and then sat still for two months has said something that no individual score on that plateau could.

A few patterns and what they usually prompt:

  • Steady climb. The plan is working. The main risk is stopping too early, since a scale still moving is a leg still improving.
  • Early jump, then flat. Common, and not automatically bad — the easy gains arrive first. It is a reasonable moment to ask whether the exercises have got harder as the leg has.
  • Sawtooth. Up, down, up. Often the measurement band rather than the leg. Watch the trend across months.
  • Genuine decline across several visits. A fall well past nine points, sustained, is worth raising rather than pushing through.

This is also why filling it in casually costs you something real. A score answered while distracted in a waiting room becomes part of your run, and a fictional dot bends the line.

When the score and the leg disagree

Take the disagreement seriously, because it is usually the most useful thing on the sheet. A score that climbs while the leg feels no better most often means the listed activities are not the activities that matter to you. The scale asks about a generic lower limb; you live in a specific one. If your life turns on kneeling in a garden, or on carrying a toddler downstairs, a scale that never mentions either can rise while your actual week does not.

Some instruments hand that choice back to you — the patient-specific functional scale is built around activities you nominate yourself — and there is no reason both cannot be tracked at once. Saying so out loud costs nothing: the number is meant to serve the conversation rather than replace it.

The same logic runs across the rest of the family. A dash score for the arm, an ndi score for the neck, a womac score for an arthritic hip or knee — each is a total, a threshold, and a run of readings. What none of them is, is a verdict on the limb. Yours is the leg. The scale is only how it gets written down.

Common questions

Around nine scale points. The scale's development work reports a minimal clinically important difference of 9 points and a minimal detectable change of 9 points, both at 90% confidence. Below that, a change cannot be separated from measurement error. At or above it, the change is real and large enough that patients themselves notice it, which is what makes nine the number worth remembering.

Yes. The total rises as difficulty falls, so a higher number means more function. That direction is worth confirming whenever you compare instruments rather than assuming it, because knee and back questionnaires do not all run the same way — several published scales are the reverse, where a higher number means worse symptoms. Scores from different instruments cannot be compared regardless.

The scale does not define one. It carries no normal range, no severity bands, and no line above which a leg counts as recovered, because it was built to detect change rather than to grade a limb. Your own earlier scores are the only meaningful comparison. Somebody else's number tells you about their leg, their life, and their expectations, not about yours.

A drop of a few points is within the instrument's own error, which is about five points either side of any single reading. Two scores four points apart are effectively the same score. What deserves attention is a fall well past nine points that holds across several visits, or a decline alongside new symptoms, both of which are worth raising with your therapist.

Usually because the activities on the list are not the activities that matter in your life. The scale asks about a generic lower limb, and it can climb while the specific thing you care about stays impossible. It also asks how hard something is without asking why, so a leg limited by fear rather than tissue can look identical on paper.

No. It is a record of self-reported difficulty, with no threshold that points toward or away from an operation. That decision rests on the diagnosis, imaging that matches your symptoms, how the leg has responded to treatment so far, and what your life needs from it. A score contributes one honest data point to that discussion and settles none of it.

Related

Say it back

How would you explain this to someone you love?

Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.

Talk to a clinician

Gale can help you find a clinician in your state and request a visit.

Find care →

Symptoms a score will not capture

  • A leg that gives way without warning, or a foot that catches on the floor because you cannot lift it
  • New numbness in the groin or inner thighs, or loss of bladder or bowel control alongside leg symptoms
  • Calf pain, swelling, warmth, or redness in one leg, particularly with shortness of breath or chest pain
  • Sudden severe pain after a pop or snap, with an inability to bear weight on that leg

Numbness in the groin or inner thighs with loss of bladder or bowel control, or calf swelling with shortness of breath or chest pain, are emergencies — go to an emergency department or call 911 the same day rather than waiting for a scheduled appointment.

This article explains how a patient-reported outcome score is read. It is general education, not medical advice, and no questionnaire can assess your leg. Questions about your own score, your symptoms, or your treatment belong with the clinician who is treating you.

References

  1. 1.Binkley JM, Stratford PW, Lott SA, Riddle DL (1999). The Lower Extremity Functional Scale (LEFS): Scale Development, Measurement Properties, and Clinical Application. Physical Therapy, 79(4), 371-383. doi:10.1093/ptj/79.4.371The LEFS was developed for outpatients with lower-extremity musculoskeletal dysfunction, and this paper reports its core psychometrics: a minimal clinically important difference of 9 scale points (90% CI), a minimal detectable change of 9 scale points (90% CI), point-in-time measurement error of ±5.3 scale points (90% CI), test-retest reliability of R = .94 (95% lower-limit CI = .89), and sensitivity to change exceeding that of the SF-36 physical function subscale in this population. Cited for the intended population and those figures only; the article deliberately attributes no item count, total score range, or scoring direction to this source.
  2. 2.Irrgang JJ, Anderson AF, Boland AL, Harner CD, Kurosaka M, Neyret P, Richmond JC, Shelborne KD (2001). Development and Validation of the International Knee Documentation Committee Subjective Knee Form. The American Journal of Sports Medicine, 29(5), 600–613. doi:10.1177/03635465010290051301The IKDC Subjective Knee Form reports a minimal detectable change of 9.0 points on its own scale. Cited solely as a minimal detectable change for a different instrument, to show that a numerically similar threshold does not transfer between scales; no MCID, item count, score range, or scoring direction is attributed to it.
  3. 3.Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014). The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency. Journal of Clinical Epidemiology, 67(5):516-526. doi:10.1016/j.jclinepi.2013.10.024PROMIS Physical Function is reported on a standardised T-score metric normed to a mean of 50 and standard deviation of 10 in a US general population sample, with higher scores indicating better physical function, and it is a norm-referenced continuous metric rather than a cutoff instrument with severity bands. Cited for the metric, its direction, and the absence of severity bands; no MCID or MDC is attributed to it.
  4. 4.Fairbank JCT, Pynsent PB (2000). The Oswestry Disability Index. Spine. doi:10.1097/00007632-200011150-00017The Oswestry Disability Index is a validated ten-section patient-reported measure of low-back-pain-related disability, scored as a percentage from 0 to 100%. Cited for the instrument's structure and its percentage reporting format, as a contrast to the LEFS's raw total.
  5. 5.Sullivan MJL, Bishop SR, Pivik J (1995). The Pain Catastrophizing Scale: Development and validation. Psychological Assessment 1995;7(4):524-532. doi:10.1037/1040-3590.7.4.524The Pain Catastrophizing Scale is a 13-item self-report measure with items rated 0 to 4 for a total range of 0–52, higher scores indicating greater catastrophizing, resolving into three factors (rumination, magnification, helplessness); high scorers report more negative pain-related thoughts, greater emotional distress, and greater pain intensity under a cold-pressor procedure; and the originating paper establishes no clinical cutoff. Cited for the instrument's construction, its construct-validity findings, and the absence of a cutoff in the original.

5 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — every citation independently verified. Editorial policy