Reading a Lower Extremity Functional Scale Score
SaveYour physical therapist wrote a number at the top of the sheet and moved on. There is no band it falls into, no line marking normal, and no way to read it against anybody else's leg. What it can do is tell you whether you have genuinely improved since the last time you filled it in, and there is a specific size that change has to reach before it counts.
Last updated: July 2026
What a LEFS score is a score of
It is a total of your own difficulty ratings, and nothing beyond that. The scale hands you a list of ordinary lower-limb activities — walking, stairs, squatting, getting in and out of a car, standing for an hour — and asks how much difficulty each gives you today. The ratings are summed into one total, where a higher total means more function. The maximum is printed on your clinic's scoring sheet.
So the number is a self-report, and it inherits everything about the day you reported it. It was developed for outpatients with lower-extremity musculoskeletal dysfunction 1Ref 1Binkley JM, Stratford PW, Lott SA, Riddle DL (1999).The Lower Extremity Functional Scale (LEFS): Scale Development, Measurement Properties, and Clinical Application.The LEFS was developed for outpatients with lower-extremity musculoskeletal dysfunction, and this paper reports its core psychometrics: a minimal clinically important difference of 9 scale points (90% CI), a minimal detectable change of 9 scale points (90% CI), point-in-time measurement error of ±5.3 scale points (90% CI), test-retest reliability of R = .94 (95% lower-limit CI = .89), and sensitivity to change exceeding that of the SF-36 physical function subscale in this population. Cited for the intended population and those figures only; the article deliberately attributes no item count, total score range, or scoring direction to this source., which is a broad population on purpose: the same scale follows an ankle sprain, a knee replacement, and a hip that has hurt for a decade.
The scale measures difficulty. It does not measure damage, and the two come apart all the time.
The nine-point threshold
The scale's development paper reports both of the figures anyone needs here, and they happen to land on the same value. The minimal clinically important difference is 9 scale points at a 90% confidence interval, and the minimal detectable change is also 9 scale points at the same confidence 1Ref 1Binkley JM, Stratford PW, Lott SA, Riddle DL (1999).The Lower Extremity Functional Scale (LEFS): Scale Development, Measurement Properties, and Clinical Application.The LEFS was developed for outpatients with lower-extremity musculoskeletal dysfunction, and this paper reports its core psychometrics: a minimal clinically important difference of 9 scale points (90% CI), a minimal detectable change of 9 scale points (90% CI), point-in-time measurement error of ±5.3 scale points (90% CI), test-retest reliability of R = .94 (95% lower-limit CI = .89), and sensitivity to change exceeding that of the SF-36 physical function subscale in this population. Cited for the intended population and those figures only; the article deliberately attributes no item count, total score range, or scoring direction to this source.. Nine is the number to hold on to. Below it, a change cannot be told apart from noise. At or above it, the change is both real and large enough to be worth something.
Those are two distinct ideas that happen to coincide here, and the distinction is worth keeping straight.
- Minimal detectable change is a statement about the instrument. It is the smallest change that exceeds the measurement error — the point past which you can say something moved rather than the ruler wobbling.
- Minimal clinically important difference is a statement about people. It is the smallest change patients themselves notice as meaningful.
When the two coincide, as they do here, the instrument is only just sensitive enough for the job. There is no zone where a change is measurable but too small to care about, and no zone where a change matters but hides inside the error.
A six-point gain on the LEFS is not a small improvement. It is a result you cannot distinguish from noise 1Ref 1Binkley JM, Stratford PW, Lott SA, Riddle DL (1999).The Lower Extremity Functional Scale (LEFS): Scale Development, Measurement Properties, and Clinical Application.The LEFS was developed for outpatients with lower-extremity musculoskeletal dysfunction, and this paper reports its core psychometrics: a minimal clinically important difference of 9 scale points (90% CI), a minimal detectable change of 9 scale points (90% CI), point-in-time measurement error of ±5.3 scale points (90% CI), test-retest reliability of R = .94 (95% lower-limit CI = .89), and sensitivity to change exceeding that of the SF-36 physical function subscale in this population. Cited for the intended population and those figures only; the article deliberately attributes no item count, total score range, or scoring direction to this source..
Nine is also specific to this scale, which is easy to forget when instruments start piling up. The IKDC Subjective Knee Form reports a minimal detectable change of 9.0 points 2Ref 2Irrgang JJ, Anderson AF, Boland AL, Harner CD, Kurosaka M, Neyret P, Richmond JC, Shelborne KD (2001).Development and Validation of the International Knee Documentation Committee Subjective Knee Form.The IKDC Subjective Knee Form reports a minimal detectable change of 9.0 points on its own scale. Cited solely as a minimal detectable change for a different instrument, to show that a numerically similar threshold does not transfer between scales; no MCID, item count, score range, or scoring direction is attributed to it. — an identical-looking number on a different scale, measuring a different thing. The coincidence is meaningless. Thresholds never travel between instruments.
How much error sits in a single reading
The same paper puts a figure on this, and it is the most clarifying number on the page. Point-in-time measurement error is ±5.3 scale points at a 90% confidence interval 1Ref 1Binkley JM, Stratford PW, Lott SA, Riddle DL (1999).The Lower Extremity Functional Scale (LEFS): Scale Development, Measurement Properties, and Clinical Application.The LEFS was developed for outpatients with lower-extremity musculoskeletal dysfunction, and this paper reports its core psychometrics: a minimal clinically important difference of 9 scale points (90% CI), a minimal detectable change of 9 scale points (90% CI), point-in-time measurement error of ±5.3 scale points (90% CI), test-retest reliability of R = .94 (95% lower-limit CI = .89), and sensitivity to change exceeding that of the SF-36 physical function subscale in this population. Cited for the intended population and those figures only; the article deliberately attributes no item count, total score range, or scoring direction to this source.. Which means a score of 52 is not really 52. It is a statement that the true value sits somewhere in a band roughly five points either side, and the band belongs to the instrument rather than to you.
This is not a flaw. Every self-report measure has one, and most never tell you its size. The LEFS does, and knowing it changes how you read your own sheet: two scores four points apart are the same score.
Reliability, meanwhile, is high. Test-retest reliability came out at R = .94, with a 95% lower-limit confidence interval of .89 1Ref 1Binkley JM, Stratford PW, Lott SA, Riddle DL (1999).The Lower Extremity Functional Scale (LEFS): Scale Development, Measurement Properties, and Clinical Application.The LEFS was developed for outpatients with lower-extremity musculoskeletal dysfunction, and this paper reports its core psychometrics: a minimal clinically important difference of 9 scale points (90% CI), a minimal detectable change of 9 scale points (90% CI), point-in-time measurement error of ±5.3 scale points (90% CI), test-retest reliability of R = .94 (95% lower-limit CI = .89), and sensitivity to change exceeding that of the SF-36 physical function subscale in this population. Cited for the intended population and those figures only; the article deliberately attributes no item count, total score range, or scoring direction to this source.. A scale can be both highly reliable and carry a five-point band, and holding both facts at once is the whole skill of reading these numbers.
A score that dipped a few points since last visit has not necessarily gone anywhere. That is what the band is for.
Is there a good LEFS score?
Not in the sense people are asking. The scale carries no severity bands, no normal range, and no line above which a leg counts as recovered. That absence is deliberate rather than an oversight — it is what a change-detecting instrument is. Built to be sensitive to change, it outperformed the SF-36 physical function subscale on exactly that in the population it was developed in 1Ref 1Binkley JM, Stratford PW, Lott SA, Riddle DL (1999).The Lower Extremity Functional Scale (LEFS): Scale Development, Measurement Properties, and Clinical Application.The LEFS was developed for outpatients with lower-extremity musculoskeletal dysfunction, and this paper reports its core psychometrics: a minimal clinically important difference of 9 scale points (90% CI), a minimal detectable change of 9 scale points (90% CI), point-in-time measurement error of ±5.3 scale points (90% CI), test-retest reliability of R = .94 (95% lower-limit CI = .89), and sensitivity to change exceeding that of the SF-36 physical function subscale in this population. Cited for the intended population and those figures only; the article deliberately attributes no item count, total score range, or scoring direction to this source..
Plenty of respected instruments are built this way. PROMIS Physical Function reports on a standardised T-score metric normed to a mean of 50 and standard deviation of 10 in a US general population, higher meaning better function, and it is explicitly a norm-referenced continuous metric rather than a cutoff instrument with severity bands 3Ref 3Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014).The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency.PROMIS Physical Function is reported on a standardised T-score metric normed to a mean of 50 and standard deviation of 10 in a US general population sample, with higher scores indicating better physical function, and it is a norm-referenced continuous metric rather than a cutoff instrument with severity bands. Cited for the metric, its direction, and the absence of severity bands; no MCID or MDC is attributed to it.. Different design, same refusal to hand out grades.
Others do carry an interpretive format. The Oswestry Disability Index, for the low back, is a ten-section measure reported as a percentage from 0 to 100% 4Ref 4Fairbank JCT, Pynsent PB (2000).The Oswestry Disability Index.The Oswestry Disability Index is a validated ten-section patient-reported measure of low-back-pain-related disability, scored as a percentage from 0 to 100%. Cited for the instrument's structure and its percentage reporting format, as a contrast to the LEFS's raw total., and a percentage invites banding in a way a raw total does not. If you have met an oswestry score, that is why it came with a label attached and your LEFS did not.
| What you want to know | Does the LEFS answer it? |
|---|---|
| Am I better than last month? | Yes — if the gap clears nine points |
| Is my leg normal for my age? | No |
| Am I worse than the person next to me? | No, and the comparison is meaningless |
| Has my recovery stalled? | Yes, over a run of scores |
The fourth row is the one worth taking seriously, and it needs more than one visit's number to answer.
What the scale never asks
It asks how much difficulty an activity gives you. It never asks why. Whether stairs are hard because the knee will not bend, or hard because you are afraid of what bending it might do, the scale records the same answer — and those two legs need completely different treatment.
That second possibility is measurable, just not here. The Pain Catastrophizing Scale is a 13-item self-report measure, each item rated from 0 to 4, giving a total from 0 to 52 where higher means more catastrophizing, and it resolves into three factors: rumination, magnification, and helplessness 5Ref 5Sullivan MJL, Bishop SR, Pivik J (1995).The Pain Catastrophizing Scale: Development and validation.The Pain Catastrophizing Scale is a 13-item self-report measure with items rated 0 to 4 for a total range of 0–52, higher scores indicating greater catastrophizing, resolving into three factors (rumination, magnification, helplessness); high scorers report more negative pain-related thoughts, greater emotional distress, and greater pain intensity under a cold-pressor procedure; and the originating paper establishes no clinical cutoff. Cited for the instrument's construction, its construct-validity findings, and the absence of a cutoff in the original.. People scoring high on it report more negative pain-related thoughts, greater emotional distress, and greater pain intensity when pain is induced experimentally 5Ref 5Sullivan MJL, Bishop SR, Pivik J (1995).The Pain Catastrophizing Scale: Development and validation.The Pain Catastrophizing Scale is a 13-item self-report measure with items rated 0 to 4 for a total range of 0–52, higher scores indicating greater catastrophizing, resolving into three factors (rumination, magnification, helplessness); high scorers report more negative pain-related thoughts, greater emotional distress, and greater pain intensity under a cold-pressor procedure; and the originating paper establishes no clinical cutoff. Cited for the instrument's construction, its construct-validity findings, and the absence of a cutoff in the original.. It is worth knowing that its original paper set no clinical cutoff 5Ref 5Sullivan MJL, Bishop SR, Pivik J (1995).The Pain Catastrophizing Scale: Development and validation.The Pain Catastrophizing Scale is a 13-item self-report measure with items rated 0 to 4 for a total range of 0–52, higher scores indicating greater catastrophizing, resolving into three factors (rumination, magnification, helplessness); high scorers report more negative pain-related thoughts, greater emotional distress, and greater pain intensity under a cold-pressor procedure; and the originating paper establishes no clinical cutoff. Cited for the instrument's construction, its construct-validity findings, and the absence of a cutoff in the original., which puts it in the same family as the LEFS — a measure, not a verdict.
The practical version: a LEFS that will not move, in a leg that examines well, is a reason to ask a wider question rather than to repeat the same exercises harder. Fear of movement is common after an injury, it responds to being addressed directly, and no functional score will surface it on its own.
Read the run, not the reading
One score is a dot. Three scores are a direction, and a direction is the thing anybody treating you actually acts on. A leg gaining four points a fortnight is not stalling merely because no single fortnight clears nine — over six weeks the run clears it comfortably. A leg that gained nine points early and then sat still for two months has said something that no individual score on that plateau could.
A few patterns and what they usually prompt:
- Steady climb. The plan is working. The main risk is stopping too early, since a scale still moving is a leg still improving.
- Early jump, then flat. Common, and not automatically bad — the easy gains arrive first. It is a reasonable moment to ask whether the exercises have got harder as the leg has.
- Sawtooth. Up, down, up. Often the measurement band rather than the leg. Watch the trend across months.
- Genuine decline across several visits. A fall well past nine points, sustained, is worth raising rather than pushing through.
This is also why filling it in casually costs you something real. A score answered while distracted in a waiting room becomes part of your run, and a fictional dot bends the line.
When the score and the leg disagree
Take the disagreement seriously, because it is usually the most useful thing on the sheet. A score that climbs while the leg feels no better most often means the listed activities are not the activities that matter to you. The scale asks about a generic lower limb; you live in a specific one. If your life turns on kneeling in a garden, or on carrying a toddler downstairs, a scale that never mentions either can rise while your actual week does not.
Some instruments hand that choice back to you — the patient-specific functional scale is built around activities you nominate yourself — and there is no reason both cannot be tracked at once. Saying so out loud costs nothing: the number is meant to serve the conversation rather than replace it.
The same logic runs across the rest of the family. A dash score for the arm, an ndi score for the neck, a womac score for an arthritic hip or knee — each is a total, a threshold, and a run of readings. What none of them is, is a verdict on the limb. Yours is the leg. The scale is only how it gets written down.
Common questions
Related
Muscle, joint & pain
The Patient-Specific Functional Scale: Define Your Own GoalsMuscle, joint & pain
The Lower Extremity Functional Scale, ExplainedMuscle, joint & pain
The Foot and Ankle Ability Measure, in Plain Language
Say it back
How would you explain this to someone you love?
Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.
Symptoms a score will not capture
- —A leg that gives way without warning, or a foot that catches on the floor because you cannot lift it
- —New numbness in the groin or inner thighs, or loss of bladder or bowel control alongside leg symptoms
- —Calf pain, swelling, warmth, or redness in one leg, particularly with shortness of breath or chest pain
- —Sudden severe pain after a pop or snap, with an inability to bear weight on that leg
Numbness in the groin or inner thighs with loss of bladder or bowel control, or calf swelling with shortness of breath or chest pain, are emergencies — go to an emergency department or call 911 the same day rather than waiting for a scheduled appointment.
This article explains how a patient-reported outcome score is read. It is general education, not medical advice, and no questionnaire can assess your leg. Questions about your own score, your symptoms, or your treatment belong with the clinician who is treating you.
References
- 1.Binkley JM, Stratford PW, Lott SA, Riddle DL (1999). The Lower Extremity Functional Scale (LEFS): Scale Development, Measurement Properties, and Clinical Application. Physical Therapy, 79(4), 371-383. doi:10.1093/ptj/79.4.371 ✓The LEFS was developed for outpatients with lower-extremity musculoskeletal dysfunction, and this paper reports its core psychometrics: a minimal clinically important difference of 9 scale points (90% CI), a minimal detectable change of 9 scale points (90% CI), point-in-time measurement error of ±5.3 scale points (90% CI), test-retest reliability of R = .94 (95% lower-limit CI = .89), and sensitivity to change exceeding that of the SF-36 physical function subscale in this population. Cited for the intended population and those figures only; the article deliberately attributes no item count, total score range, or scoring direction to this source.
- 2.Irrgang JJ, Anderson AF, Boland AL, Harner CD, Kurosaka M, Neyret P, Richmond JC, Shelborne KD (2001). Development and Validation of the International Knee Documentation Committee Subjective Knee Form. The American Journal of Sports Medicine, 29(5), 600–613. doi:10.1177/03635465010290051301 ✓The IKDC Subjective Knee Form reports a minimal detectable change of 9.0 points on its own scale. Cited solely as a minimal detectable change for a different instrument, to show that a numerically similar threshold does not transfer between scales; no MCID, item count, score range, or scoring direction is attributed to it.
- 3.Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014). The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency. Journal of Clinical Epidemiology, 67(5):516-526. doi:10.1016/j.jclinepi.2013.10.024 ✓PROMIS Physical Function is reported on a standardised T-score metric normed to a mean of 50 and standard deviation of 10 in a US general population sample, with higher scores indicating better physical function, and it is a norm-referenced continuous metric rather than a cutoff instrument with severity bands. Cited for the metric, its direction, and the absence of severity bands; no MCID or MDC is attributed to it.
- 4.Fairbank JCT, Pynsent PB (2000). The Oswestry Disability Index. Spine. doi:10.1097/00007632-200011150-00017 ✓The Oswestry Disability Index is a validated ten-section patient-reported measure of low-back-pain-related disability, scored as a percentage from 0 to 100%. Cited for the instrument's structure and its percentage reporting format, as a contrast to the LEFS's raw total.
- 5.Sullivan MJL, Bishop SR, Pivik J (1995). The Pain Catastrophizing Scale: Development and validation. Psychological Assessment 1995;7(4):524-532. doi:10.1037/1040-3590.7.4.524 ✓The Pain Catastrophizing Scale is a 13-item self-report measure with items rated 0 to 4 for a total range of 0–52, higher scores indicating greater catastrophizing, resolving into three factors (rumination, magnification, helplessness); high scorers report more negative pain-related thoughts, greater emotional distress, and greater pain intensity under a cold-pressor procedure; and the originating paper establishes no clinical cutoff. Cited for the instrument's construction, its construct-validity findings, and the absence of a cutoff in the original.
5 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — every citation independently verified. Editorial policy