PROMIS Physical Function, the Universal Yardstick
SaveMost joint questionnaires score you against the form itself: against the worst thing that form can describe. PROMIS Physical Function does something different. It scores you against the general population, on a metric built to carry the same meaning whether the trouble is a knee, a shoulder, or a lung. That design choice explains almost everything about how the number reads.
Last updated: July 2026
What does a PROMIS Physical Function score mean?
It is a position on a distribution. The metric is standardized so that 50 is the mean of a US general-population sample and 10 points is one standard deviation, with higher scores indicating better physical function; the score is continuous and is conventionally read across roughly four standard deviations of range 1Ref 1Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014).The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency.The defining calibration of the adult PROMIS Physical Function item bank: the 124-item bank spanning upper-extremity, central and lower-extremity function plus instrumental activities of daily living; the standardized T-score metric normed to mean 50 and SD 10 in a US general population sample with higher scores indicating better function, read across roughly four standard deviations; the development sample of 149 candidate items tested with 10 SF-36 and 20 HAQ-DI items in 16,065 adults oversampling the chronically ill; the simulated 10-item CAT eliminating floor effects, reducing ceiling effects, achieving higher precision than any comparable-length static tool and discriminating age and disease groups better than legacy instruments; and the absence of any MCID, MDC, cutoff or severity band in this source..
So 40 sits one standard deviation below the general population, 60 sits one above, and the working span most people fall in runs from about 30 to about 70. There is no maximum you are climbing toward and no percentage of a perfect knee anywhere in it.
50 is not a passing grade. It is average — the middle of the general population, including everyone in that population who is already unwell.
That is why the number is written without units, and why a clinician reads it by asking where you sit relative to fifty rather than how close you are getting to a hundred.
Why a T-score instead of a percentage out of a hundred?
Because a percentage is trapped inside its own instrument. A total on a back questionnaire and a total on a hip questionnaire cannot be laid against each other, since each one is scaled to the form that produced it and to the population that form was built for. A norm-referenced metric moves the reference point outside the questionnaire entirely — out to the general population — and the number then means the same thing wherever it turns up.
The disease-specific instruments are built the opposite way, deliberately. The womac index was developed and validated as a disease-specific health status instrument for people with osteoarthritis of the hip or knee, reporting pain, stiffness and physical function 2Ref 2Bellamy N, Buchanan WW, Goldsmith CH, et al. (1988).Validation study of WOMAC: a health status instrument for measuring clinically important patient relevant outcomes to antirheumatic drug therapy in patients with osteoarthritis of the hip or knee.The WOMAC as a disease-specific, self-administered health status instrument validated for osteoarthritis of the hip or knee, and its three-subscale structure holding pain, stiffness and physical function apart rather than merging them.. Inside that population it is precise and meaningful. Ask it about a wrist and it has nothing to say at all.
The PROMIS Physical Function bank was calibrated to a standardized metric against a general-population reference precisely so that its scores would not be locked to one diagnosis or one joint 1Ref 1Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014).The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency.The defining calibration of the adult PROMIS Physical Function item bank: the 124-item bank spanning upper-extremity, central and lower-extremity function plus instrumental activities of daily living; the standardized T-score metric normed to mean 50 and SD 10 in a US general population sample with higher scores indicating better function, read across roughly four standard deviations; the development sample of 149 candidate items tested with 10 SF-36 and 20 HAQ-DI items in 16,065 adults oversampling the chronically ill; the simulated 10-item CAT eliminating floor effects, reducing ceiling effects, achieving higher precision than any comparable-length static tool and discriminating age and disease groups better than legacy instruments; and the absence of any MCID, MDC, cutoff or severity band in this source..
Where the questions come from: one bank, many forms
There is no single PROMIS Physical Function questionnaire. There is an item bank: 124 items calibrated by item response theory, spanning upper-extremity function, central function, lower-extremity function, and instrumental activities of daily living 1Ref 1Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014).The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency.The defining calibration of the adult PROMIS Physical Function item bank: the 124-item bank spanning upper-extremity, central and lower-extremity function plus instrumental activities of daily living; the standardized T-score metric normed to mean 50 and SD 10 in a US general population sample with higher scores indicating better function, read across roughly four standard deviations; the development sample of 149 candidate items tested with 10 SF-36 and 20 HAQ-DI items in 16,065 adults oversampling the chronically ill; the simulated 10-item CAT eliminating floor effects, reducing ceiling effects, achieving higher precision than any comparable-length static tool and discriminating age and disease groups better than legacy instruments; and the absence of any MCID, MDC, cutoff or severity band in this source.. Every short form and every adaptive test draws its questions from that pool, which is how two forms with different wording can report onto one shared scale.
Building it took 149 candidate items, evaluated alongside 10 items from the SF-36 and 20 from the HAQ-DI, using both classical test theory and item response theory 1Ref 1Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014).The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency.The defining calibration of the adult PROMIS Physical Function item bank: the 124-item bank spanning upper-extremity, central and lower-extremity function plus instrumental activities of daily living; the standardized T-score metric normed to mean 50 and SD 10 in a US general population sample with higher scores indicating better function, read across roughly four standard deviations; the development sample of 149 candidate items tested with 10 SF-36 and 20 HAQ-DI items in 16,065 adults oversampling the chronically ill; the simulated 10-item CAT eliminating floor effects, reducing ceiling effects, achieving higher precision than any comparable-length static tool and discriminating age and disease groups better than legacy instruments; and the absence of any MCID, MDC, cutoff or severity band in this source.. The calibration sample was 16,065 adults, deliberately oversampling people with chronic illness 1Ref 1Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014).The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency.The defining calibration of the adult PROMIS Physical Function item bank: the 124-item bank spanning upper-extremity, central and lower-extremity function plus instrumental activities of daily living; the standardized T-score metric normed to mean 50 and SD 10 in a US general population sample with higher scores indicating better function, read across roughly four standard deviations; the development sample of 149 candidate items tested with 10 SF-36 and 20 HAQ-DI items in 16,065 adults oversampling the chronically ill; the simulated 10-item CAT eliminating floor effects, reducing ceiling effects, achieving higher precision than any comparable-length static tool and discriminating age and disease groups better than legacy instruments; and the absence of any MCID, MDC, cutoff or severity band in this source..
That oversampling is not a footnote. An item bank calibrated only on healthy volunteers would be at its blurriest exactly where patients live — at the low-function end, where the clinically important differences are.
How ten questions can outperform a longer questionnaire
Item response theory is what makes the shortcut legitimate rather than lazy. Each calibrated item has a known position on the underlying scale, so a person's answers are not counted up — they are used to locate that person. A small set of well-targeted items can pin a position about as tightly as a long set of items that mostly repeat one another.
In the calibration study, a simulated ten-item computerized adaptive test — one that chooses each next question from the answer before it — eliminated floor effects, reduced ceiling effects, and achieved higher precision than any comparable-length static tool, while discriminating age groups and disease groups better than the legacy instruments it was tested against 1Ref 1Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014).The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency.The defining calibration of the adult PROMIS Physical Function item bank: the 124-item bank spanning upper-extremity, central and lower-extremity function plus instrumental activities of daily living; the standardized T-score metric normed to mean 50 and SD 10 in a US general population sample with higher scores indicating better function, read across roughly four standard deviations; the development sample of 149 candidate items tested with 10 SF-36 and 20 HAQ-DI items in 16,065 adults oversampling the chronically ill; the simulated 10-item CAT eliminating floor effects, reducing ceiling effects, achieving higher precision than any comparable-length static tool and discriminating age and disease groups better than legacy instruments; and the absence of any MCID, MDC, cutoff or severity band in this source..
One caveat belongs with that result, and it is regularly dropped: it describes the adaptive version. A printed short form is static — everyone answers the same items in the same order. It draws on the same calibrated bank and reports on the same metric, but it does not adapt to you 1Ref 1Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014).The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency.The defining calibration of the adult PROMIS Physical Function item bank: the 124-item bank spanning upper-extremity, central and lower-extremity function plus instrumental activities of daily living; the standardized T-score metric normed to mean 50 and SD 10 in a US general population sample with higher scores indicating better function, read across roughly four standard deviations; the development sample of 149 candidate items tested with 10 SF-36 and 20 HAQ-DI items in 16,065 adults oversampling the chronically ill; the simulated 10-item CAT eliminating floor effects, reducing ceiling effects, achieving higher precision than any comparable-length static tool and discriminating age and disease groups better than legacy instruments; and the absence of any MCID, MDC, cutoff or severity band in this source.. Same yardstick, simpler machinery.
What the score does not tell you
How much change counts. The paper that defines this metric reports no minimal clinically important difference and no minimal detectable change, and it establishes no diagnostic cutoffs and no severity bands: PROMIS Physical Function is a norm-referenced continuous measure, not a threshold instrument 1Ref 1Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014).The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency.The defining calibration of the adult PROMIS Physical Function item bank: the 124-item bank spanning upper-extremity, central and lower-extremity function plus instrumental activities of daily living; the standardized T-score metric normed to mean 50 and SD 10 in a US general population sample with higher scores indicating better function, read across roughly four standard deviations; the development sample of 149 candidate items tested with 10 SF-36 and 20 HAQ-DI items in 16,065 adults oversampling the chronically ill; the simulated 10-item CAT eliminating floor effects, reducing ceiling effects, achieving higher precision than any comparable-length static tool and discriminating age and disease groups better than legacy instruments; and the absence of any MCID, MDC, cutoff or severity band in this source.. There is no score at which you become a surgical candidate and none at which you are discharged.
Any figure quoted to you as "the MCID for PROMIS" therefore comes from some other study, in some particular population, after some particular treatment — and it belongs to that population, not to the metric in general.
Compare an instrument that published its thresholds in its own development paper. The Lower Extremity Functional Scale reports a minimal detectable change of 9 scale points, a minimal clinically important difference of the same 9 points, and roughly ±5.3 points of measurement error around any single score 3Ref 3Binkley JM, Stratford PW, Lott SA, Riddle DL (1999).The Lower Extremity Functional Scale (LEFS): Scale Development, Measurement Properties, and Clinical Application.The contrast case of an instrument that published change thresholds in its own development paper: a minimal detectable change and a minimal clinically important difference of 9 scale points each, with a point-in-time measurement error of ±5.3 scale points.. Those numbers are the property of that scale and travel nowhere else.
A norm-referenced score tells you where you stand. On its own it does not tell you how far you have moved.
Universal yardstick, narrow blind spot
The price of comparability is specificity. A general physical-function metric asks about ordinary movement; it does not ask whether your knee gives way when you pivot, or whether the shoulder catches at the top of the reach. Joint-specific instruments ask precisely those questions, and that is sometimes what lets them detect improvement first.
The evidence is concrete rather than theoretical. When the hoos questionnaire was validated in patients having total hip replacement, it proved more responsive than the WOMAC on its pain and symptom subscales — a hip-specific instrument outperforming a broader osteoarthritis instrument on the outcome those patients cared about 4Ref 4Nilsdotter AK, Lohmander LS, Klässbo M, Roos EM (2003).Hip disability and osteoarthritis outcome score (HOOS)—validity and responsiveness in total hip replacement.The HOOS as a validated hip-specific patient-reported outcome, and its greater responsiveness than the WOMAC on the pain and symptom subscales in total hip replacement — the evidence that a joint-specific instrument can detect change a broader one misses.. The koos questionnaire does the equivalent job for the knee, reporting pain, other symptoms, daily activities, sport and recreation, and knee-related quality of life as five separate readings, because a knee that behaves on a pavement and fails on grass needs somewhere to say so 5Ref 5Roos EM, Roos HP, Lohmander LS, Ekdahl C, Beynnon BD (1998).Knee Injury and Osteoarthritis Outcome Score (KOOS)—Development of a Self-Administered Outcome Measure.The KOOS as a validated self-administered knee instrument reporting five separate subscales: pain, other symptoms, daily activities, sport and recreation, and knee-related quality of life.. Above the waist, the dash questionnaire takes the entire arm as one region 6Ref 6Hudak PL, Amadio PC, Bombardier C (Upper Extremity Collaborative Group) (1996).Development of an upper extremity outcome measure: the DASH (disabilities of the arm, shoulder, and hand).The DASH as a self-reported measure of symptoms and physical function treating the upper extremity as a single region across musculoskeletal disorders., and the quickdash covers the same ground by a shorter route.
This is not a contest between instruments. A clinic reasonably runs both: the specific form to steer the treatment week to week, the general metric to answer the larger question of whether a life is coming back.
Function is not pain, and the score knows the difference
Physical function and pain intensity are separate constructs, and this instrument measures the first. It records what you can do — reaching, carrying, climbing, dressing — rather than how much it hurts while you do it 1Ref 1Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014).The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency.The defining calibration of the adult PROMIS Physical Function item bank: the 124-item bank spanning upper-extremity, central and lower-extremity function plus instrumental activities of daily living; the standardized T-score metric normed to mean 50 and SD 10 in a US general population sample with higher scores indicating better function, read across roughly four standard deviations; the development sample of 149 candidate items tested with 10 SF-36 and 20 HAQ-DI items in 16,065 adults oversampling the chronically ill; the simulated 10-item CAT eliminating floor effects, reducing ceiling effects, achieving higher precision than any comparable-length static tool and discriminating age and disease groups better than legacy instruments; and the absence of any MCID, MDC, cutoff or severity band in this source.. The two can move independently, and a single figure that averages them conceals which one actually moved.
This is why pain versus function outcomes are usually collected alongside each other rather than merged. The womac index handles the problem internally, holding pain, stiffness and physical function apart as three subscales rather than delivering one verdict 2Ref 2Bellamy N, Buchanan WW, Goldsmith CH, et al. (1988).Validation study of WOMAC: a health status instrument for measuring clinically important patient relevant outcomes to antirheumatic drug therapy in patients with osteoarthritis of the hip or knee.The WOMAC as a disease-specific, self-administered health status instrument validated for osteoarthritis of the hip or knee, and its three-subscale structure holding pain, stiffness and physical function apart rather than merging them.. PROMIS handles it by keeping physical function as its own metric.
A function score that climbs while pain has not yet fallen is not a contradiction or a failure. Doing more is a real outcome, and for a great many people it is the one they came in for.
Common questions
Related
Muscle, joint & pain
The QuickDASH, a Shorter Arm-Function CheckMuscle, joint & pain
How to Fill Out a Function Questionnaire HonestlyMuscle, joint & pain
The ASES Shoulder Score Your Surgeon Records
Say it back
How would you explain this to someone you love?
Two or three sentences, just as you’d say it. Gale reflects back what you focused on — a mirror, not a quiz.
What a function score will miss
- —Function that collapses over hours or a single day rather than declining across weeks
- —A joint that turns hot, swollen and unusable, particularly with fever or shaking chills
- —New weakness or numbness affecting both legs, or any change in bladder or bowel control
- —Pain that is worst at night, unrelieved by any position, alongside weight loss you did not intend
Sudden loss of strength or sensation, or a change in bladder or bowel control, needs assessment the same day — call 911 or go to the nearest emergency department rather than waiting for a scheduled review.
This page explains how the PROMIS Physical Function metric was built and what its own defining research supports. It is not an interpretation of your score and not a substitute for the clinician who collected it. Nothing here is medical advice.
References
- 1.Rose M, Bjorner JB, Gandek B, Bruce B, Fries JF, Ware JE Jr. (2014). The PROMIS Physical Function item bank was calibrated to a standardized metric and shown to improve measurement efficiency. Journal of Clinical Epidemiology, 67(5):516-526. doi:10.1016/j.jclinepi.2013.10.024 ✓The defining calibration of the adult PROMIS Physical Function item bank: the 124-item bank spanning upper-extremity, central and lower-extremity function plus instrumental activities of daily living; the standardized T-score metric normed to mean 50 and SD 10 in a US general population sample with higher scores indicating better function, read across roughly four standard deviations; the development sample of 149 candidate items tested with 10 SF-36 and 20 HAQ-DI items in 16,065 adults oversampling the chronically ill; the simulated 10-item CAT eliminating floor effects, reducing ceiling effects, achieving higher precision than any comparable-length static tool and discriminating age and disease groups better than legacy instruments; and the absence of any MCID, MDC, cutoff or severity band in this source.
- 2.Bellamy N, Buchanan WW, Goldsmith CH, et al. (1988). Validation study of WOMAC: a health status instrument for measuring clinically important patient relevant outcomes to antirheumatic drug therapy in patients with osteoarthritis of the hip or knee. J Rheumatol. PMID 3068365 ✓The WOMAC as a disease-specific, self-administered health status instrument validated for osteoarthritis of the hip or knee, and its three-subscale structure holding pain, stiffness and physical function apart rather than merging them.
- 3.Binkley JM, Stratford PW, Lott SA, Riddle DL (1999). The Lower Extremity Functional Scale (LEFS): Scale Development, Measurement Properties, and Clinical Application. Physical Therapy, 79(4), 371-383. doi:10.1093/ptj/79.4.371 ✓The contrast case of an instrument that published change thresholds in its own development paper: a minimal detectable change and a minimal clinically important difference of 9 scale points each, with a point-in-time measurement error of ±5.3 scale points.
- 4.Nilsdotter AK, Lohmander LS, Klässbo M, Roos EM (2003). Hip disability and osteoarthritis outcome score (HOOS)—validity and responsiveness in total hip replacement. BMC Musculoskeletal Disorders. PMID 12777182 ✓The HOOS as a validated hip-specific patient-reported outcome, and its greater responsiveness than the WOMAC on the pain and symptom subscales in total hip replacement — the evidence that a joint-specific instrument can detect change a broader one misses.
- 5.Roos EM, Roos HP, Lohmander LS, Ekdahl C, Beynnon BD (1998). Knee Injury and Osteoarthritis Outcome Score (KOOS)—Development of a Self-Administered Outcome Measure. Journal of Orthopaedic & Sports Physical Therapy. doi:10.2519/jospt.1998.28.2.88The KOOS as a validated self-administered knee instrument reporting five separate subscales: pain, other symptoms, daily activities, sport and recreation, and knee-related quality of life.
- 6.Hudak PL, Amadio PC, Bombardier C (Upper Extremity Collaborative Group) (1996). Development of an upper extremity outcome measure: the DASH (disabilities of the arm, shoulder, and hand). American Journal of Industrial Medicine. doi:10.1002/(SICI)1097-0274(199606)29:6<602::AID-AJIM4>3.0.CO;2-LThe DASH as a self-reported measure of symptoms and physical function treating the upper extremity as a single region across musculoskeletal disorders.
6 sources, numbered by first appearance. General health information, not medical advice. AI-assisted editorial content — citations link their sources. Editorial policy