What a recovery score can and cannot know

Patricia Bedoya
·
·
7
min read

Cover image generated with AI — © Chubby Studio S.L.
KEY TAKEAWAY
A recovery score is a summary of your overnight autonomic state, and heart rate variability supplies most of it. That is a real read on how your nervous system handled the last day, but it is not a measurement of muscle damage, stress load or fitness, and it recovers on a different timeline from all three. Trust it when it moves a long way off your own baseline, not when it shifts a few points.
A recovery score knows roughly one thing: what your autonomic nervous system was doing while you slept. It reads that through heart rate variability, resting heart rate, breathing rate and how long you were in bed, compares the night against your own recent nights, and prints the result as a percentage. That is a real measurement, and it is worth having.
It is also much narrower than the word recovery suggests. The gap between those two things is where almost every argument with the number comes from: you feel wrecked and it says 82, or you feel fine and it says 41. Neither reading is a malfunction. The score is answering a smaller question than the one you asked it.
What is a recovery score actually made of?
In 2025 a team led by Cailbhe Doherty catalogued 14 of these composite scores across 10 major wearable brands, including Fitbit, Garmin, Oura, WHOOP, Polar, Samsung, Suunto, Coros and Withings. They then counted which signals fed into them. Heart rate variability appeared in 86% of the scores, resting heart rate in 79%, physical activity in 71% and sleep duration in 71%. Strip the packaging away and most recovery scores are heart rate variability wearing a percentage sign, with a few other measurements adjusting it at the edges.
That matters because it tells you what the number can be sensitive to. HRV is a read on the balance between the two halves of your autonomic nervous system, and it moves with things like sleep, alcohol, illness and stress. If you want the mechanism rather than the summary, what HRV actually tells you is the longer version. What the number is compared against is not a population table either, it is your own rolling average, which is the only comparison that makes any sense for a metric this personal. We have written separately about why personal baselines beat normal ranges.
Why does every app give a different number for the same night?
The same review found significant discrepancies between brands in how long a window they look at, how they weight each contributor, and what the scoring method does with them. And then the part that decides how much trust the number deserves: not one of the ten manufacturers published the formula behind its score, and few offered any peer-reviewed evidence that the score is accurate or clinically meaningful.
So a recovery score is only ever comparable to itself. Two apps reading the same wrist on the same night can legitimately disagree by twenty points, and neither is lying. It also means a score can shift because the company changed its weighting, not because anything changed in you.
What does a recovery score genuinely know?
More than the sceptics allow. A heavy drinking night, the day a fever starts, a session finished at eleven at night, four hours in bed instead of seven: these move overnight heart rate variability and resting heart rate by amounts far larger than the ordinary night-to-night wobble. On days like that the score is not guessing, it is reporting something that measurably happened. The number is far better at flagging an extreme than at grading the middle.
The middle is where it turns vague. A 68 and a 74 are not two different physiological states you can act on differently. They are two points inside the noise band of a measurement that varies night to night for reasons nobody can name. Most of your days land there, which is why most days the score should not be the loudest voice in the room.

See which signal actually moved
FitWoody reads five overnight signals against your own baseline and tells you whether each one is above, below or in range, instead of collapsing your night into a single percentage. Comes with a 7-day free trial.
Download the app
★★★★★
4.5 · 732 ratings on the App Store
What can the score not see at all?
This is the part worth understanding, because it explains the mornings when the number is generous and your legs are not. In 2019 Andrew Flatt and colleagues put ten trained men through six sets to failure on squat, bench press and pull-down, then tracked three separate things for the next two days: heart rate variability, actual neuromuscular performance, and how recovered the lifters said they felt.
The three did not move together at all. HRV was crushed immediately after training and was already back at baseline by the next morning. Countermovement jump power and barbell velocity took a full 48 hours to return. Perceived recovery was still suppressed at 48 hours. And when the researchers tested for associations between the measures, none reached significance. The autonomic system had signed off a full day before the muscles had. It is one small study, ten people, and it deserves to be read as an illustration rather than a law, but the direction it points is the one that matches how people actually feel.
Line up what is invisible to a wrist at night and the list gets long: muscle damage and connective tissue load, depleted glycogen, how sore you will be on the stairs, the deadline you are carrying, the argument you had at dinner, the fact that you were ill last week and this is your first session back. None of that reaches the sensor. The score is not ignoring those things. It was never able to see them.
Does the score know how you will perform today?
Here the honest answer is uncomfortable for anyone selling a readiness number, and it comes from the largest review on the question. Saw, Main and Gastin went through 56 studies that measured athlete well-being both objectively and subjectively, and reported in the British Journal of Sports Medicine that the two kinds of measure generally did not correlate with each other. More pointedly, the subjective measures tracked both acute and chronic training load with better sensitivity and consistency than the objective ones did.
Read that carefully, because it is easy to over-claim in either direction. It does not mean the sensors are wrong. It means that when the question is how the training is landing on a person, the most sensitive instrument in the literature is a short question you answer yourself, and the physiological measurement is the supporting evidence rather than the verdict.
Why does the number argue with how you feel?
Because the two are measuring different things and only one of them has access to the whole picture. Feeling flat on a green morning usually means something real that the sensor cannot reach, and feeling strong on a red one usually means the overnight deviation had a cause that has already passed. Neither is a reason to distrust your own body. If today is one of those low days, adapting the session instead of skipping it is almost always the better move, and if the pattern has lasted a couple of weeks then the question is no longer about one score but about whether you are overreaching or simply tired.
People who live with these scores work this out on their own. When researchers interviewed 17 regular exercisers who had used a Whoop band or an Oura ring for at least three months, one of the three themes that came out of the interviews was the users recognising the limits and the errors of the devices, and leaning on their own judgement to decide what to do. One of them put it as not being able to capture the complexities of a human on a device. They kept the number. They just stopped letting it have the final word.
So what is the number good for?
Treat it as one instrument on the dashboard rather than the verdict on your day. It is a genuine read on how your nervous system handled the last 24 hours, it is at its most useful when it moves a long way off your own normal, and it is close to meaningless as a way of separating a decent day from a slightly better one. A score is a measurement, not a grade, and nothing about a low one says you failed.
This is the reason FitWoody does not collapse your night into a single recovery percentage. It reads five overnight signals, heart rate variability, resting heart rate, body temperature, blood oxygen and respiratory rate, and for each one tells you whether you are above, below or in range against your own baseline. Five plain readings take a moment longer to take in than one number does, and they are much harder to misread, because you can see which signal actually moved instead of guessing what the percentage is upset about.
Frequently asked questions
Is a recovery score accurate?
It is an accurate summary of a narrow thing: your overnight autonomic state, which is mostly heart rate variability. It is reliable when it moves a long way off your own baseline and close to meaningless when it shifts a few points. No major manufacturer publishes its formula, so a score can only be compared with itself, never with another app's.
Why is my recovery score always low?
The score compares your night with your own recent nights, so a persistently low reading usually means whatever is depressing your baseline is still present: short or irregular sleep, alcohol, a heavy training block, illness or ongoing stress. If it stays low for weeks with no cause you can point to, that is a conversation for your doctor rather than for an app.
Do two apps give the same recovery score for the same night?
No, and they are not built to. A 2025 review of 14 composite scores across 10 brands found they differ in the window of data they analyse, which signals they include and how each one is weighted, and none of the manufacturers disclosed the exact formula. Two apps can read the same wrist on the same night and disagree widely without either being wrong.