How Sleep Scores Are Calculated, and What They Miss

Asier G. Morato
·
·
6
min read

Cover image generated with AI — © Chubby Studio S.L.
KEY TAKEAWAY
A sleep score is a weighted summary, not a measurement. It combines how long you slept, how broken the night was, the estimated share of deep and REM sleep, and your timing, into a single number. Because the stage estimates are inferred from movement and heart rate rather than brain activity, the score is far more useful as a trend against your own history than as a grade for one night.
A sleep score is a weighted summary, not a reading off an instrument. Almost every version of it uses the same four ingredients: how long you slept, how that time was divided between light, deep and REM sleep, how broken the night was, and when it happened relative to your usual schedule. Each ingredient is compared against a target, scored, weighted, and added together. The result is printed as a number out of 100, which makes it look far more precise than it actually is.
What goes into a sleep score?
Duration is the backbone. Almost every scoring model starts with total time asleep and compares it against a target, either a fixed adult range or a goal you set yourself. Continuity comes next: how many times you woke, and how much of the night you spent awake after first falling asleep. Then composition, meaning the share of the night estimated as deep and REM sleep. Timing usually contributes something too, since going to bed three hours later than usual is treated as a cost even when the total is fine. Some models add a physiological layer on top, usually resting heart rate or heart rate variability measured while you slept.
That list is fairly consistent across apps. The weights are not. No two companies combine those ingredients the same way, and almost none of them publish the formula, which is why the same night can produce noticeably different scores on two devices without either one being broken. It also means a score is not comparable between people, or even between two apps on the same wrist.
Where do the stage numbers come from?
This is the part most people skip, and it is the part that matters most. In a sleep lab, stages are read from electrical activity: an EEG for brain waves, plus sensors for eye movement and muscle tone. A watch has none of that. It has an accelerometer that senses movement and an optical sensor that reads your pulse at the wrist.
The inference is not arbitrary, to be fair to it. Your heart slows and steadies in deep sleep, breathing becomes regular, and you barely move. In REM your pulse turns erratic again, closer to a waking pattern, while your body stays still. Those signatures are real, and a model can learn them. The trouble starts at the edges, where light sleep and quiet wakefulness look almost identical from the outside: lying still, breathing calmly, heart rate low. From the wrist, resting quietly with your eyes shut and being lightly asleep are close to the same event.
So the stages in your app are not observed. Your watch is not measuring sleep stages, it is predicting them, using a model trained to guess what a sleep technician would have written down given that pattern of movement and beat-to-beat heart rate. It is a reasonable prediction. It is still a prediction, and everything built on top of it inherits that uncertainty. If you want the underlying picture first, here is what light, deep and REM sleep actually do.

Score your night against your own baseline
FitWoody scores your sleep against your own sleep goal and your usual patterns rather than a population ideal, then shows it next to your metabolic baselines and training load. Plus comes with a 7-day free trial.
Download the app
★★★★★
4.5 · 732 ratings on the App Store
How close does the prediction get?
Closer than critics claim, further than marketing implies. In a 2025 validation study, 62 adults slept a night wired to full polysomnography while wearing up to four consumer devices at once. For four-stage classification, agreement with the lab ranged from a Cohen kappa of 0.21 to 0.53, with the Apple Watch Series 8 at the top of that range. In plain terms: fair to moderate agreement, and not interchangeable with a lab.
The more revealing number is the split between detecting sleep and detecting wake. Every device tested was much better at noticing you were asleep than at noticing you were awake, with sensitivity above 90% but specificity somewhere between roughly 29% and 52%. Most also overestimated total sleep time. The practical effect is that brief awakenings tend to get absorbed into sleep, so the continuity part of your score is usually flattered rather than punished.
Even the gold standard disagrees with itself
It is worth being fair to the algorithms here, because the target they are trained on is blurrier than most people assume. In the Sleep Revolution project, ten scorers across seven European sleep centres scored the same 50 overnight recordings. Overall agreement was a Fleiss kappa of 0.714, which is good. But all ten agreed on the exact stage in only 48% of the thirty-second windows they scored. Agreement was highest for REM sleep and worst for N1, the lightest stage of all.
That reframes the whole question. If trained humans reading actual brain waves reach unanimity on fewer than half of a night, a wrist-worn model predicting those same labels from movement and pulse is never going to be exact. A score that shifts from 84 to 87 is reporting a difference far smaller than the noise sitting underneath it.
What a sleep score cannot tell you
It cannot tell you why. A score can tell you a night was unusual, but it has no idea what made it unusual: the second glass of wine, the room being two degrees too warm, the conversation you were still rehearsing at 2am. That interpretation is yours, and it is the part that actually changes anything.
It also cannot tell you whether something is medically wrong. The American Academy of Sleep Medicine is direct about this: consumer sleep devices cannot diagnose sleep disorders, and their data belongs inside a proper evaluation rather than instead of one. If you snore heavily, wake up gasping, or feel wrecked despite nights that look fine on paper, that is a conversation with a doctor, not with an app.
When the score and your body disagree
Believe your body. Clinicians have a name for what happens when people stop doing that: orthosomnia, described in 2017 after patients began arriving for treatment of sleep problems they had diagnosed from their tracker. Some were spending extra hours in bed purely to raise the number the device reported, which is one of the more reliable ways to make insomnia worse. If you feel rested and the app says 62, the useful piece of information is that you feel rested.
How to actually use the number
Read the trend, not the night. A single score is mostly noise, but a two-week average that keeps sliding is signal. One bad score is weather, three weeks of falling scores is climate. Compare it against your own history rather than against a friend, for the same reason your own baseline beats a generic normal range everywhere else in health data.
It also helps to read the score next to metrics that are measured more directly than stages are inferred. Overnight heart rate comes from your actual pulse rather than from a stage classifier, which is why a rising resting heart rate and what HRV actually tells you will often flag a problem before a sleep score does.
This is how FitWoody handles it. Your night is scored against your own sleep goal and your usual patterns rather than a population ideal, and it sits alongside your metabolic baselines and your training load, so the number becomes a starting point for a decision about today instead of a grade for last night.
The number only means something next to your own history, which is why we score sleep against your goal and your usual patterns instead of a fixed ideal.
Frequently asked questions
How is a sleep score calculated?
Most sleep scores combine four inputs: total time asleep, how continuous the night was, the estimated share of deep and REM sleep, and how well the timing matched your usual schedule. Each input is compared against a target, weighted, and added into one number. The weightings differ between apps and are rarely published, so scores are not comparable across devices.
Are sleep scores accurate?
The arithmetic is exact. What is uncertain is what goes into it. Wrist devices infer sleep stages from movement and heart rate rather than brain activity, and a 2025 validation of six consumer wearables found four-stage agreement with a sleep lab ranging from a Cohen kappa of 0.21 to 0.53. Treat a score as a trend indicator, not a measurement.
What is a good sleep score?
There is no universal answer, because every app scales its number differently. The more useful question is whether your score is drifting against your own two-week average, and whether it agrees with how you actually feel and with your resting heart rate and HRV.