What Your Sleep Score Actually Measures And What It Misses

What Your Sleep Score Actually Measures And What It Misses

You wake up feeling fine, reach for the phone, and a small number tells you that you slept badly. Or you wake up exhausted and the app awards you an 88. Either way the sleep score has become the first thing millions of people read each morning, and for a metric that governs so many moods it is remarkably poorly understood. It is not a measurement. It is an opinion, generated by a proprietary formula that no two manufacturers agree on.

Where the number actually comes from

A wrist tracker or ring has three or four sensors to work with. An accelerometer notices movement. A photoplethysmography sensor shines green or infrared light through the skin and infers heart rate and heart rate variability from changes in blood volume. Newer devices add skin temperature and blood oxygen estimates. From these signals an algorithm guesses when you fell asleep, when you woke, and how the night divided into light, deep and REM stages. Those guesses are then blended with duration, consistency and disturbance counts into a single figure between zero and one hundred.

The weighting is where the vendors diverge. One brand punishes a late bedtime severely. Another cares mostly about total hours. A third folds in your recent activity load, so an identical night can score 71 on one device and 84 on another worn on the opposite wrist. Neither is lying. They are answering slightly different questions and presenting the answer with the same false precision.

What a wristband cannot see

The clinical reference standard is polysomnography, an overnight study that records brain activity with scalp electrodes alongside eye movement, muscle tone, airflow and blood oxygen. Sleep stages are defined by electrical patterns in the brain, which is precisely the thing a consumer wearable has no access to. The Wikipedia overview of polysomnography gives a sense of how much instrumentation the real measurement requires. Wearables infer brain states from the body instead, and they do it reasonably well for the simple question of asleep versus awake. Stage classification is much shakier, and deep sleep in particular is the figure people should treat with the most caution.

Trackers also miss things that matter clinically. Loud snoring with repeated breathing pauses, restless legs, or a partner who wakes you six times can all coexist with a perfectly respectable score. If you are persistently tired despite good numbers, the numbers are the wrong thing to trust. The NHS guidance on sleep and tiredness is a better starting point than any dashboard, and daytime sleepiness that lasts weeks is a reason to see a doctor rather than to buy a better ring.

When the score starts causing the problem

Sleep researchers have coined the term orthosomnia for anxiety generated by sleep tracking itself. The pattern is familiar to anyone who has done it: you lie awake calculating whether you can still hit seven hours, which guarantees you cannot. People check the app before they check how they feel, and then spend the day behaving according to the number rather than the body. A tracker that makes you tense at bedtime is producing a worse outcome than no tracker at all, and the sensible response is to hide the score for a fortnight and see whether anything changes.

Read the trend, not the night

Single-night scores are noise. Weekly and monthly averages are signal. A tracker earns its price when it shows you that your sleep collapses every Thursday, or that the nights after evening alcohol cost you forty minutes of deep sleep on average across three months. That is an experiment worth running because the sample size is large enough to mean something.

The same logic applies to training. Endurance athletes have learned to treat recovery data as a rolling average rather than a verdict, which is why heart rate variability and zone 2 cardio pair so naturally in modern coaching apps. Low intensity work builds the aerobic base that shows up months later in resting heart rate, and sleep is the other half of that same slow accounting.

The algorithm is a product decision

It helps to remember that a sleep score is designed, not discovered. Someone chose the thresholds, someone decided that eight hours deserves a green ring, and someone decided how encouraging the wording should be. Those choices carry cultural assumptions that travel badly. A siesta culture, a country where the working day starts at seven, and a market where a nine hour night is normal all need different baselines, which is why serious wellness apps localise the interpretation layer and not just the interface strings. The overlap between cognitive science and artificial intelligence is exactly where those judgements get made, and it is a more interesting problem than the hardware.

Getting something useful out of it

Pick one device and stay with it, because cross-brand comparison is meaningless. Ignore the daily figure and look at the weekly one. Treat bedtime consistency as the metric worth optimising, since it is the one wearables measure accurately and the one that responds fastest to behaviour change. And if your best sleep tracking app tells you that you slept badly on a morning when you feel rested, believe your body. The sensor on your wrist is guessing. You are not.