Every consumer sleep tracker promises to break your night into stages, hand you a score, and tell you how you slept. Almost none of them tell you how accurate any of that is against the actual clinical measurement — polysomnography — that sleep medicine considers the ground truth.
That is not because the data does not exist. It does. There have been at least a dozen serious validation studies published in peer-reviewed sleep journals over the last decade, and the picture they paint is more nuanced than either the marketing or the skeptics suggest. Consumer trackers are neither useless nor accurate. They are good at some things, bad at others, and the honest way to use them depends on knowing which is which.
This piece walks through what polysomnography actually measures, what the validation studies found, why REM detection is the hardest problem in the space, and which devices in our own catalog have published k-coefficient work worth trusting.
What polysomnography measures
Polysomnography, or PSG, is the overnight study a sleep doctor orders when they need to diagnose a sleep disorder. It happens in a lab (or occasionally at home with a portable rig) and simultaneously records:
- EEG — electroencephalography, brain wave activity through scalp electrodes. This is how sleep stages (N1, N2, N3 slow-wave, REM) are actually identified. Everything else in a sleep study is context.
- EOG — electrooculography, eye movement, which distinguishes REM sleep from other stages.
- EMG — electromyography at the chin and legs, which drops in REM (muscle atonia) and spikes during arousals or periodic limb movements.
- Respiratory effort and airflow — to catch apneas and hypopneas.
- Pulse oximetry — SpO2 to detect desaturations.
- ECG — single-lead cardiac rhythm.
A trained polysomnographic technologist scores the night in 30-second epochs, assigning each one to a stage (Wake, N1, N2, N3, or REM) according to the AASM scoring manual. The output is a hypnogram plus a report. It is a labor-intensive process, and it is the reference standard everything else gets compared against.
What a consumer wearable actually measures
Your wrist tracker, ring, or under-mattress sensor sees a much thinner slice of physiology. Depending on the device, it captures some combination of:
- Pulse rate via photoplethysmography (PPG) — an LED shines through your tissue, a photodetector measures the pulsatile change in blood volume with each heartbeat.
- Pulse rate variability, derived from beat-to-beat PPG intervals.
- Body movement via a 3-axis accelerometer.
- SpO2 via multi-wavelength PPG (higher-end devices only).
- Skin temperature (some devices).
- Ballistocardiography via a load-cell strip under the mattress (Withings Sleep Mat, some Beddit devices).
From those signals, a proprietary algorithm attempts to classify each 30-second window into a sleep stage. There is no direct measurement of brain activity. The algorithm is essentially a statistical inference from heart-rate patterns, movement, and (on better devices) breathing patterns.
The interesting question is: how well does that inference work?
The Chinoy validation studies (Sleep journal, 2020 and 2021)
The most cited independent validation work in this space is a pair of studies from Evan Chinoy and colleagues published in the journal Sleep — the flagship journal of the American Academy of Sleep Medicine.
The 2020 paper ("Performance of Seven Consumer Sleep-Tracking Devices Compared with Polysomnography") tested seven then-current consumer devices against gold-standard PSG in a controlled lab environment. Devices included Fitbit models, Oura Ring, Whoop, Garmin, and Apple Watch generations available at the time. The 2021 follow-up broadened the population and looked at newer models.
Two headline findings are worth quoting in plain language.
Total sleep time and sleep efficiency are the easiest to get right. Most consumer devices agreed with PSG within 5 to 15 minutes on total sleep time and within a few percentage points on sleep efficiency. If you want to know roughly how much you slept, most modern trackers are competent.
Stage classification is harder — meaningfully so. Epoch-by-epoch stage agreement, quantified by Cohen's kappa (a statistic that corrects agreement for chance), landed in the 0.4 to 0.6 range for most devices across most stages. Kappa of 1.0 is perfect agreement; 0.0 is chance-level; 0.4 to 0.6 is conventionally called "moderate" agreement. For context, PSG-to-PSG scorer agreement between two trained human technologists on the same night is roughly kappa 0.75 to 0.80 — not perfect either, but the ceiling.
Consumer devices are meaningfully below the human-scorer ceiling. They are also well above chance. That is the honest state of the art.
Why REM detection is the hardest
Across nearly every published validation study, REM is the stage consumer devices miss most often.
This is not accidental. REM sleep is characterized by rapid eye movement (visible on EOG), muscle atonia (visible on EMG), and a specific low-amplitude mixed-frequency EEG signature. A wrist tracker cannot see any of those signals directly. It has to infer REM from indirect proxies — typically an increase in heart-rate variability with maintained low motion. Those proxies work reasonably well in typical adults but break down in several populations:
- Older adults, whose autonomic modulation of heart rate is dampened.
- People on beta-blockers or other cardio-active medications, which flatten the HR signal REM detection depends on.
- People with sleep apnea, whose arousals confound the movement-plus-HR pattern.
- People with REM sleep behavior disorder, who move during REM.
The upshot: your device probably knows within a few minutes how long you slept, has a reasonable guess at how much of that was NREM versus REM, and is genuinely bad at telling you which specific 30-second window was REM versus N2 or N3.
Why deep sleep is often overestimated
Another consistent finding across validation studies: many wrist-based devices overestimate deep (slow-wave, N3) sleep compared to PSG.
The reason is mechanical. When you are motionless with low heart rate and stable breathing, the algorithm's best guess is deep sleep. But motionless-with-low-HR is also what a long stretch of stage N2 looks like. And so a fair chunk of what your wearable labels "deep sleep" is actually N2 in the underlying PSG.
This matters because deep sleep is one of the metrics people fixate on. Marketing tells you that deep sleep is when you recover, and your watch tells you that you got 18 percent of it, and both of those things are approximately true and approximately misleading. The 18 percent is a noisy estimate; the truth might be 11 percent deep and 7 percent N2 that got classified as deep. For a healthy adult, this rarely changes any actionable decision. It does mean that fetishizing a specific deep-sleep percentage is more theater than science.
The k-coefficient — what it actually means
Cohen's kappa is worth understanding because it is the statistic you should actually look for when a device manufacturer claims accuracy.
Kappa measures agreement between two raters (or a device and a reference) after correcting for the amount of agreement that would happen by pure chance. A device that classifies everything as "asleep" would get very high raw accuracy for people who spend most of their night asleep, but kappa would correctly penalize that non-informative agreement.
Conventional interpretation, from the Landis and Koch 1977 paper:
- < 0.20 — Slight agreement (basically noise).
- 0.21 to 0.40 — Fair.
- 0.41 to 0.60 — Moderate.
- 0.61 to 0.80 — Substantial.
- 0.81 to 1.00 — Almost perfect.
Most consumer wearables land in Moderate. A few of the best under-mattress and finger-ring devices land at the low end of Substantial for some stages. This is the honest ceiling of the current generation of consumer sleep hardware.
Devices in the catalog with honest validation
Three devices we carry have published sleep-related validation work worth citing.
The Sleep Mat has published validation for its sleep-staging algorithm against PSG in the Journal of Clinical Sleep Medicine and other venues. The kappa values reported are in the low-to-mid Substantial range for total sleep time and mid-Moderate for stage-specific classification — the best of the three form factors we carry. It also has a peer-reviewed breathing-disturbance-index validation used in some countries for at-home sleep apnea screening.
The ScanWatch 2 is not going to beat the Sleep Mat on staging accuracy — no wrist device will, for the physical reasons described above. But it is competitive within the wrist-PPG category and includes overnight HRV and cardiac screening in one device.
Fitbit's algorithm was among the earlier consumer stage-classification models and has been iterated for over a decade. Independent validation puts it in the moderate-agreement range against PSG, similar to other mature wrist devices.
How to actually use a sleep tracker knowing all this
If the epoch-by-epoch stage classification is only Moderately accurate, why bother?
Because you are almost never going to make a health decision based on epoch-by-epoch scoring. What matters is:
- Total sleep time. Most trackers get this within 15 minutes. Good enough to matter.
- Sleep efficiency and WASO. The percentage of time in bed you were actually asleep, and how much time you spent awake after first falling asleep. Both are reliably measured.
- Trend in nightly averages. Whether your seven-day rolling average is trending up or down over months is where the real signal lives. Any one night is noise; a persistent trend is a real change.
- Overnight HRV as a stress marker. Not perfectly accurate but internally consistent enough for trend-tracking.
The parts of the sleep score people fixate on — "you only got 8 percent deep sleep last night" — are the parts that are least reliable. The parts they ignore are usually the ones worth watching.
The honest bottom line
Peer-reviewed validation shows that current-generation consumer sleep trackers achieve moderate agreement with polysomnography on stage classification (typically kappa 0.4 to 0.6), good agreement on total sleep time and efficiency, and consistent under-detection of REM and over-detection of deep sleep. That is the honest range. Anyone claiming higher without citing a peer-reviewed source is doing marketing, not measurement.
Within those limits, a sleep tracker worn every night for a year is a real tool. Under-mattress form factors have the strongest independent validation. Wrist devices are useful trend-trackers if you accept the limits of what they can see. Neither is a substitute for a formal sleep study if you suspect you have a sleep disorder — that is still a clinical question that requires PSG.
Buy a tracker. Wear it consistently. Watch the seven-day trend, not any single night. And when the marketing tells you that your "sleep score" is 73 out of 100, remember that number was invented by product designers and does not appear in any peer-reviewed sleep journal. The kappa does. That is the number to look for.