Wearable sleep tracking accuracy vs polysomnography
Consumer wearables track total sleep time within about 30 minutes of PSG but get deep and REM stages right only 50 to 70 percent of the time. The data.

For research and educational purposes only. Not medical advice.
Category: Sleep. 6 min read. By pepSmart Editorial. . .
Key takeaways
- Wrist wearables estimate sleep from motion (accelerometer), heart rate and HRV (photoplethysmography), and skin temperature. They do not record brain activity (EEG), which is what defines sleep stages .
- Total sleep time is the strong number: wearables land within roughly 15 to 30 minutes of polysomnography (Apple Watch Series 8 ran 19.6 minutes long, Fitbit Sense 6.3 minutes long in a 2025 six-device test). Deep (N3) and REM staging is the weak number, right about 50 to 70 percent of the time .
- Wearables systematically miss brief awakenings. Wake-detection specificity ran 0.18 to 0.54 in Chinoy 2021 and 29 to 52 percent in the 2025 test, so devices overestimate how long you slept .
- Sleep scores (Whoop, Oura, Apple, Fitbit, Garmin) are proprietary composites with different weighting. A given night scores differently on different brands, so the number does not compare across devices.
- No consumer wrist wearable diagnoses sleep apnea. If apnea is suspected (loud snoring, choking awakenings, daytime sleepiness, witnessed breathing pauses), the next step is home sleep apnea testing or in-lab polysomnography .
What wearables actually measure
Wrist wearables estimate sleep from a small set of inputs: a 3-axis accelerometer (wrist motion), photoplethysmography (PPG) for heart rate and heart-rate variability, skin temperature in newer devices, and sometimes blood-oxygen estimates from reflective pulse oximetry. They do not measure brain activity. The clinical reference standard for sleep staging is polysomnography (PSG), which combines electroencephalography (EEG), electrooculography (EOG), submental electromyography (EMG), respiratory effort, airflow, and oxygen saturation .
So the sleep-stage chart on your wrist is a model output: an estimate of EEG-derived stages built from non-EEG signals. The model is trained against PSG-labeled sleep, but the inputs are indirect. Different vendors use different proprietary algorithms, and even the same hardware can produce different stage breakdowns across firmware versions.
The AASM scoring standard (what the wearable is trying to estimate)
The American Academy of Sleep Medicine (AASM) scoring manual is the international reference for staging sleep into N1, N2, N3 (slow-wave, or deep), and REM, in 30-second epochs. It specifies the EEG, EOG, EMG, and respiratory criteria for each stage plus the arousal-scoring rules . Even with full PSG data, trained human scorers do not agree perfectly: inter-scorer agreement averages about 83 percent epoch by epoch, and it drops to the mid-60s for the hardest stages, N1 and N3 .
That is the ceiling. Even the gold standard has irreducible noise at the epoch level, so a wearable working from motion and pulse cannot beat humans reading full PSG. The real question is how close it gets with far sparser inputs.
What validation studies actually show: total sleep time close, stages 50 to 70 percent
A widely cited lab validation, Chinoy and colleagues 2021, put seven consumer sleep trackers head to head against polysomnography in 34 healthy adults over three nights, including one disrupted-sleep night. The seven were four wearables (Fatigue Science Readiband, Fitbit Alta HR, Garmin Fenix 5S, Garmin Vivosmart 3) and three non-wearables (EarlySense Live, ResMed S+, SleepScore Max) .
The pattern was consistent. Every device was good at spotting sleep (epoch-by-epoch sensitivity 0.93 or higher) and bad at spotting wake (specificity 0.18 to 0.54), so they credited you with sleep you were not getting and missed brief awakenings. Sleep-latency estimates were off by under 5 minutes. Stage detection was mixed .
The devices people actually wear now do better on stages, but not by much. Here is the current picture from a 2025 lab test of six current wrist devices (Fitbit Charge 5, Fitbit Sense, Withings ScanWatch, Garmin Vivosmart 4, Whoop 4.0, Apple Watch Series 8) and a 2024 test of the Oura Ring Gen3, Fitbit Sense 2, and Apple Watch Series 8:
- Total sleep time: within roughly 15 to 30 minutes of PSG. Apple Watch Series 8 ran 19.6 minutes long, Fitbit Sense 6.3 minutes long .
- Deep sleep (N3): about 50 to 70 percent epoch accuracy (Apple Watch Series 8 50.7 percent, Whoop 4.0 69.6 percent) .
- REM: about 60 to 69 percent (Fitbit Charge 5 60.0 percent, Apple Watch Series 8 68.6 percent) .
- Wake detection: still the weak spot at 29 to 52 percent specificity, so devices overestimate total sleep .
- The Oura Ring Gen3, Fitbit Sense 2, and Apple Watch Series 8 matched PSG on sleep duration, with stage sensitivity 50 to 86 percent and Oura the most consistent (76 to 80 percent) .
- Performance is worse on disrupted-sleep nights, and validation in people who actually have sleep disorders is thin .
None of this is new. De Zambotti and colleagues flagged the same problems back in 2019: proprietary algorithms nobody outside the company can inspect, firmware updates that change the output overnight, and thin independent validation .
The sleep score is a vendor-defined composite, not a measurement
The single sleep score on a wearable is a vendor-defined composite. It blends total time, stage estimates, heart rate, HRV, respiratory rate, skin-temperature trend, and sometimes timing consistency. Because the input signals are noisy and the weighting is proprietary, the same night of sleep can produce different scores across devices.
- Whoop 'Sleep Performance' weights total time and stage breakdown against an internal sleep-need calculation.
- Oura 'Sleep Score' weights total sleep time, efficiency, restfulness, REM, deep sleep, latency, and timing.
- Apple Watch reports stage estimates plus a sleep-duration goal, with no single proprietary score.
- Fitbit 'Sleep Score' (0-100) weights duration, sleep stages, and restoration heart-rate metrics.
- Garmin 'Sleep Score' (0-100) blends similar inputs with stress-balance trends.
The score correlates with subjective recovery in many users, but it does not compare across brands, or even across firmware versions of one brand. A good night on Oura is a different number than a good night on Whoop, even when the underlying physiology is identical.
What wearables can and cannot do clinically
- Can do: track total sleep time trends, catch bedtime-consistency drift, surface low-HRV mornings that many users find track with how recovered they feel, and nudge sleep-hygiene behavior change.
- Cannot do: diagnose obstructive sleep apnea (OSA). Wearable SpO2 is reflective rather than transmissive and lower signal quality than fingertip oximetry, and current consumer wearables are not validated as diagnostic for OSA.
- Cannot do: replace polysomnography for clinical sleep evaluation. Night-to-night variability is high, and a single wearable trace does not establish a chronic pattern at clinical evidence standards.
- Cannot do: pick out micro-arousals or sleep fragmentation at the granularity that matters for restless legs syndrome, periodic limb movements, or REM behavior disorder.
Home sleep apnea testing is the clinical alternative
If sleep apnea is the question, home sleep apnea testing (HSAT) is the appropriate next step. HSAT devices (WatchPAT, ApneaLink, and similar) record airflow, respiratory effort, oxygen saturation, and pulse rate over one or two nights at home, and a sleep physician reads the result. AASM guidelines give a strong recommendation for HSAT with a technically adequate device, or in-lab PSG, to diagnose OSA in uncomplicated adults who show signs and symptoms of moderate-to-severe disease .
HSAT does not replace in-lab PSG when central sleep apnea, complex nocturnal arousals, or pediatric questions are involved. But for the common adult OSA workflow, the HSAT-then-CPAP-titration pathway is well established. Wearable data can flag who should pursue HSAT (loud snoring, choking awakenings, daytime sleepiness, witnessed breathing pauses); it does not substitute for the test .
What this changes for how you use the device
If you use a wearable to track sleep, the defensible read is simple: did total sleep time go up or down, and did your bedtime get more or less consistent. Those are the numbers the devices get right. Stage-by-stage, night-to-night comparisons are noisier than the chart implies, so watch trends over weeks, not single nights.
Persistent symptoms (loud snoring, choking awakenings, daytime sleepiness despite enough hours in bed, insomnia, witnessed breathing pauses) are clinical questions that deserve evaluation whatever the wearable shows . The device gives you cheap behavioral feedback. It does not diagnose anything.
The bottom line on wearable sleep data
Trust the wearable on how long and how regularly you sleep. Treat the stage pie chart and the recovery score as rough estimates from a device that cannot see your brain. If the question is whether you have a sleep disorder, use the wearable as the prompt to go get a real sleep study.
For research and educational purposes only. Not medical advice.
pepSmart has not commissioned independent clinical review of this article.
More on how we write and source these pieces: Editorial process and contributor disclosure and Sourcing posture.
Spot an error? Email corrections via /about.
Sources: 7 entries, all primary canon (peer-reviewed sleep-medicine journals and AASM clinical guidelines), last reviewed 2026-07-08.
References
- [1] Berry RB et al. J Clin Sleep Med 2017: AASM Scoring Manual Updates for 2017 (Version 2.4), the standard rules defining polysomnography channels and sleep staging (PMID 28416048) (PubMed)
- [2] Rosenberg RS & Van Hout S, J Clin Sleep Med 2013: the AASM inter-scorer reliability program for sleep stage scoring, overall agreement 82.6 percent (PMID 23319910) (PubMed)
- [3] Chinoy ED et al. Sleep 2021: performance of seven consumer sleep-tracking devices compared with polysomnography in 34 healthy adults (PMID 33378539) (PubMed)
- [4] Schyvens AM et al. Sleep Adv 2025: performance validation of six commercial wrist-worn wearable devices for sleep stage scoring compared to polysomnography (PMID 40303381) (PubMed)
- [5] Robbins R et al. Sensors (Basel) 2024: accuracy of three commercial wearable devices (Oura Ring Gen3, Fitbit Sense 2, Apple Watch Series 8) for sleep tracking versus polysomnography in healthy adults (PMID 39460013) (PubMed)
- [6] de Zambotti M et al. Med Sci Sports Exerc 2019: wearable sleep technology in clinical and research settings (PMID 30789439) (PubMed)
- [7] Kapur VK et al. J Clin Sleep Med 2017: AASM clinical practice guideline for diagnostic testing for adult obstructive sleep apnea (PMID 28162150) (PubMed)
For research and educational purposes only. Not medical advice.