Apple Watch sleep tracking accuracy in perimenopause
Ask how accurate Apple Watch sleep tracking is and you will usually get one reassuring number back: it identifies sleep correctly more than 90% of the time. That figure is real. It is also the wrong half of the answer for anyone whose nights have started coming apart in perimenopause — because the single thing a wrist device is worst at happens to be the thing your night is now largely made of.
The number everyone quotes, and the one they leave out
Accuracy against a sleep lab is really two separate skills, and they are usually reported as one.
The first is sensitivity: of all the minutes you were genuinely asleep, how many did the device call sleep? Wrist wearables are excellent at this. In a 2025 laboratory comparison of six commercial devices against polysomnography — the Apple Watch Series 8 among them — every device correctly labeled more than 90% of the sleep epochs.
The second is specificity: of all the minutes you were genuinely awake, how many did the device call awake? In that same study, the devices managed somewhere between roughly 29% and 52%. Half of your wakefulness, or more, was filed as sleep.
The reason is not a flaw anyone can patch out. A wrist device infers your state from movement and heart signals. Lying still in the dark at 3 a.m. with your mind running produces almost no movement and a calm heart rate — which is, to a sensor, very close to what sleeping produces. Quiet wakefulness is the blind spot, and quiet wakefulness is not a rare edge case in midlife.
Why that particular blind spot lands here
Sleep during the menopause transition does not usually fail at the start of the night. It fails in the middle. Across cohorts, reported sleep disturbance rises from roughly 16–42% before the transition to 39–47% during it, and it breaks in a specific way — waking up, rather than being unable to fall asleep.
So the measurement error and the symptom overlap almost perfectly.
It gets sharper than that. Wake detection does not fail at a constant rate; it fails progressively. When researchers compared actigraphy against polysomnography across conditions containing different amounts of wakefulness, accuracy fell as wakefulness rose, and total sleep time and sleep efficiency were overestimated more strongly in the more disturbed conditions.
Read that slowly, because it inverts the usual assumption. The error is not random noise that averages out. It scales with the very thing you opened the app to look at, and it runs in the flattering direction. The worse the night, the more generous the estimate is likely to be — which is one reason a night that felt shattered can come back looking unremarkable. A night sweat you slept through and a stretch of still, open-eyed wakefulness can both leave a thinner trace than they deserve.
There is no single accuracy number to quote
The comparison also assumes the reference standard is fixed. It is not.
When the American Academy of Sleep Medicine had more than 2,500 scorers read the same laboratory recordings, agreement on sleep stages averaged 82.6% — good, not perfect. Agreement was lowest exactly where consumer interest is highest: about 67% for stage N3, the deep sleep everyone screenshots, and about 63% for stage N1. Trained humans, same recording, one in three disagreeing about deep sleep.
Meanwhile the pooled picture points a different way again. A meta-analysis of 24 studies covering 798 participants found wrist devices estimated total sleep time about 17 minutes below polysomnography on average, not above it. That is not a contradiction of the epoch-level finding — it is a different measurement, averaged across devices, algorithms and populations, where opposing errors partly cancel.
The honest conclusion is the one almost no review states plainly: there is no single accuracy figure for this. There is a device, an algorithm version, a study population and a specific question, and the answer changes with all four.
What survives all of that
Quite a lot, once you stop asking the wrong question of it.
There is a difference between a measurement being wrong and a measurement being inconsistent. A tape measure that reads two centimeters long is useless for telling you your height and perfectly good for telling you whether you have grown. Systematic bias cancels when you compare a thing to itself.
That is the use that survives the validation literature, and the researchers say so themselves. The 2025 comparison found agreement with the lab ranged from fair to moderate across devices, with the Apple Watch Series 8 at the top of that range, and concluded that the better-agreeing devices could be used to track prolonged and significant changes in sleep architecture — changes over time, not a verdict on Tuesday. The seven-device study before it reached the same practical conclusion from the other end: use these for sleep-and-wake outcomes rather than stage breakdowns.
This is baseline-not-average stated as a measurement fact rather than a philosophy. Your watch is a mediocre instrument for answering “how much deep sleep did I get” and a genuinely useful one for answering “are my nights breaking more often this month than last.” The same is true of the other overnight numbers, including sleeping respiratory rate, and it is why we treat the deep sleep percentage as the softest figure on the screen.
One more thing worth knowing before you decide the watch is simply wrong and your memory is right. When 323 midlife women were recorded at home with in-lab-quality polysomnography, a research wrist monitor and a sleep diary on the same nights, the diaries were the outlier: on average they reported more total sleep, higher efficiency and less time awake than the lab found. Both accounts of a night drift. Neither is the truth.
How Perigee reads a night it cannot fully trust
We built Perigee around what the sensor can actually support. It reads your night against your own recent nights rather than a population norm, keeps the emphasis on timing and continuity — when you settled, when the night broke, whether that is shifting across weeks — and treats the stage breakdown as a sketch rather than a measurement. It is never a score, and it will say when the data is thin instead of filling the gap with a confident story.
The mechanics of what your watch records overnight, and where each signal stops, are laid out on the Sleep signal page; the whole approach starts from the same premise.
One small thing to try this week
For the next seven mornings, before you open any app, write down one word for the night you just had. Rough. Broken. Fine. Then look at the record.
You are not checking which one is correct. You are looking for the nights where they disagree — the shattered night the watch recorded as unremarkable, or the calm reading after a night you barely remember. Those mismatches are the most informative thing this pairing produces, and they are worth taking to a clinician if broken sleep is wearing you down.
Your sleep record shows more time awake than on a typical night. That can leave a full night feeling less restorative. Keep tonight simple and note whether heat, sweating or a racing heart breaks sleep again.
Questions, answered
How accurate is Apple Watch sleep tracking?
It depends entirely on which question you mean. In a 2025 laboratory comparison of six wrist wearables, every device correctly labeled more than 90% of the epochs a sleep lab scored as sleep — but correctly labeled the awake epochs only about 29% to 52% of the time. It is strong at recognizing sleep and weak at recognizing wakefulness, and no single percentage captures both.
Does my watch overestimate how much I slept?
Often, and most on the nights you care about. Actigraphy research shows wake detection degrades as the amount of wakefulness in a night rises, with total sleep time and sleep efficiency overestimated more strongly in the most disturbed conditions. Lying still and awake looks a great deal like sleeping to a sensor reading movement and heart rate.
Can I trust the deep sleep percentage on my Apple Watch?
Treat it as the softest number on the screen. Stage classification varies widely between consumer devices, and even trained human scorers agree on stage N3 only about 67% of the time when reading the same laboratory recording. The broad shape of several nights carries more information than any single night's percentage.
Is a sleep tracker still worth wearing during perimenopause?
Yes, for a narrower purpose than the marketing suggests. What holds up is timing and continuity read against your own recent weeks — when you settled, when the night broke, whether that is changing. What does not hold up is treating one night's stage breakdown as a measurement of how you slept.
- Schyvens AM, Peters B, Van Oost NC, et al. A performance validation of six commercial wrist-worn wearable sleep-tracking devices for sleep stage scoring compared to polysomnography. SLEEP Advances. 2025. PMC12038347. pmc.ncbi.nlm.nih.gov/articles/PMC12038347
- Paquet J, Kawinska A, Carrier J. Wake detection capacity of actigraphy during sleep. Sleep. 2007. PMID 17969470. PMC2266273. pubmed.ncbi.nlm.nih.gov/17969470
- Chinoy ED, Cuellar JA, Huwa KE, et al. Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep. 2021. PMID 33378539. PMC8120339. pubmed.ncbi.nlm.nih.gov/33378539
- Rosenberg RS, Van Hout S. The American Academy of Sleep Medicine inter-scorer reliability program: sleep stage scoring. Journal of Clinical Sleep Medicine. 2013. PMID 23319910. PMC3525994. pubmed.ncbi.nlm.nih.gov/23319910
- Performance of consumer wrist-worn sleep tracking devices compared to polysomnography: a meta-analysis. Journal of Clinical Sleep Medicine. 2025. PMID 39484805. PMC11874098. pubmed.ncbi.nlm.nih.gov/39484805
- Lehrer HM, et al. Comparing polysomnography, actigraphy, and sleep diary in the home environment: the Study of Women’s Health Across the Nation (SWAN) Sleep Study. SLEEP Advances. 2022. PMC8918428. pmc.ncbi.nlm.nih.gov/articles/PMC8918428
- Kravitz HM, Joffe H. Sleep during the perimenopause: a SWAN story. Obstetrics and Gynecology Clinics of North America. PMC3185248. pmc.ncbi.nlm.nih.gov/articles/PMC3185248
Perigee doesn’t provide medical advice or diagnose any condition. It organizes your Watch readings so you and your doctor can review them together.