Reference · Devices

How each wearable computes HRV, and why the numbers do not match

Two people compare their “HRV” and one is 95, one is 38. Nothing about their bodies may differ. What differs is what their devices are counting.

All guides

The same word for different measurements

Every major wearable shows a number labelled HRV. Behind the label sit at least three decisions the manufacturer made for you: which statistic to compute, which part of the day to compute it over, and how to filter the raw signal. Different brands have made different choices, and none of them is wrong — but it means a reading from one device has no defined relationship to a reading from another, or to the published norms, which are mostly from short seated electrocardiogram recordings.

This guide summarises what each device does, to the best of publicly available documentation at the time of writing. Manufacturers change their algorithms; treat the specifics as a starting point and check your own device's documentation if the detail matters to you.

Device-by-device summary

DeviceMetricWhen sampledReported as
Oura RingRMSSDOvernight, 5-minute windows across sleepNightly average plus a trend line
WhoopRMSSDOvernight, weighted to the final slow-wave sleep periodSingle nightly value feeding “Recovery”
GarminRMSSDOvernight, 5-minute windows across sleepNightly average and 7-day average against a personal baseline (“HRV Status”)
Apple WatchSDNNSporadic ~1-minute samples during the day and during Mindfulness sessionsIndividual readings in the Health app
Fitbit / Pixel WatchRMSSDOvernight, during sleepNightly value and 30-day trend
PolarRMSSDRoughly the first four hours of sleep (Nightly Recharge)“ANS charge” relative to a 28-day baseline
Chest strap + appRMSSD (usually)Whenever you record, typically 1–5 min on wakingRaw ms value, usually with a rolling baseline

RMSSD versus SDNN

RMSSD — the root mean square of successive differences between beats — is dominated by fast, breath-to-breath changes in the interval, which are produced almost entirely by the vagus nerve. It is stable in short recordings and is what almost every overnight device uses. SDNN is the standard deviation of all intervals in the recording; it picks up slower rhythms driven by blood pressure regulation, temperature and the sympathetic branch, and it grows with recording length, so a 1-minute SDNN and a 24-hour SDNN are different quantities. Apple's choice of SDNN from short daytime samples makes its numbers the least comparable of the group — not less valid, but describing something else.

The practical consequence: an Apple Watch SDNN of 45 and an Oura RMSSD of 45 share a unit and nothing else.

When the sample is taken

Overnight values are systematically higher than waking ones, and deep-sleep values higher still, because lying down, slow regular breathing and the absence of cognitive load all push vagal tone to its daily peak. That is why a device weighting the last slow-wave period will often read higher than one averaging the whole night, which will read higher than a seated morning recording with a chest strap — all in the same person, the same night.

Overnight sampling also brings a confound that is easy to miss: alcohol, late meals and late exercise have their largest HRV effects in the first half of the night. A device sampling early sleep will show them strongly; one weighting late sleep will partly miss them. Neither is lying.

Optical sensors and the accuracy question

Rings and wristbands measure pulse optically (photoplethysmography), inferring beats from blood-volume changes, rather than electrically. At rest, with the device snug and the wearer still, agreement with ECG-derived RMSSD is good in validation studies. Agreement deteriorates with movement, poor fit, cold extremities and some skin tones, and it deteriorates asymmetrically — artefacts almost always inflate RMSSD rather than lower it, because a missed or spurious beat creates a huge “successive difference”. Devices filter for this, but a single-night spike with no plausible cause is more often a sensor problem than a physiological one.

A chest strap measures the electrical signal directly and is the consumer gold standard. Several validation studies have found that recent strap models agree with clinical ECG closely enough for the difference to be irrelevant for personal tracking.

What to do with this

  1. Never compare across devices or people. Compare yourself to yourself on one device.
  2. Use the device's own baseline feature. Garmin, Polar, Whoop and Oura all compute a personal baseline; that comparison is the only one that is methodologically clean.
  3. If you switch devices, start a new history. Do not try to translate old values into new ones.
  4. When entering HRV into the Calmspan test, use a seven-day average of an overnight RMSSD value, and expect the implied age to run slightly young if your device weights deep sleep — the population curve we use is built from shorter, seated recordings. If your device reports SDNN, leave the overlay off; the numbers are not comparable.

Sources

  • Shaffer, F. & Ginsberg, J. P. An overview of heart rate variability metrics and norms. Frontiers in Public Health, 2017; 5:258. Link
  • Stone, J. D. et al. Assessing the accuracy of popular commercial technologies that measure resting heart rate and heart rate variability. Frontiers in Sports and Active Living, 2021; 3:585870. Link
  • Plews, D. J. et al. Comparison of heart-rate-variability recording with smartphone photoplethysmography, Polar H7 chest strap, and electrocardiography. International Journal of Sports Physiology and Performance, 2017; 12:1324–1328. Link
  • Task Force of the European Society of Cardiology and NASPE. Heart rate variability: standards of measurement, physiological interpretation and clinical use. Circulation, 1996; 93:1043–1065. Link
Next: Measuring HRV without a wearable — a phone camera and a protocol are enough.