How Wearable Technology Is Changing Health and Fitness

Some numbers on your wrist are accurate enough to make decisions on, and some are not even close. Heart rate at rest, sleep duration and blood oxygen hold up well against laboratory equipment. Calories burned and sleep staging do not, and the gap is much larger than most people assume.

Below is what the validation literature actually reports, metric by metric, with the size of the error and a plain verdict on each.

How accurate is each wearable metric

Metric

Typical error against the laboratory reference

Source

Good enough to decide on?

Heart rate at rest

Pooled mean bias 0.21 bpm, limits of agreement -8.1 to +8.6 bpm. Mean absolute percentage error across studies 1.7% to 7.2%

Lambe et al., npj Digital Medicine, 2026 (Apple Watch, 22 studies, 1,247 people)

Yes. Treat it as a real number.

Heart rate during exercise

Pooled mean bias -0.63 bpm, limits of agreement -6.9 to +5.6 bpm. Ten of eleven studies reported error under 10%. Absolute error during activity averaged 30% higher than at rest

Lambe et al., 2026; Bent et al., npj Digital Medicine, 2020 (53 people, six devices)

Yes for steady effort. Be careful during intervals and lifting.

Step count, normal walking (4.0 to 6.4 km/h)

Wrist-worn devices 15% mean absolute percentage error. The better consumer watches sat around 4%

Mora-Gonzalez et al., IJBNPA, 2022 (CADENCE-adults, 258 adults, 21 devices)

Yes for tracking your own consistency. No for a precise daily total.

Step count, slow walking (0.8 to 3.2 km/h)

Wrist-worn devices 41% mean absolute percentage error. Averaged across all devices and wear sites, 40%

Mora-Gonzalez et al., IJBNPA, 2022

No. Slow or shuffling walking is badly undercounted.

Energy expenditure (calories burned)

Median error across seven wrist devices ranged from 27.4% to 92.6%. No device achieved error below 20%

Shcherbina et al., Journal of Personalized Medicine, 2017 (60 people, indirect calorimetry)

No. This is the weakest number your device produces.

Sleep duration (total sleep time)

Bias versus polysomnography ranged from -0.3 to +46.8 minutes across seven devices. The best wearable was +2.6 minutes, limits of agreement -42 to +47 minutes

Chinoy et al., Sleep, 2021 (34 adults, three nights each)

Yes, roughly. Good to within about half an hour on a given night.

Sleep staging (light, deep, REM)

Deep sleep detected with sensitivity 0.53 to 0.68, REM 0.49 to 0.69. Between 40% and 43% of true deep sleep was scored as light sleep

Chinoy et al., Sleep, 2021

No. Do not change anything based on one night's deep sleep figure.

Blood oxygen (SpO2)

Pooled mean bias -0.04%, limits of agreement -4.0% to +3.9%. Error grew as true oxygen saturation fell

Lambe et al., npj Digital Medicine, 2026 (9 studies, 969 people)

Yes on average, no for an individual reading. A single low value means very little.

VO2 max estimate

Exercise-based algorithms: bias -0.09 mL/kg/min, limits of agreement -9.9 to +9.7. Resting-based algorithms: bias +2.17, limits of agreement -13.1 to +17.4. Apple Watch specifically: underestimated by 6.07 mL/kg/min, mean absolute percentage error 13.3%

Molina-Garcia et al., Sports Medicine, 2022 (INTERLIVE network, 14 studies, 403 people); Lambe et al., PLoS One, 2025 (30 people)

Accurate for a population, not for you. Use the trend, not the digit.

Sources: full citations are listed at the end of this article. Where a review pooled several devices, the figure shown is the pooled result rather than the best device in the set.

Why heart rate is the metric that holds up

Your watch measures heart rate with photoplethysmography. Green light is shone into the skin, some of it is absorbed by blood, and the amount reflected back changes with each pulse of blood through the vessels in your wrist. Count those changes and you have a pulse rate.

At rest this works well. In the 2026 living systematic review of Apple Watch accuracy, which pooled 22 studies and 1,247 participants, the mean bias for resting heart rate was 0.21 beats per minute. That is effectively zero. The limits of agreement, which describe where an individual reading is likely to land, ran from about 8 beats below to 8 beats above the electrocardiogram value.

Movement is what breaks it. Bent and colleagues at Duke tested six devices against ECG in 53 people and found that absolute error during activity was on average 30% higher than at rest. Three things go wrong. The watch shifts against the skin, so the light path changes. Muscle contraction and arm swing add their own rhythm to the signal, which the algorithm can mistake for a pulse. And the cadence of your arm movement can lock onto the same frequency as your heart rate, so the device tracks your stride rather than your pulse.

This is why wrist heart rate is best during running and cycling at a steady pace, and worst during anything with irregular, high force arm movement. Weights, rowing, boxing and racquet sports are the hard cases. In the Stanford validation of seven devices, error was lowest on the cycle ergometer at 1.8% and highest during walking at 5.5%.

One finding is worth stating clearly, because it is often assumed the other way. Bent and colleagues found no statistically significant difference in accuracy across the full range of Fitzpatrick skin tones. Differences between devices were far larger than differences between skin tones.

Why calories burned is the number to ignore

Energy expenditure is the weakest metric in consumer wearables, and it is not close. The Stanford group tested seven wrist devices against indirect calorimetry in 60 adults across sitting, walking, running and cycling. Median error ranged from 27.4% for the best device to 92.6% for the worst. Not one device came in under 20%.

The 2026 Apple Watch review found the same pattern in newer hardware. Across eight studies, every study that calculated percentage error reported 20% or higher in at least one test condition, with values running from 9.7% during running to 151.7% during walking. The reviewers described the error as inconsistent and frequently large.

The larger meta-analysis by O'Driscoll and colleagues in the British Journal of Sports Medicine pooled 60 studies and 104 effect sizes. Devices tended to underestimate, error in individual studies ranged from -21% to +15%, and heterogeneity was severe. Their own conclusion was that they would be hesitant to consider any device sufficiently accurate.

The reason is structural. Your watch cannot see what you are actually burning. It infers energy cost from motion and heart rate, then runs that through a model built on a population that may not resemble you. Resting metabolic rate varies between people of the same height and weight by a meaningful margin, and that variation is baked into every estimate that follows.

The practical consequence: if you are eating to a calorie target and your watch says you burned 700 calories in a session, the honest range on that number could be 400 to 1,100. Building a daily energy deficit on top of that estimate is how people end up confused about why the scales are not moving. Set intake from body weight, activity level and measured change over time. Do not set it from what your watch says you burned.

Sleep: the duration is fine, the stages are not

These two things sit on the same screen and they deserve very different levels of trust.

Chinoy and colleagues tested seven consumer sleep trackers alongside research grade actigraphy against polysomnography, the clinical standard, in 34 healthy adults across three nights each. For total sleep time, several devices were close. The best wearable was within 2.6 minutes of the polysomnography value on average, and most devices performed as well as or better than actigraphy. Devices detected sleep very reliably, with epoch by epoch sensitivity of 0.93 or higher across the board.

The weakness is wake. Specificity, the ability to correctly identify an epoch where you were actually awake, ranged from 0.18 to 0.54. In plain terms, your device is good at knowing you were asleep and poor at knowing you were awake, which is why it tends to overstate total sleep and understate the time you spent lying there awake.

Sleep staging is a different order of problem. Sensitivity for detecting deep sleep ranged from 0.53 to 0.68 and REM from 0.49 to 0.69. The misclassification pattern is the telling part. Across devices, 40% to 43% of epochs that polysomnography scored as deep sleep were labelled light sleep by the wearable. A quarter to nearly half of true REM was also called light sleep.

So the deep sleep figure you woke up disappointed by this morning is, on the evidence, about as likely to be a labelling error as a real shortfall. Watch your sleep duration and your bed and wake times. Treat the coloured stage bars as decoration.

Step count, and what slow walking does to it

The CADENCE-adults study is the cleanest data here. Researchers put 258 adults aged 21 to 85 on treadmills and compared 21 devices against directly observed step counts across speeds from 0.8 to 8.0 km/h.

At normal walking speeds of 4.0 to 6.4 km/h, fifteen devices came in under 5% error. Wrist-worn devices as a group averaged 15%, which is worse than the ankle and thigh at 1% and the waist at 3%, though individual consumer watches did much better than the group average.

Slow walking is where it falls apart. Across all devices and wear locations, mean absolute percentage error at 0.8 to 3.2 km/h was 40%. For wrist-worn devices specifically it was 41%. That matters more than it sounds. Slow walking is what happens in a supermarket, in a hospital corridor and around an office, and it is the dominant walking speed for many older adults and for anyone recovering from injury or surgery. The people whose step counts matter most clinically are the people whose step counts are least accurate.

There is a second, quieter problem. Wrist devices count arm movement, so drumming your fingers, driving on a rough road and folding washing all add phantom steps. Fuller and colleagues, reviewing 158 publications, found wrist wearables underestimated step count by a mean of 9% and a median of 2%, with wide variation by brand. The direction of the error is not consistent enough to correct for.

VO2 max estimates and what the INTERLIVE network found

The INTERLIVE network, a consortium set up to standardise how consumer wearables get validated, published a systematic review with meta-analysis covering 14 validation studies and 403 participants.

The headline is a split. Devices that estimate VO2 max from resting data alone significantly overestimated it, by 2.17 mL/kg/min, with limits of agreement from 13.07 below to 17.41 above. Devices that use exercise data, meaning the relationship between your heart rate and your running pace, were far better calibrated, with a bias of -0.09 mL/kg/min and limits of agreement of roughly plus or minus 10.

That near zero bias looks impressive and it is easy to misread. A bias of -0.09 means the errors cancel out across a group. The limits of agreement are what apply to you as an individual, and a range of 20 mL/kg/min is enormous, roughly the difference between the 20th and the 90th percentile for a given age. The INTERLIVE authors said it directly: exercise-based estimation is suitable at the population level, but the individual error is large enough that these methods still need improvement for sport and clinical use.

A single device study puts a number on what that means in practice. Lambe and colleagues had 30 participants wear an Apple Watch for five to ten days, then measured them directly on a treadmill using the modified Astrand protocol. The watch underestimated VO2 max by an average of 6.07 mL/kg/min, with a mean absolute percentage error of 13.3% and limits of agreement running from 6.11 below to 18.26 above. The authors concluded that the estimates need refinement before they could be used clinically, while noting they remain practically useful as an alternative to conventional submaximal prediction equations.

What wearables are genuinely good at

Everything above concerns absolute accuracy, meaning how close your device gets to a laboratory measurement of the same thing at the same moment. That is the wrong question for most of what you want to do with the data.

A device with a consistent bias is still useful for detecting change. If your watch reads your resting heart rate 3 beats high every single morning, it will still show you clearly when your resting heart rate climbs 8 beats above your own normal. The bias largely cancels out of the comparison, because you are comparing yourself to yourself, on the same device, worn the same way.

This is not a consolation prize, it is how the useful applications work. Mishra and colleagues analysed smartwatch data from a cohort of nearly 5,300 people and identified 32 with confirmed COVID-19. Of those, 26 showed detectable changes in heart rate, daily steps or sleep, and 22 of the 25 with symptom data available showed those changes at or before symptom onset, four of them at least nine days earlier. The detection worked on deviations from each person's own baseline, not on any absolute threshold.

So the rule is simple. Same device, same wrist, same conditions, watch the trend over weeks. Do not compare your absolute numbers against a friend's different device, and do not compare a watch estimate against a laboratory value.

What a sensible week with a wearable looks like

Here is how the numbers above translate into actual use for someone training four times a week.

  • Monday: resting heart rate reads 54, against a 30 day average of 52. That is inside normal variation and inside the device's own measurement error. Train as planned.

  • Tuesday: sleep duration 5 hours 40 minutes, against a usual 7 hours 20. Duration is a metric you can trust to roughly half an hour, so this is a real shortfall of about 100 minutes. Keep the session, drop the top set, skip the finisher.

  • Wednesday: the watch reports 42 minutes of deep sleep, well below your usual 75. Ignore it. Between 40% and 43% of true deep sleep gets misfiled as light sleep, so a single night's staging figure carries almost no signal.

  • Thursday: heart rate during a 40 minute steady run averages 152. Trust it. Steady running is exactly the condition where wrist optical sensors perform best.

  • Friday: a heavy lower body session shows an average heart rate of 148 and 780 calories burned. Treat both sceptically. Bar work produces motion artefact at the wrist, and the calorie figure could plausibly sit anywhere between roughly 450 and 1,100.

  • Saturday: step count 4,200 on a day spent mostly walking slowly around shops. The real figure is likely higher, because slow walking is undercounted by around 40%. Do not add a compensatory session over it.

  • Sunday: resting heart rate has averaged 57 for four consecutive days, up 5 from your baseline, alongside shorter sleep. That is a genuine signal. Take an easy week.

The pattern across the week is that single readings rarely change anything, and multi day trends in the metrics that are measured well change quite a lot.

Where your own data fits

Reading a wearable well means weighting each number by how much it deserves to be trusted, then watching how it moves against your own history rather than against a population average. That is tedious to do by hand every morning. hlth. coach reads your wearable, health and body composition data and adapts your training, nutrition and recovery around the metrics that carry signal, and around your own baseline rather than someone else's.

The short version

  • Heart rate at rest is accurate. Pooled bias in the 2026 Apple Watch meta-analysis was 0.21 bpm.

  • Heart rate during steady exercise is accurate. During irregular arm movement it is not, and error during activity runs about 30% higher than at rest.

  • Calories burned is the worst metric on your device. No wrist device in the Stanford study achieved error under 20%, and the worst was 92.6%.

  • Sleep duration is measured well. Sleep staging is not, with 40% to 43% of true deep sleep labelled as light sleep.

  • Step count is reasonable at normal walking pace and poor below about 3 km/h, where error reaches 40%.

  • VO2 max from exercise-based algorithms is unbiased across a population but carries individual limits of agreement of roughly plus or minus 10 mL/kg/min. Apple Watch estimates came in 6.07 mL/kg/min low against a laboratory test.

  • The trend on your own device over weeks is worth far more than any single absolute reading.

Common questions

Are fitness trackers accurate?

It depends entirely on which metric. Resting and exercise heart rate are close to laboratory grade on modern wrist devices. Energy expenditure is not, with errors that run into double and sometimes triple digits. Sleep duration is measured well, sleep staging poorly. The table above gives each metric with its measured error.

How accurate is the Apple Watch for calories?

Poorly, and this is consistent across the literature rather than a fault of one brand. Validation studies report mean absolute percentage errors above 20 percent for energy expenditure, with some conditions far worse. Do not build a nutrition plan on the calorie figure your watch reports.

Can I trust my watch's sleep stages?

Sleep and wake detection is reliable. Staging is not. Consumer devices frequently misclassify a large share of deep sleep as light, and the specificity for detecting wake is low. Read total sleep time and consistency, and treat the deep and REM breakdown as a rough impression.

Why does my watch show different steps from my phone?

Wrist and pocket sensors see different movement, and step algorithms differ between devices. Errors are largest at slow walking speeds. If you want a step trend, pick one device and stay with it rather than reconciling two.

Can a wearable diagnose a health condition?

No. Consumer wearables are not diagnostic devices and are not regulated as such. They can flag something worth investigating, which is genuinely useful, but the diagnosis belongs to a clinician with proper equipment.

Which fitness tracker is the most accurate?

The honest answer is that the gap between brands is smaller than the gap between metrics. Heart rate is good on most current devices. Energy expenditure is unreliable on all of them. Choose on fit, battery and whether you will actually wear it.

Sources

  • Lambe R, Baldwin M, O'Grady B, Schumann M, Caulfield B, Doherty C. The accuracy of Apple Watch measurements: a living systematic review and meta-analysis. npj Digital Medicine, 2026;9:63.

  • Doherty C, Baldwin M, Keogh A, Caulfield B, Argent R. Keeping Pace with Wearables: A Living Umbrella Review of Systematic Reviews Evaluating the Accuracy of Consumer Wearable Technologies in Health Measurement. Sports Medicine, 2024;54(11):2907 to 2926.

  • Bent B, Goldstein BA, Kibbe WA, Dunn JP. Investigating sources of inaccuracy in wearable optical heart rate sensors. npj Digital Medicine, 2020;3:18.

  • Shcherbina A, Mattsson CM, Waggott D, Salisbury H, Christle JW, Hastie T, Wheeler MT, Ashley EA. Accuracy in Wrist-Worn, Sensor-Based Measurements of Heart Rate and Energy Expenditure in a Diverse Cohort. Journal of Personalized Medicine, 2017;7(2):3.

  • O'Driscoll R, Turicchi J, Beaulieu K, Scott S, Matu J, Deighton K, Finlayson G, Stubbs J. How well do activity monitors estimate energy expenditure? A systematic review and meta-analysis of the validity of current technologies. British Journal of Sports Medicine, 2020;54(6):332 to 340.

  • Chinoy ED, Cuellar JA, Huwa KE, Jameson JT, Watson CH, Bessman SC, Hirsch DA, Cooper AD, Drummond SPA, Markwald RR. Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep, 2021;44(5):zsaa291.

  • Mora-Gonzalez J, Gould ZR, Moore CC, Aguiar EJ, Ducharme SW, Schuna JM, Barreira TV, Staudenmayer J, McAvoy CR, Boikova M, Miller TA, Tudor-Locke C. A catalog of validity indices for step counting wearable technologies during treadmill walking: the CADENCE-adults study. International Journal of Behavioral Nutrition and Physical Activity, 2022;19:117.

  • Molina-Garcia P, Notbohm HL, Schumann M, Argent R, Hetherington-Rauth M, Stang J, Bloch W, Cheng S, Ekelund U, Sardinha LB, Caulfield B, Brønd JC, Grøntved A, Ortega FB. Validity of Estimating the Maximal Oxygen Consumption by Consumer Wearables: A Systematic Review with Meta-analysis and Expert Statement of the INTERLIVE Network. Sports Medicine, 2022;52(7):1577 to 1597.

  • Lambe R, O'Grady B, Baldwin M, Doherty C. Investigating the accuracy of Apple Watch VO2 max measurements: A validation study. PLoS One, 2025;20(5):e0323741.

  • Fuller D, Colwell E, Low J, Orychock K, Tobin MA, Simango B, Buote R, Van Heerden D, Luan H, Cullen K, Slade L, Taylor NGA. Reliability and Validity of Commercially Available Wearable Devices for Measuring Steps, Energy Expenditure, and Heart Rate: Systematic Review. JMIR mHealth and uHealth, 2020;8(9):e18694.

  • Mishra T, Wang M, Metwally AA, Bogu GK, Brooks AW, Bahmani A, et al. Pre-symptomatic detection of COVID-19 from smartwatch data. Nature Biomedical Engineering, 2020;4:1208 to 1220.

This article is general information, not medical advice. A wearable is not a diagnostic device, and if a reading concerns you or you have symptoms, speak to your doctor rather than your watch.

Related reading