smartwatch

(Credit: Photo by Luke Chesser)

In a Nutshell

  • Readiness and recovery scores blend several signals (resting heart rate, heart rate variability, sleep, activity, training load) into one number, but a review in Sensors found these proprietary scores lack an agreed-upon standard and remain largely unvalidated.
  • Because each brand uses its own private formula, similar body data can produce different readiness verdicts from one watch to another, with no way to check which is right.
  • Many everyday metrics — calorie burn, sleep stages, fitness level, blood pressure, glucose — are algorithm-based estimates, not direct measurements, and their accuracy varies widely by metric and device.

Millions of people glance at their wrist each morning and let a single number, something like a “readiness” or “recovery” score, decide whether to push through a hard workout or take it easy. That number feels scientific. It sits next to a heart rate reading and a sleep chart, dressed up in clean graphics and confident language. A new review of smartwatch research suggests that confidence may not be earned. Many of these composite scores have never been independently checked against real medical measurements, and different brands can turn the same body data into completely different verdicts about how “ready” someone is.

Researchers from the University of Michigan and Hope College combed through the science behind consumer smartwatches, examining what these devices actually measure versus what they merely estimate. Their conclusion, published in the journal Sensors, is blunt: most of the numbers on a smartwatch screen are not direct readings of what’s happening inside the body. They are modeled predictions built from sensor signals, private math formulas, and assumptions about the person wearing the device. Some of those predictions hold up reasonably well. Others, especially the flashy composite scores meant to sum up a person’s health in one tidy figure, rest on far shakier ground.

Nearly one in three Americans report regularly using a wearable device, according to figures cited in the review. Athletes plan training around recovery scores. People manage chronic conditions using sleep and activity trends. With that much daily decision-making riding on these gadgets, it matters what the numbers actually mean, and where they fall short.

What a Smartwatch Readiness Score Actually Measures

Composite scores marketed under brand-specific names try to answer a simple question: how prepared is a person for the day’s physical or mental demands? To generate that answer, the software blends overnight resting heart rate, heart rate variability (how much the time between heartbeats fluctuates), sleep duration, recent activity, and training load into one combined score, then sorts it into color-coded zones like “low,” “moderate,” or “high.”

Trouble starts with what happens behind that math. According to the review, a systematic evaluation of 14 composite health scores across 10 consumer wearable manufacturers found that proprietary weighting schemes “generally lack published validation” and that theoretical rationale is often emphasized over actual evidence. In plain terms, companies explain why their formula should work, but rarely publish proof that it does. Because each manufacturer weighs the ingredients differently and keeps that formula private, similar body data fed into two different watches could produce two different readiness verdicts. No agreed-upon standard exists to check these scores, so researchers currently have no reliable way to say whether a given score reflects true recovery.

How a Smartwatch Turns Raw Data Into a Score

To understand why readiness scores sit on such shaky footing, it helps to know how smartwatches gather information in the first place. Inside the device, motion sensors track movement, GPS chips estimate location and distance, and optical sensors shine light into the skin to detect blood flow changes tied to heartbeats. None of these sensors directly measure things like calories burned, sleep stage, or fitness level. Instead, a chain of algorithms converts raw signals into a final number that appears on the screen.

That chain introduces multiple points where accuracy can slip. Heart rate readings tend to hold up fairly well at rest and during steady exercise, but errors generally increase with movement and rapidly changing exercise intensity. Calorie-burn estimates fare worse, with errors frequently topping 10% to 20% because they depend on a stack of assumptions about body weight, activity type, and resting metabolism. Fitness-level estimates, meant to approximate a lab-measured test of how well the body uses oxygen during exercise, perform reasonably well for casual exercisers but become less accurate for highly trained athletes.

Sleep tracking shows the same pattern of strength and weakness. Watches are quite good at telling whether someone is asleep or awake, matching lab-based brain-wave testing with high agreement. Sorting sleep into specific stages such as light, deep, or REM is tougher, with accuracy ranging roughly from 50% to 86% depending on the device, and a consistent tendency to mistake quiet wakefulness for light sleep.

Infographic comparing the accuracy and limitations of smartwatch health metrics, including heart rate, sleep, calories, readiness, VO2 max, blood pressure, and glucose
Infographic by StudyFinds

Where Smartwatch Health Tracking Falls Apart Fastest

Some newer smartwatch features push even further into territory where the science hasn’t caught up. Cuffless blood pressure monitoring, which estimates pressure from how quickly a pulse wave travels through the body rather than measuring pressure directly, has shown inconsistent results across studies. Medical groups have taken notice: a 2025 American Heart Association scientific statement and updated national blood pressure guidelines both advise against using cuffless devices to diagnose or manage high blood pressure, even as one wrist-worn system received limited regulatory clearance for home use with frequent recalibration required.

Blood sugar tracking sits even further behind. No smartwatch currently on the market has been authorized to measure glucose levels, and the Food and Drug Administration has explicitly warned people not to rely on smartwatches for that purpose.

Skin tone adds another layer of concern. Because optical sensors rely on light passing through skin, the amount of pigment in a person’s skin could distort readings, which may contribute to heart rate underestimation and oxygen level overestimation in people with darker skin. That pattern prompted the FDA to issue updated 2025 guidance requiring companies to report how their oxygen sensor results vary across skin tone groups during testing.

What This Means for Smartwatch Habits

None of this means smartwatches are worthless. Researchers point out that these devices are genuinely good at one thing: tracking how a single person’s own numbers shift over weeks and months. A gradual rise in resting heart rate, a steady decline in sleep duration, or a drop in daily step count can flag meaningful change even if the exact figures behind them carry some built-in error. Watches also let people gather information in daily life rather than only during a doctor’s visit, giving them an ongoing snapshot that traditional checkups can’t capture.

Trouble arises when those same numbers get treated as precise medical facts rather than rough estimates. Comparing one person’s calorie burn to another’s, using a sleep score to diagnose insomnia, or trusting a readiness score to make a serious training or health decision goes beyond what the current technology can reliably support. The review’s authors argue that smartwatch data should support lab tests, doctor visits, and self-reported information, rather than stand in for them.

Smartwatches aren’t lying to their wearers, but they are simplifying a lot of complicated biology into a single glanceable number, and that simplification comes at a cost. Anyone who checks a readiness score before deciding whether to skip a workout is trusting a formula that companies have not shown actually predicts anything meaningful. Until independent testing and more transparency arrive, the smartest use of a smartwatch may be watching the trend line, not chasing the daily score.

Disclaimer: This article summarizes findings from a narrative review published in the journal Sensors. A narrative review synthesizes existing research rather than conducting new experiments or pooled statistical analysis, so its conclusions reflect the authors’ interpretation of the available literature. Smartwatch technology, validation evidence, and regulatory status change quickly, and details described here may shift over time. This content is for general informational purposes only and is not medical advice. It should not be used to diagnose, treat, or manage any health condition. Anyone with questions about their health, or about a specific device or metric, should consult a qualified healthcare professional.

Paper Notes

Limitations

The authors describe this work as a narrative review rather than a systematic one, meaning the literature search was not exhaustive and no formal quality scoring or pooled statistical analysis was performed. They note that the wearable technology literature varies widely in device brand, software version, reference standards, populations studied, and testing conditions, which limits direct comparisons across studies. Proprietary algorithms and incomplete company reporting also prevented the authors from fully tracing how raw sensor signals get converted into the numbers users see. Because smartwatch technology changes quickly, the authors caution that validation findings, regulatory approvals, and manufacturer practices described in the paper may shift after publication, and the review should be treated as a framework for evaluating devices rather than a fixed ranking of specific products.

Funding and Disclosures

The authors state that the research received no external funding and report no conflicts of interest. The paper notes that generative AI tools (Google Gemini and Claude) were used to refine drafts, condense information into a table, and support editorial clarity, and the authors state they reviewed and take full responsibility for the final content. A figure in the paper was created using BioRender, including AI-assisted layout tools within that platform, which the authors also reviewed and approved.

Publication Details

Paper Title: “Consumer Smartwatch Technology in Health and Performance Research: Validity, Limitations, and Real-World Applications”

Authors: Adam S. Lepley, Fiddy Davis, Amanda C. Melvin, and Zheng-Yang Zhao. Lepley, Melvin, and Zhao

Author Affiliations: School of Kinesiology at the University of Michigan in Ann Arbor. Davis is affiliated with the Department of Kinesiology, Division of Social Sciences, at Hope College in Holland, Michigan.

Journal: Sensors (2026, Volume 26, Article 4486)

DOI: 10.3390/s26144486

About StudyFinds Analysis

Called "brilliant," "fantastic," and "spot on" by scientists and researchers, our acclaimed StudyFinds Analysis articles are created using an exclusive AI-based model with complete human oversight by the StudyFinds Editorial Team. For these articles, we use an unparalleled LLM process across multiple systems to analyze entire journal papers, extract data, and create accurate, accessible content. Our writing and editing team proofreads and polishes each and every article before publishing. With recent studies showing that artificial intelligence can interpret scientific research as well as (or even better) than field experts and specialists, StudyFinds was among the earliest to adopt and test this technology before approving its widespread use on our site. We stand by our practice and continuously update our processes to ensure the very highest level of accuracy. Read our AI Policy (link below) for more information.

Our Editorial Process

StudyFinds publishes digestible, agenda-free, transparent research summaries that are intended to inform the reader as well as stir civil, educated debate. We do not agree nor disagree with any of the studies we post, rather, we encourage our readers to debate the veracity of the findings themselves. All articles published on StudyFinds are vetted by our editors prior to publication and include links back to the source or corresponding journal article, if possible.

Our Editorial Team

Steve Fink

Editor-in-Chief

John Anderer

Associate Editor