Every morning millions of people pick up a phone and read a two-digit number that tells them how they slept, how ready they are to train, and by implication whether today is going to be a good day. Whoop calls it Recovery. Oura calls it Readiness. Garmin calls it Body Battery. Fitbit calls it Daily Readiness. The numbers do not agree with each other and none of the vendors will tell you exactly how they are computed.
This is an honest walk-through of what is inside those scores, what makes them feel wrong on the mornings they feel wrong, and how to use them without treating a proprietary composite as a medical diagnosis.
The category is inherently opinionated
Sleep score, recovery score, readiness score, strain score, body battery — these are all composites. They take multiple sensor streams and collapse them into one number on a 0 to 100 or 0 to 20 or 0 to 5 scale. The individual sensor readings behind them are real physical measurements: heart rate, heart rate variability, respiratory rate, movement, sometimes skin temperature and blood oxygen. The score is a decision by product designers about how to weight and combine those measurements.
Two vendors making different weighting decisions will produce different scores from the same physiology. This is not a bug. It is the definition of what a composite score is. When your Whoop says 42 and your Oura says 78 on the same morning, both are correct within their own framework. Neither is telling you the truth about your body.
What is actually inside the black box
The vendors do not publish algorithms, but between filed patents, blog posts, engineering interviews, and reverse-engineering by independent researchers, the broad inputs are known. In descending order of typical weight:
- Overnight HRV — Heart rate variability during sleep is the single largest input for every score in this category. HRV is a millisecond-level fluctuation in the interval between heartbeats. Higher HRV during sleep generally reflects a body that is not fighting off illness, not carrying accumulated training stress, and not stuck in a sympathetic (fight-or-flight) mode. It is not a direct measurement of readiness for anything. It is a marker that correlates with recovery in trained athletes and with autonomic balance in general adult populations.
- Resting heart rate — Your minimum overnight heart rate, or the average during your deepest sleep window. Elevated RHR against your baseline can signal illness onset, training overload, alcohol, or heat.
- Sleep duration and continuity — Total sleep time and the number of nighttime awakenings. Longer sleep and fewer arousals generally push scores up.
- Sleep architecture (stages) — Estimated deep sleep and REM percentages. These are already noisy estimates (60 to 75 percent agreement with polysomnography, as we have written elsewhere), so they contribute but not decisively.
- Respiratory rate — Breaths per minute during sleep. Elevated respiratory rate is one of the more sensitive early markers of respiratory illness and metabolic stress.
- Prior day activity or strain — Whoop and Garmin explicitly factor yesterday's training load into today's score. Oura and Fitbit weight it less directly.
- Skin temperature deviation — Included on devices that have the sensor. Small weight, but present.
Everything above is well-established across the category. What is different between vendors is the weighting. Whoop leans heavily on HRV and on prior-day strain. Oura leans on HRV plus sleep architecture and temperature. Garmin's Body Battery emphasizes stress-derived heart rate patterns across the day, not just overnight. Fitbit's Daily Readiness weights activity load and sleep debt.
The most transparent scores in the market
Most of these vendors publish nothing. Two are more forthcoming than the rest.
Withings — of the wearables in our catalog, Withings has the most published methodological detail. Its Health Improvement Score and Vantage-related metrics have accompanying peer-reviewed validation work and a design philosophy the company will discuss on the record. This is unusual in the category and a reason the ScanWatch 2 lands at 85 Verified in our system.
Fitbit — has published more validation methodology than Whoop or Oura, though the Daily Readiness score itself remains proprietary. Fitbit Health Solutions has a research arm that has produced peer-reviewed work on sleep staging, HRV estimation, and afib detection with published sensitivity and specificity numbers.
Everyone else is essentially opaque. Whoop's marketing implies scientific rigor and publishes a handful of internal validation reports; independent peer-reviewed validation of the Recovery score specifically is thin. Oura is somewhere in between — the underlying sensors have been validated in the literature, but the Readiness algorithm itself is proprietary.
Why the number feels wrong on the mornings it feels wrong
Almost everyone who wears one of these devices has experienced this: you feel fine, you slept fine, you woke up ready to train — and the wearable tells you your recovery is 34, do not push it. Or the reverse: you feel wrecked, and the app says 91, go crush a workout. What is going on?
1. Overnight HRV overweighting
Every score in this category leans hard on overnight HRV. HRV is a real physiological signal, but it is noisier at the individual level than the marketing suggests. Studies of nightly HRV variation in healthy adults show large day-to-day swings driven by prior alcohol, prior heat exposure, prior meal size and timing, prior arousal state at bedtime, and even body position. A single low-HRV night is often not a real signal about recovery — it is a real signal about last night's dinner.
When a score puts 40 to 60 percent of its weight on a single noisy measurement, the composite inherits the noise.
2. Recency bias
Most vendor scores are built to react quickly. This is a UX decision — a score that changed once a month would feel useless. So the algorithm heavily weights the last one to three nights. A single bad night after two great weeks will crush your score in a way that does not reflect the actual state of your accumulated fitness. Two great nights after a hard month will inflate it beyond your real readiness.
3. No training-load context
Some scores factor yesterday's activity. Some do not. And even the ones that do generally do not know your training history. A CrossFit athlete's normal Monday-morning HRV is not the same as a sedentary office worker's, and a good algorithm should not treat them the same way. Most current algorithms only lightly account for this.
4. Menstrual cycle phase
Female physiology has a large, real, well-documented HRV variation across the cycle. HRV is typically higher in the follicular phase and lower in the luteal phase, and resting heart rate rises in the luteal phase. Older versions of these scores penalized women for normal luteal-phase drops. Newer versions of Oura and Fitbit have added cycle-aware adjustments; Whoop lagged and has been catching up. If your score is systematically lower in the second half of your cycle, that is a known limitation of the algorithm and not a signal about your body.
5. Position of the sensor
Photoplethysmography (the green-light optical heart sensor at the back of your watch) is easily contaminated by wrist position, tattoos, hair, cold peripheries, and how tightly the band is worn. Sleep in a cold room with the watch loose and you may lose signal quality overnight. The algorithm still gives you a score. The score is not what you think it is.
What these scores are and are not
Here is a plain framing that helps.
- A recovery or readiness score is a directional check, not a mandate. When it is high, it is a mild vote for training as planned. When it is low, it is a mild vote for pulling back a bit or paying attention.
- The trend is more real than the number. A single 42 tells you almost nothing. A 30-day decline in your average readiness from 78 to 62 while your training load stays constant is a real signal — usually of accumulating stress, illness, or life load.
- It is not a medical score. It is not calibrated against clinical outcomes. It does not diagnose overtraining. It does not diagnose depression or burnout, though it may sometimes track with them.
- Do not let it override how you feel. Subjective readiness — sleep quality, mood, motivation, soreness — is a validated construct in sports science that often outperforms wearable scores at predicting training tolerance. If your body says fine and your watch says 34, your body is usually right.
The Fitbit and Withings picks in our catalog
Two devices in our catalog give you a well-validated overnight signal without pretending the composite score is a diagnosis.
How to actually use a recovery score
Try this framework instead of the default "look at the number and react" pattern.
- Ignore the score for the first four weeks. The score needs your baseline to stabilize. Numbers before that are guesses around a moving average.
- Look at the seven-day moving average, not the daily number. Most apps let you set a rolling view. Do that. This washes out overnight HRV noise and shows the actual trend.
- Correlate the trend with your own subjective log. A cheap paper journal with three morning ratings — sleep quality, mood, motivation, each 1 to 5 — will pair usefully with the wearable trend. Over months, you will see which one leads which.
- Use it to catch the drift, not the day. The valuable use case is noticing that your average recovery has quietly slipped 15 points over three weeks while nothing has visibly changed. That is the signal that predates most illnesses and burnouts.
- Do not train solely to the score. If the algorithm becomes the coach, you have handed a proprietary black box the authority to make decisions about your body. That is not the deal.
The honest bottom line
Recovery and readiness scores are opinionated composites built by product designers to make noisy physiological data feel actionable. The inputs are real. The weighting is a business decision. The score is not a medical instrument, and no vendor in the category — including the ones we recommend — has published a peer-reviewed, prospectively validated study showing that following the score produces better outcomes than not following it.
Use them for what they are: a directional second opinion, most valuable on the monthly trend, least valuable on any single morning. Pay for the underlying sensor quality and manufacturer transparency, not for the score itself. The ScanWatch 2 and the Fitbit Sense line are the two picks in our catalog that hold up on those grounds.
Trust the seven-day average. Ignore the daily verdict. And when the watch tells you 42 and you feel fine, go train.