How we grade
How we grade the evidence.
Every tier on this site comes from fixed rules, applied the same way to every watch. This page is generated from those rules, so it says exactly what they do. These are rules version r1, in force since .
What a tier means
Each watch gets a tier for each thing it measures, from the published studies of that exact model.
- Validated
- At least one top-grade (A) study meets our threshold, at least one independently funded A or B study meets it, and no A-grade study misses it.
- Limited
- There is usable evidence, but not enough to call it validated: a single study, a smaller or weaker one, maker funding, or evidence carried over from an earlier model.
- Contested
- At least two good (A or B) studies disagree: one meets the threshold and one misses it. We show both and never average them away.
- Unvalidated
- No usable published test. That's not proof a watch is poor; there's no evidence either way.
- Not covered
- Something we don't cover. We never grade medical measures such as ECG, blood oxygen or blood pressure.
How a study is graded
- A
- A criterion reference (an ECG chest strap for heart rate, polysomnography for sleep, indirect calorimetry for energy, a metabolic cart for fitness estimates, a surveyed course for GPS), at least 20 people, peer-reviewed, a stated protocol and a standard error statistic.
- B
- As A, except one of: fewer than 20 people (but at least 10), a preprint, or another device used as the reference.
- C
- Data are reported, but something A needs is missing: no reference standard, an unclear protocol, no standard error statistic, fewer than 10 people, or neither peer-reviewed nor a preprint.
- D
- Abstract only, or methods not reported. Recorded, never used for a tier.
Each finding uses its own reference and sample size where the study reports them separately.
What counts as accurate enough
A finding meets our threshold when it shows:
- Heart rate at rest
- an error (mean absolute percentage error) of 10% or less.
- Heart rate in steady exercise
- an error (mean absolute percentage error) of 10% or less.
- Heart rate in interval exercise
- an error (mean absolute percentage error) of 10% or less.
- Step count
- an error (mean absolute percentage error) of 10% or less.
- Energy expenditure
- an error (mean absolute percentage error) of 20% or less.
- Sleep duration
- an average bias within 30 minutes, either way.
- Sleep staging
- agreement on stages of kappa 0.6 or more, or epoch-by-epoch agreement of 75% or more.
- VO2max estimate
- an error (mean absolute percentage error) of 10% or less.
- GPS distance and pace
- an error (mean absolute percentage error) of 5% or less.
Correlation or intraclass correlation alone doesn't show agreement, so a finding that reports only those counts as usable evidence but never as meeting or missing a threshold. These are starting values; any change is a new rules version.
Independence
Only a study that was independently funded, and that no brand commissioned, counts towards Validated. A study a brand commissioned is labelled as such wherever it appears and graded like any other.
Evidence on one model is shown on another only when the maker states the sensor and algorithm are unchanged and a person has confirmed it. It is labelled, and it caps the tier at Limited.
We're paid by retailer commission, never by brands for placement. Affiliate links are labelled Ad, at or before the link.
How fresh each fact is
- Prices and stock
- Current for up to 6 hours after we check. A price we checked more than 6 hours ago is shown as last seen, not as current.
- Phone compatibility
- Up to 7 days, and out of date at once when a new phone operating system comes out. Until we check again, the checker says we don't know.
- Specifications and subscriptions
- Checked at least every 30 days; always shown with the date.
- Studies
- We check each study hasn't been retracted at least every 180 days. When a study is retracted, we remove it from the evidence and the tiers.
How evidence is weighed for a purpose
When we choose watches for you, each purpose weighs the measurements differently. This is the table selection reads (weights version w1). It never includes anything a retailer pays.
For each measurement, the watch's tier counts as Validated 1, Limited 0.5, Contested 0.25 and Unvalidated 0, times the weight below. Some purposes add a bonus for the maker's specification, up to the weight shown, shared between the rows named.
| Purpose | Measurements and weights | Specification bonus |
|---|---|---|
| Running |
|
None |
| Cycling |
|
None |
| Swimming |
|
Up to 2, from a water rating (0 to 10 ATM) and pool swim tracking |
| Gym and classes |
|
None |
| Sleep |
|
None |
| General fitness |
|
None |
| An everyday smartwatch |
|
Up to 2, from how fully it works with your phone (1 to 2 features) and battery life (1 to 14 days) |
| Outdoors and navigation |
|
Up to 2, from battery life (1 to 14 days) and offline maps |
How the three watches are chosen
Selection is fixed rules, not a judgement call: the same answers and the same register always give the same three watches. These are selection rules s1.
- First, what's ruled out
- A watch without a current, sourced check that it fully pairs with your phone. One whose price is outside your budget: we use the middle of the prices we checked in the last few hours at shops with it in stock, or the maker's price where no shop has it; never an out-of-date price. One that can't meet your battery answer (we use the lower of the maker's claim and any independent measurement), your swimming answer (a water rating of 5 ATM or more, or a maker's rating for diving) or a must-have. And, if you'd rather not pay a subscription, one that puts a feature your purpose depends on behind one.
- Our pick
- The highest score. A tie goes to the watch with more Validated measurements, then the lower price, then a fixed order. If no watch that fits has any accuracy evidence for your purpose, we don't pick one.
- A good alternative
- The best watch from a different maker, if it scores within 15% of our pick. Otherwise the next best that has something sourced to say for it.
- The popular choice we'd skip
- The best-selling of the remaining watches that fit your answers (by the sales rank in a retailer feed), if it scores no more than 75% of our pick's score and both watches have their own studies of the same measurement, reported the same way, so the comparison is like for like. If none qualifies, we name none and say so.
- Reasons
- Each card shows the two or three rows that counted most, each with its tier and the date of its source. A written summary, where there is one, is checked sentence by sentence against those rows before it is shown.
- s1, : First selection rules: the filters, weights, pick, alternative and skip rules as set out on this page.
Versions
Any change to a grade, a threshold or a tier rule is a new rules version, listed here with what changed.
- r1, : First rules: study grades A to D, the thresholds per metric, the tiers and the freshness limits as set out on this page.