TechSussd

How we grade

How we grade the evidence.

Every tier on this site comes from fixed rules, applied the same way to every watch. This page is generated from those rules, so it says exactly what they do. These are rules version r1, in force since .

What a tier means

Each watch gets a tier for each thing it measures, from the published studies of that exact model.

Validated
At least one top-grade (A) study meets our threshold, at least one independently funded A or B study meets it, and no A-grade study misses it.
Limited
There is usable evidence, but not enough to call it validated: a single study, a smaller or weaker one, maker funding, or evidence carried over from an earlier model.
Contested
At least two good (A or B) studies disagree: one meets the threshold and one misses it. We show both and never average them away.
Unvalidated
No usable published test. That's not proof a watch is poor; there's no evidence either way.
Not covered
Something we don't cover. We never grade medical measures such as ECG, blood oxygen or blood pressure.

How a study is graded

A
A criterion reference (an ECG chest strap for heart rate, polysomnography for sleep, indirect calorimetry for energy, a metabolic cart for fitness estimates, a surveyed course for GPS), at least 20 people, peer-reviewed, a stated protocol and a standard error statistic.
B
As A, except one of: fewer than 20 people (but at least 10), a preprint, or another device used as the reference.
C
Data are reported, but something A needs is missing: no reference standard, an unclear protocol, no standard error statistic, fewer than 10 people, or neither peer-reviewed nor a preprint.
D
Abstract only, or methods not reported. Recorded, never used for a tier.

Each finding uses its own reference and sample size where the study reports them separately.

What counts as accurate enough

A finding meets our threshold when it shows:

Heart rate at rest
an error (mean absolute percentage error) of 10% or less.
Heart rate in steady exercise
an error (mean absolute percentage error) of 10% or less.
Heart rate in interval exercise
an error (mean absolute percentage error) of 10% or less.
Step count
an error (mean absolute percentage error) of 10% or less.
Energy expenditure
an error (mean absolute percentage error) of 20% or less.
Sleep duration
an average bias within 30 minutes, either way.
Sleep staging
agreement on stages of kappa 0.6 or more, or epoch-by-epoch agreement of 75% or more.
VO2max estimate
an error (mean absolute percentage error) of 10% or less.
GPS distance and pace
an error (mean absolute percentage error) of 5% or less.

Correlation or intraclass correlation alone doesn't show agreement, so a finding that reports only those counts as usable evidence but never as meeting or missing a threshold. These are starting values; any change is a new rules version.

Independence

Only a study that was independently funded, and that no brand commissioned, counts towards Validated. A study a brand commissioned is labelled as such wherever it appears and graded like any other.

Evidence on one model is shown on another only when the maker states the sensor and algorithm are unchanged and a person has confirmed it. It is labelled, and it caps the tier at Limited.

We're paid by retailer commission, never by brands for placement. Affiliate links are labelled Ad, at or before the link.

How fresh each fact is

Prices and stock
Current for up to 6 hours after we check. A price we checked more than 6 hours ago is shown as last seen, not as current.
Phone compatibility
Up to 7 days, and out of date at once when a new phone operating system comes out. Until we check again, the checker says we don't know.
Specifications and subscriptions
Checked at least every 30 days; always shown with the date.
Studies
We check each study hasn't been retracted at least every 180 days. When a study is retracted, we remove it from the evidence and the tiers.

How evidence is weighed for a purpose

When we choose watches for you, each purpose weighs the measurements differently. This is the table selection reads (weights version w1). It never includes anything a retailer pays.

For each measurement, the watch's tier counts as Validated 1, Limited 0.5, Contested 0.25 and Unvalidated 0, times the weight below. Some purposes add a bonus for the maker's specification, up to the weight shown, shared between the rows named.

Weights per purpose
PurposeMeasurements and weightsSpecification bonus
Running
  • Heart rate in steady exercise: 3
  • Heart rate in interval exercise: 3
  • GPS distance and pace: 3
  • VO2max estimate: 2
  • Energy expenditure: 1
None
Cycling
  • Heart rate in steady exercise: 3
  • Heart rate in interval exercise: 2
  • GPS distance and pace: 2
  • Energy expenditure: 1
None
Swimming
  • Heart rate in steady exercise: 2
Up to 2, from a water rating (0 to 10 ATM) and pool swim tracking
Gym and classes
  • Heart rate in interval exercise: 3
  • Energy expenditure: 2
None
Sleep
  • Sleep duration: 3
  • Sleep staging: 3
  • Heart rate at rest: 2
None
General fitness
  • Step count: 3
  • Heart rate at rest: 2
  • Energy expenditure: 2
  • Sleep duration: 1
None
An everyday smartwatch
  • Step count: 1
  • Heart rate at rest: 1
Up to 2, from how fully it works with your phone (1 to 2 features) and battery life (1 to 14 days)
Outdoors and navigation
  • GPS distance and pace: 3
Up to 2, from battery life (1 to 14 days) and offline maps

How the three watches are chosen

Selection is fixed rules, not a judgement call: the same answers and the same register always give the same three watches. These are selection rules s1.

First, what's ruled out
A watch without a current, sourced check that it fully pairs with your phone. One whose price is outside your budget: we use the middle of the prices we checked in the last few hours at shops with it in stock, or the maker's price where no shop has it; never an out-of-date price. One that can't meet your battery answer (we use the lower of the maker's claim and any independent measurement), your swimming answer (a water rating of 5 ATM or more, or a maker's rating for diving) or a must-have. And, if you'd rather not pay a subscription, one that puts a feature your purpose depends on behind one.
Our pick
The highest score. A tie goes to the watch with more Validated measurements, then the lower price, then a fixed order. If no watch that fits has any accuracy evidence for your purpose, we don't pick one.
A good alternative
The best watch from a different maker, if it scores within 15% of our pick. Otherwise the next best that has something sourced to say for it.
The popular choice we'd skip
The best-selling of the remaining watches that fit your answers (by the sales rank in a retailer feed), if it scores no more than 75% of our pick's score and both watches have their own studies of the same measurement, reported the same way, so the comparison is like for like. If none qualifies, we name none and say so.
Reasons
Each card shows the two or three rows that counted most, each with its tier and the date of its source. A written summary, where there is one, is checked sentence by sentence against those rows before it is shown.
  1. s1, : First selection rules: the filters, weights, pick, alternative and skip rules as set out on this page.

Versions

Any change to a grade, a threshold or a tier rule is a new rules version, listed here with what changed.

  1. r1, : First rules: study grades A to D, the thresholds per metric, the tiers and the freshness limits as set out on this page.