Methodology

How the numbers work

Every figure the app works out rather than reads, with the formula, the window and the source behind it.

Last updated

The short version. Almost everything on screen is a reading, shown as it arrived. The handful of figures the app calculates are listed below with their arithmetic in full. Three rules run through all of them: a day with no reading is absent, never zero; the app never grades you, only describes what was recorded; and it will not show a figure it cannot explain — which is why there is no sleep quality score anywhere in the app.

Not a medical device

Metrics Nest is a personal tracking and motivation tool. It does not diagnose, treat, cure or prevent any condition, and nothing in it — including the Balance score, the detail pages and any notification it sends — is medical advice. Where a figure below is compared against a healthy range, that range is broad public guidance of the kind a coach would cite, not an assessment of you. Talk to a qualified healthcare professional about anything that concerns you, and about any change to your exercise, diet or sleep. See the Terms of Use for the full statement.

Read, computed, or estimated

Every number in the app is one of three things, and the app is careful about which:

  • Read. A measurement from Apple Health, a connected source, a heart rate monitor or your phone’s GPS, shown as it arrived. Steps, heart rate, sleep stages, weight, calories logged, distance. The app converts units and nothing else.
  • Computed. Arithmetic over readings that adds no assumption — a sum, a mean, a difference, a rolling average. The window is always named on screen.
  • Estimated. A figure produced by a published formula from a population study, which is true of people in general and only approximately of you. There are four in the whole app: maximum heart rate, workout calories, estimated 1RM, and the Balance score. Each is labelled where it appears, and each is below.

Balance

Balance is the one score the app invents, and the only place it allows itself a judgement. It answers was this a good day — against published healthy marks, the same for everyone — and the app shows the whole formula on the Balance detail screen as well as here.

The weights

Four contributors, each scored 0–100, combined by weight:

ContributorWeightWhat it reads
Sleep40%Time asleep, and time in deep sleep
Food25%Energy logged, water logged, and portions of fruit and veg
Movement20%Steps, and active energy
Training15%Recorded workouts, and exercise minutes

A contributor’s score is the mean of whichever of its inputs your account actually carries. Training reads a merged list of sessions — workouts recorded in this app plus workouts your sources reported, with anything both of them saw counted once — and only counts sessions of ten minutes or more.

The marks

Each input is scored against a healthy range. Between the points below the score moves in a straight line; beyond the ends it stays flat, so nothing is extrapolated into a penalty the table does not state.

InputHealthy markThe curve, as value → score
Sleep7–9h0h→0, 4h→25, 6h→60, 7h→100, 9h→100, 11h→85
Deep sleep1h+0→0, 30m→45, 1h→85, 90m→100
Steps8–10k0→0, 3k→35, 6k→65, 8k→85, 10k→100
Active energy~500 kcal0→0, 150→40, 300→70, 500→100
Training20–40m0→40, 20m→85, 40m→100
Exercise minutes20–40m0→40, 20m→85, 40m→100
Energy logged1,500–2,800 kcal0→0, 800→40, 1,500→100, 2,800→100, 3,600→70
Water2L+0→0, 500ml→35, 1L→60, 2L→100
Fruit and veg5 portions0→25, 2→60, 5→100

Two of those are deliberately asymmetric. Training starts at 40, not 0 — a rest day is part of a healthy week, not a failure of one. And food is a wide band with gentle shoulders: undereating falls faster than overeating and neither edge reaches the floor, because judging food intake hard is how an app stops being useful and starts being a scold.

The marks come from broad public guidance — sleep-duration consensus, WHO activity minutes, published step-count studies — and they are population marks, not a clinical assessment of you. They are printed on screen precisely so they can be argued with.

Missing data, and what a score is out of

An input you have no reading for drops out of its contributor, with one exception: a day with meals logged and no fruit or veg among them counts as none of the five, not as unknown — the meals were logged, and the count is read off them. A contributor with nothing at all drops out of the total, and the remaining weights are scaled back up to 100% — so the app states the share each part actually carried rather than the share it would have. A day carrying nothing has no score rather than a zero: no score is an absence, zero is a verdict. Balance needs at least two contributors before it appears at all, because a composite of one thing is that thing wearing a composite’s clothes.

The word under the score

ScoreWord
85–100Excellent
70–84Good
55–69Fair
40–54Needs attention
below 40Low

“Up on your usual”

Beside the score, the app may say a contributor is up or down on your usual. That comparison is the mean of the last 30 days, ending the day before the day being read, and it needs at least 7 days of readings before it is said at all. A move of more than 10% either way is what gets called. None of it changes the score. An earlier version scored every day against your own trailing average, which meant somebody averaging two hours of sleep who then slept five scored a perfect 100 — improving and being well are different facts, and the app now tells both instead of letting one impersonate the other.

Sleep

Which day a night belongs to. Sources date a sleep session by when it started, which files last night under yesterday. The app uses the rule Apple and every sleep app uses: a night belongs to the day you wake up. Expressed on the start time, a session starting at or after 18:00 counts towards the next day, and one starting after midnight counts towards the day it started in. Six in the evening is the boundary because it is the only gap between the latest anyone wakes and the earliest anyone goes to bed — so an afternoon nap stays on its own day.

There is no sleep score. Stage totals and the timeline are shown as recorded. No efficiency percentage, no quality figure out of 100, no target hours on the sleep page — because there is no way to show you how such a number was reached. The app allows itself a figure it can fully explain and refuses one it cannot.

Heart

The Heart page leads with resting heart rate and heart rate variability, because those are the two that are measured overnight and mean something across weeks. The day’s average heart rate is a byproduct of how active you happened to be, so it is context underneath rather than the headline.

  • The window is 30 days, and the change reported is simply the first reading in it against the last — the same two points the chart’s own end labels show, so the figure can be checked against the picture.
  • Nothing is drawn below 7 nights of readings. Both signals move with a late meal, a warm room, a glass of wine; four points of that is a line whose direction is decided by which night happened to be last. Seven is the smallest window that spans a whole week, so weekday training and weekend late nights are compared like with like.
  • A move under 1 bpm of resting heart rate, or 2 ms of variability, reads as level — that is inside the night-to-night wobble of both.
  • The pair moving apart — resting rate down, variability up — is named as the favourable pattern. Every other case is described without being graded, because a fortnight of poor sleep, a cold and a hard training block all look identical here.

Maximum heart rate, and the zones

Zones need a ceiling to be a percentage of. If you have entered your own maximum heart rate, the app uses it and never overwrites it. Otherwise it estimates from the birth year you gave at setup:

max HR ≈ 211 − 0.64 × age

That is Nes et al. (2013), fitted on 3,320 healthy adults aged 19–89 — not the 220 − age everyone quotes, which comes from a 1971 literature review that never fitted a line to its own data. At 25 the two agree within a beat; by 55 the old one reads 165 against 176, and eleven beats is most of a zone. The estimate is clamped to ages 13–100, which is the range the formula was fitted over. Both formulas are population estimates with a spread of around 10 bpm, which is why the app never presents the result as measured and keeps it editable everywhere it appears. With no birth year and no entered value, the fallback is 190.

The five zones are the standard percentages of that ceiling:

ZoneNameOf max HR
1Warm-up50–60%
2Fat burn60–70%
3Aerobic70–80%
4Threshold80–90%
5Maximum90%+

What a workout cost

A heart rate monitor reports beats and nothing else, so calories are never in what it sent. The app answers from whichever source can answer honestly, and says which one it used:

  1. Apple Health’s active energy over the workout’s own start and end. A measurement of a sort — the phone’s accelerometry, and a watch’s heart rate where there was one. When it is a measurement of the workout, it wins: at least a kilocalorie a minute, which any real session clears and a phone lying on a bench does not. Below that, the estimate answers.
  2. An estimate, from the standard metabolic-equivalent arithmetic: kcal = (MET − 1) × 3.5 × kg ÷ 200 × minutes.

Why the minus one. MET tables describe total energy, resting included — sitting still is one MET, and your body spends that whether or not you go for a walk. Apple’s active energy is the other quantity: what the activity cost on top of being alive. Subtracting one MET makes the app’s estimate and Apple’s figure the same kind of number, where a gross figure would read about forty per cent higher for the same walk.

MET values are the Compendium of Physical Activities (Ainsworth et al.), used at the tenth they are published at — they are not this app’s opinion about anybody’s metabolism. Where your phone recorded a track, the pace picks the band, since “walking” spans a stroll with a pram and a march up a hill; where it did not, the activity’s middle band is used. The estimate needs a body mass, and returns nothing at all when the app has never been given one — a calorie figure computed from a guessed weight is exactly the invented number the rest of the app refuses.

Strength

Estimated one-rep max uses Epley: weight × (1 + reps ÷ 30), the convention every strength app shares. It is least reliable at high reps, and the answer to that is the word “est” beside it rather than a second formula. Tonnage is weight × reps × sets, and a set logged without both a weight and reps contributes to neither.

A change between two runs of the same lift is reported as arithmetic — up 2.5, held, down 2 reps — and never as an opinion about it. Whether two fewer reps at the end of a longer session is fatigue, form work or a deliberate deload is not something the app can know. Below half a kilo, a change reads as “held”: that is the smallest plate most gyms own, and rounding noise below it is not news. Nothing is called a direction from a single run.

Food

Daily totals are read, not computed. The macro ring is proportioned by energy rather than by grams, using the Atwater factors printed on every nutrition label: carbohydrate 4 kcal/g, protein 4 kcal/g, fat 9 kcal/g. 61 g of fat is more of a day’s calories than 115 g of protein, and a ring drawn by grams says the opposite.

Where a source keeps the individual entries a total was summed from, the day is grouped into morning, midday and evening — named for the clock, which is a fact, rather than for a meal, which would be a guess. What was eaten is not available from any source the app reads, so the app does not name foods.

Weight and body composition

A weight is not a fact about a day — it is a standing figure measured occasionally on a scale that disagrees with itself by more than a week of real change. So body metrics lead with a 7-day average, not the last reading, and the same average is what goal progress is measured against. The morning’s reading is still shown underneath, visibly not the headline. The tile’s trend line covers 30 days.

For the same reason, body metrics cannot be given a daily pass/fail target: day-to-day movement there is water and measurement noise, and a daily goal would manufacture failure out of nothing.

“These two line up”

On a tile combining two metrics, the app may say they move together. That is Pearson’s correlation coefficient over the days where both were recorded, with three gates:

  • Fewer than 10 paired days — no claim at all, just how many more it needs.
  • |r| < 0.3 — it says the two have not moved together, rather than reaching for a story.
  • |r| ≥ 0.3 moderate, ≥ 0.6 clear — it states the direction, and the typical size of the difference where the units allow.

The wording is always “line up with” or “tend to”, never “causes” or “because”. A correlation on a phone’s worth of data is an observation about two lines, not a finding about your body.

Days, windows and averages

Everything above sits on one daily series, built the same way for every metric:

  • Several readings on one day collapse to one, by the rule that suits the metric: steps and calories sum, heart rate averages, a resting low takes the minimum, a weight takes the latest.
  • A day with no reading is null, not zero. It is left out of averages rather than dragging them down, and no bar is drawn for it. A dashboard that says “no data” is the app reporting a fact, not failing.
  • A rolling 7-day average is the mean of the days that carry a reading inside the window — not a mean over seven slots with the gaps counted as nothing.
  • Every window is named where it is used. “30-day average” means the last thirty days including today, and the screen says so.

Goals that unlock a permanent milestone always read the 7-day rolling average rather than a single reading, so one noisy weigh-in can never mint an achievement that cannot be taken back.

What the app will not do

  • Show a score whose formula is not on a screen you can reach.
  • Grade you. Colour is reserved for physiological state, never for approval or disapproval, and there are no failure streaks.
  • Render a missing reading as a zero.
  • Draw a trend from fewer readings than the trend needs.
  • Name a cause for anything it observes.
  • Present an estimate as a measurement, or a population figure as a fact about you.

Questions about any of this

If a number here disagrees with what the app shows you, that is a bug and worth an email: hello@metricsnest.app. If you think a mark is set wrong, that is worth an email too — they are published so they can be argued with.