← Journal

Same plan, two athletes, six weeks: 100% different

18 July 2026

Every training app claims it adapts to you. It's an easy claim — nobody ever checks. Granite's entire pitch is that it rewrites your programme around what actually happened, so before asking anyone to pay for that, I built a test designed to embarrass it.

The setup

Two simulated athletes. Identical profile, identical twelve-week block — byte-for-byte the same starting programme. Then six weeks of very different lives, logged session by session, with the coach reviewing each week and writing the next, exactly as it does for a real athlete.

Athlete one thrives. Hits every prescription with reps to spare, effort comfortable, energy high, nothing else competing for recovery.

Athlete two has life fall on them. Misses reps, grinds at high effort on low energy, trains their sport hard four times a week on top of the lifting, and flags a knee that's started complaining.

If the coach is real, those two people should be on different programmes by October. If it's decoration, they'll end up in roughly the same place, and the adaptive story is marketing.

The first metric was flattering, and wrong

My first instinct was to measure divergence from the original static plan. The struggling athlete's programme diverged from it almost completely — great headline. Then I measured the thriving athlete, expecting a low number, and got well over half.

The reason is mundane: the coach rebases next week on what you actually lifted, rather than following the projection drawn up on day one. Anyone who logs anything drifts from the original plan. So “diverges from static” measures that the coach recomputes — a bar every spreadsheet clears — not that it responds. A metric everyone passes is not evidence of coaching, and I nearly shipped it as proof.

The metric that matters

Compare the two athletes to each other. They started identical, so whatever separates them after six weeks is attributable to one thing: what happened to them.

The result: 100% of main-lift prescriptions differed between the two athletes by week six. Different loads everywhere; around the flagged knee, different movements entirely. And the direction was right — the struggling athlete's programme moved substantially further from the original projection than the thriving athlete's, which is what “responding” has to mean: more went wrong, more got rewritten.

What this doesn't prove

These are simulated athletes — scripted behaviour, not humans. The test proves the engine responds; it can't prove the responses are what a good coach would choose. That next question needs actual coaches, so the next step is exactly that: putting Granite's programmes in front of experienced S&C coaches, blind, and counting how often they agree with its calls.

The part I like most: this test now runs in the build pipeline. If a future change ever makes the coach stop responding to the athlete, the build fails. The central claim of the product is an assertion in the test suite, not a line of marketing.

Granite is an iOS strength coach for athletes whose sport isn't lifting. This is the sort of thing it does all day.