Most biological-age tests read DNA methylation patterns and feed them into a trained model, and the training target defines what the score means. First-generation clocks, beginning with Horvath's 2013 multi-tissue clock, were trained to predict calendar age. Second-generation clocks such as PhenoAge, published in 2018, and GrimAge, published in 2019, were trained against clinical biomarkers and mortality. Pace-of-aging measures such as DunedinPACE, published in 2022, were trained on two decades of within-person physiological decline in a single New Zealand birth cohort.
Because the training targets differ, the scores answer different questions, and the same person can receive meaningfully different results from different tests without either being wrong. A report that does not say which clock it runs is asking you to trust a number it has not defined.
At the group level, the signal is credible. Accelerated epigenetic age on second-generation clocks, particularly GrimAge, has been associated with higher mortality and disease risk across multiple cohorts since Lu and colleagues published it in 2019. People whose scores run older than their calendar age are, on average, at higher risk than people whose scores run younger.
On average is the operative phrase. Group-level association does not become an individual prediction, and no clock has been validated as a tool for telling one person how long they will live or which disease they will get. That is not what these models were built to do.
There is also a plainly technical limit that consumer reports rarely disclose. In 2022, Higgins-Chen and colleagues showed that re-processing the same blood sample could shift some widely used clock estimates by several years due to laboratory and array noise alone, and proposed principal-component versions of the clocks that substantially improved test-retest reliability. Unless a report states which version a laboratory runs, a few years of apparent aging or de-aging between two tests may be measurement noise rather than biology.
The best randomized evidence so far is modest. In the two-year CALERIE trial of caloric restriction in healthy adults, an analysis published in 2023 by Waziry and colleagues found that DunedinPACE slowed by roughly 2 to 3 percent in the restriction group, while PhenoAge and GrimAge showed no significant change. One trial, one population, a small effect on one measure: enough to make the field interesting, nowhere near enough to make any score a validated outcome for judging whether an intervention is working in you.
- There is no single definition of biological age; every test operationalizes it differently.
- Different clocks can produce different estimates for the same person, by design.
- Technical noise can move some scores by years, so small changes are often meaningless.
- A change in score does not prove which behavior or intervention caused it.
- A younger score does not rule out disease, and no score predicts an individual lifespan.
Used well, a biological-age score is a conversation starter and, with a consistent method over time, a rough trend signal. Keep the same test and laboratory, ignore small movements, never let the score override established risk assessment, and treat marketing about reversing your age as exactly that. At The Maximum Life, biological age and Systems Age measures appear inside a broader assessment when they help answer a clinical question; they are never sold as a standalone verdict, because that is not what the evidence supports.