A Race to Get Younger: The Biological De-Aging Competition Explained
Roughly 500 people have signed up for a six-month experiment in which the goal is not to lose weight, run faster, or build muscle, but to get biologically younger. The contest, announced this week and covered by MIT Technology Review's biotech newsletter The Checkup, assigns participants the task of reducing their biological age through a range of interventions, then tracks their progress on a public leaderboard. The journalist Jessica Hamzelou, who recently turned 40, is among the entrants. Her chronological age will keep ticking forward regardless of what the scoreboard says. What she and the other competitors are chasing is a different number: an estimate of how well their organs and bodies are functioning relative to the typical health trajectory associated with their years.
The premise is provocative because it inverts the usual framing of aging research. Most clinical work asks how to slow decline, or how to extend the period of life spent in good health. A contest asks something blunter: can you move the needle backward within a single human year, let alone six months? The leaderboard format adds competitive pressure, but it also introduces a measurement problem. To rank competitors, you need a score. And the score, in this case, is a biological age estimate, a figure derived from biomarkers rather than from a birth certificate.
Around 500 entrants is a small cohort by the standards of a randomized trial, and the six-month window is short. That does not make the exercise worthless. It does mean the results will be suggestive rather than definitive, a distinction the longevity field has learned to make the hard way.
The Science Behind Reversing Biological Age
Biological age measurement rests on the idea that two people born on the same day can be aging at different rates, and that the difference shows up in molecular marks. The most widely studied of these marks are DNA methylation patterns, chemical tags attached to DNA that change predictably across the lifespan. In 2013, Steve Horvath, then at UCLA, published work on a multi-tissue epigenetic clock that could estimate age from methylation data with striking accuracy across many tissue types. Later clocks, including the PhenoAge and GrimAge models, were trained not just on chronological age but on health outcomes, making them better at predicting mortality and disease risk than a simple birthday count. The research literature now includes dozens of such clocks, validated in cohorts numbering in the tens of thousands.
Read next Laika's Wildwood: Stop-Motion Fantasy at TIFF 2026That foundation is what makes a de-aging contest conceivable. If methylation patterns shift with age, the logic goes, interventions that alter those patterns might plausibly shift the estimate. Some small trials have reported reductions in epigenetic age after intensive lifestyle programs combining diet, exercise, sleep, and stress management. Other studies have reported changes after pharmaceutical or supplement regimens. The effect sizes in these trials are typically modest, measured in a few years of estimated age, and often not replicated at scale. A six-month contest among 500 self-selected participants will generate anecdotes and perhaps a signal, but it will not resolve whether biological age reversal is durable, or whether it matters for how long someone lives.
The stakes go beyond a leaderboard. If biological age can be reliably moved, it becomes a surrogate endpoint for trials of drugs intended to prevent age-related disease. That would let researchers test interventions in months rather than decades. The FDA has not approved any epigenetic clock as a primary endpoint for a longevity drug, and the reasons are scientific, not bureaucratic caution alone.
The Controversy Around Biological Age Clocks
Here is the central problem: a clock is a statistical model, not a direct readout of aging. It learns from training data, and its predictions depend on which population that data came from, which tissues were sampled, and which algorithm was used. Horvath's original clock and its successors can disagree with one another about the same person. A 2023 analysis in a peer-reviewed geroscience journal found that different epigenetic clocks often show only moderate correlation when applied to the same samples, and that some clocks are better at predicting chronological age while others are better at predicting disease. Those are not the same task.
There is also the question of what a change in clock value actually means. If an intervention lowers your estimated biological age by three years, have you gained three years of healthy life? Not necessarily. The clock might be responding to something incidental, like a shift in white blood cell composition, rather than to a fundamental change in tissue function. Researchers at institutions including Harvard, Yale, and the National Institute on Aging have raised this concern repeatedly: an epigenetic age estimate is a biomarker, and a biomarker is only as good as the outcome it is validated against.
Then there is the consumer layer. Direct-to-consumer biological age tests have proliferated, priced from roughly $200 to several hundred dollars, with results that can swing widely between providers. Some longevity clinicians use them to guide personalized protocols. Skeptics, including several prominent geroscientists, argue that the tests are being marketed ahead of the evidence, and that a leaderboard contest risks turning a research question into a game with winners and losers but no clear clinical meaning. Hamzelou's own reporting on the contest acknowledges this tension directly. The honest position is that biological age reversal is plausible in principle, demonstrable in some small trials, and unproven as a route to longer, healthier life at population scale.
Opinion: Why LLMs Don't Actually Reason
In a separate piece published alongside the de-aging story, Thore Graepel, a researcher who helped build a program that stunned the world by beating a Go champion a decade ago, argues that large language models do not reason. The claim is not that LLMs are useless. It is that the word "reasoning" has been applied to them too loosely, and that the looseness has consequences for how the technology is deployed and regulated.
The evidence for the skeptical view is substantial. Peer-reviewed studies have shown that LLMs can fail at multi-step logical problems that a careful human solves easily, and that small changes in wording, irrelevant context, or the order of premises can flip their answers. Models that appear to perform chain-of-thought reasoning often produce correct answers through pattern matching on training data rather than through valid inference. On standard benchmarks, performance can drop sharply when a task is perturbed slightly from the format seen during training: rename the entities in a logic puzzle, or add a distractor clause, and accuracy can fall by double-digit percentages. A model may appear to reason its way to an answer while actually retrieving a memorized association.
This does not mean the systems cannot be useful for tasks that resemble reasoning. They can summarize, translate, draft, and generate code that passes tests. What Graepel's argument targets is the inference from fluent output to genuine understanding. A system that predicts the next token well enough to sound like a reasoner is not the same as a system that represents the world and draws valid conclusions about it. The distinction matters most where stakes are high: medicine, law, scientific discovery, and any setting where a plausible-sounding wrong answer can cause harm.
The counterargument, made by some AI researchers, is that the line between pattern matching and reasoning is blurrier than critics admit, and that human cognition itself relies heavily on statistical regularities. That debate is not settled. What is settled is that current benchmarks do not cleanly measure reasoning, and that marketing language has outrun the evidence.
What These Two Stories Tell Us About the Limits of Modern Science and AI
Both stories share a common shape: a measurement is treated as a proxy for something deeper, and the proxy is then mistaken for the thing itself. In the de-aging contest, an epigenetic clock stands in for the biological process of aging. The clock is a real, useful, imperfect instrument. But a leaderboard that ranks people by a clock is ranking them by a statistic, and the statistic can move without the underlying biology moving in the way competitors hope. The same pattern appears in AI. A benchmark score stands in for reasoning. The score is real. The reasoning it is supposed to measure is not necessarily present.
The responsible response in both cases is the same discipline. Ask what the measurement was validated against. Ask whether it generalizes beyond the population and conditions in which it was derived. Ask whether a change in the number corresponds to a change in the outcome anyone actually cares about. For biological age, that means large, long, controlled trials linking clock changes to disease and mortality, not a six-month scoreboard. For LLMs, it means task-specific evaluation on problems the models have not seen, with adversarial perturbations, not headline benchmark numbers.
Neither field is fraudulent. Both are early, exciting, and prone to overinterpretation. The de-aging contestants will produce data, some of it interesting. The AI systems will keep improving at tasks that look like reasoning. The job of careful readers, and of the journalists and researchers who serve them, is to keep the distinction between a number and the reality it represents clearly in view.
Source: MIT Technology Review



