A Race to Turn Back the Clock: The Biological De-Aging Contest
On a Thursday morning in early October, roughly 500 people logged on to a leaderboard that measures how quickly they are growing younger. Among them is Jessica Hamzelou, a biotech journalist who recently turned 40 and who this week officially entered a six-month competition built around a single, disputed number: her biological age. The rules are simple to state and hard to execute. For half a year, participants will attempt biological age reversal using an assortment of interventions, and the scoreboard will track their progress against one another.
The contest arrives at a moment when longevity research has moved from the fringes of wellness culture into serious laboratory science. Chronological age counts birthdays. Biological age attempts to capture something closer to the operational wear on a body — how well its organs, cells, and regulatory systems are actually functioning relative to calendar time. Those two numbers often diverge. Two people born on the same day can differ by a decade in physiological decline.
That divergence is precisely what makes a competition like this possible, and precisely what makes it contentious. If biological age were a fixed, unassailable quantity, a leaderboard would be straightforward. It is not. The measurement tools that underpin this contest are powerful, promising, and still contested — a tension Hamzelou's reporting for MIT Technology Review's The Checkup newsletter examines directly, and one she intends to test on her own body.
The Science and Controversy Behind Biological Age Metrics
The most widely cited tools for estimating biological age are epigenetic clocks. In 2013, Steve Horvath, then at UCLA, published the first multi-tissue DNA methylation clock, which predicts age from chemical tags attached to DNA. These methylation marks shift predictably across the lifespan, and the clock built from them correlates strongly with chronological age across most tissues. Subsequent clocks, including the PhenoAge and GrimAge models developed with input from researchers at institutions such as Yale and Columbia, extended the approach by incorporating clinical biomarkers and mortality risk, improving their ability to forecast health outcomes rather than merely track birthdays.
Read next Laika's Wildwood: Stop-Motion Fantasy at TIFF 2026The supporting evidence is real. A body of peer-reviewed work has linked accelerated epigenetic aging to elevated risks of cardiovascular disease, cancer, and all-cause mortality, and studies of interventions — calorie restriction, certain drugs, intensive lifestyle programs — have reported measurable, if modest, changes in clock readings over months. Those findings are why serious scientists pay attention.
But the controversy sits one layer down. An epigenetic clock is a statistical model trained on population data. It does not measure a single physical quantity the way a thermometer measures temperature. It infers a probability from patterns. That means its output can shift for reasons unrelated to any underlying rejuvenation: changes in the composition of cells in a blood sample, batch effects in the laboratory assay, natural day-to-day noise, or simply regression to the mean — the statistical tendency of extreme measurements to drift back toward average on retesting.
For the de-aging contest, that last point matters enormously. If a participant with an unusually high baseline biological age is retested six months later, part of any apparent improvement may reflect noise rather than genuine reversal. Without controls, blinding, and repeated measurements, a leaderboard can reward artifacts as easily as it rewards biology. Researchers have also disagreed sharply over whether short-term clock changes translate into anything a person would recognize as better health — fewer diseases, more functional years — or whether the clocks are capturing a signal too shallow to matter clinically.
None of this makes biological age measurement pseudoscience. It makes it provisional. The honest summary is that epigenetic clocks are among the best proxies available, and no one has demonstrated they can be moved reliably, durably, and meaningfully in the direction competitors are chasing. A six-month race among hundreds of self-directed participants will produce anecdotes and engagement. It may also produce data. It will not, on its own, settle whether biological age reversal is achievable at scale.
What it will do is expose how much of longevity science currently rests on measurement instruments that are still being calibrated — a theme that echoes, uncomfortably, in the other story of the day.
Opinion: Why LLMs Don't Actually Reason
Thore Graepel, a machine learning researcher, makes a blunt argument in MIT Technology Review: large language models do not reason. The claim lands harder because of who is making it. A decade ago, Graepel helped build a program that stunned the world by defeating a Go champion — an achievement widely treated as a landmark in machine intelligence. He is not a skeptic of AI's power. He is a skeptic of a specific word.
The technical case is well documented. LLMs generate text by predicting the next token from statistical patterns absorbed during training. That process can produce outputs that look like chains of inference. It can pass bar exams, solve textbook problems, and explain its own answers convincingly. But benchmark research has repeatedly shown fragility beneath the surface. Performance on reasoning suites such as GSM8K, a set of grade-school math word problems, can collapse when surface details change — swapping names, altering numbers, or adding an irrelevant clause — even though the underlying logic is identical. On the ARC benchmark, designed to probe abstract pattern reasoning, models have historically underperformed what their language fluency would suggest.
The pattern is consistent with pattern matching rather than formal inference. A model that has seen millions of similar problems can retrieve a plausible solution template. A reasoner would apply rules that hold regardless of phrasing. The distinction is not academic. It determines whether you can trust a system to handle a novel situation it has never encountered in training — the exact situations where reliability matters most.
None of this means LLMs are useless. They are extraordinarily effective at compression, retrieval, translation, summarization, and code generation — tasks where statistical fluency is the point. The error is categorical, not practical: calling fluent text generation "reasoning" inflates what the system does and obscures where it will fail. Graepel's warning is less about AI's limits than about ours — the human tendency to project understanding onto anything that speaks well.
Lessons from AlphaGo: A Decade of AI Hype in Perspective
In 2016, AlphaGo beat Lee Sedol. The moment is often described as AI's moon landing, and Graepel's involvement gives his current critique unusual weight. AlphaGo was a genuine breakthrough, but it was also narrow. It mastered a board game with fixed rules, perfect information, and a clean objective function. Its successor, AlphaZero, learned chess and Go from self-play with superhuman results. Those systems reason — within their domain — because they search over possible futures against a formal goal.
The lesson widely drawn was that general intelligence was around the corner. The lesson actually supported by the evidence was that search plus learned evaluation, applied to a well-specified problem, beats humans at that problem. Ten years on, the gap between board games and open-ended reasoning remains the central unresolved question in AI. Language models skipped the formal goal and scaled the pattern learning. Impressive, and different.
The same discipline applies to the de-aging contest. A decade of AI hype taught the public to overgeneralize from striking demonstrations. A six-month leaderboard risks teaching something similar about biology: that a moving number means a reversing body.
What Both Stories Tell Us About the Limits of Measurement and Intelligence
Two stories, one structure. In each, a measurement instrument produces a signal that is real but partial, and pressure builds to treat it as complete. A biological age clock condenses complex physiology into a number that can be gamed by noise. A language model condenses statistical structure into text that can be mistaken for thought.
The responsible posture in both cases is the same. Ask what the instrument actually measures. Ask how it behaves when the conditions change. Ask whether the impressive result survives a boring retest. On biological age reversal, the answer is that the clocks are credible proxies, the interventions are unproven at scale, and the six-month contest is a useful experiment and a poor proof. On LLM reasoning, the answer is that scaling produced fluency, not logic, and that the difference will show up precisely when we most need it not to.
Graepel watched a machine he helped build beat a Go champion and did not conclude that machines now think. That restraint is the most valuable thing either story offers.
Source: MIT Technology Review



