- Study finds specialized models can match or outperform larger AI systems
ALKHOBAR: The race to build increasingly large artificial intelligence models has dominated the technology industry, but new research into aging biology is testing whether scientific AI needs a different approach.
Researchers at Insilico Medicine have developed a benchmark to measure whether AI systems can reason across biological data related to aging, alongside smaller, specialized models trained for longevity research.
The results raise a broader question for scientific AI: whether increasing model size alone can deliver the specialized reasoning needed to interpret complex biological data, or whether smaller systems grounded in domain-specific evidence can sometimes perform better.
For the Gulf, the research also reflects the region’s expanding role in AI-driven science. Insilico Medicine’s Abu Dhabi research center was the lead institutional affiliation on the work, while the company is separately collaborating with Saudi Aramco and academic partners including King Abdullah University of Science and Technology on applying generative AI to materials research.
At the center of the longevity research is LongevityBench, an open benchmark covering 17 tasks across clinical, epigenetic, transcriptomic, proteomic and genetic data.
Alex Zhavoronkov, founder and CEO of Insilico Medicine, said the benchmark was developed to distinguish between AI systems that can interpret experimental biological evidence and those that produce convincing answers based largely on information encountered during training.
“What we never had was a shared, rigorous way to ask a simple question: can an AI system reason over this data the way a trained biologist would, or is it producing fluent text that only sounds right?” Zhavoronkov told Arab News.
Existing benchmarks, he said, often focus on sequence-function prediction, clinical question answering or general pharmacology rather than requiring models to reason directly from biological measurements related to aging.
LongevityBench instead uses measurable outcomes, including comparing biological age from DNA methylation data, classifying tissue samples by age from transcriptomic information and predicting chronological age using proteomic profiles.

Researchers evaluated systems from six AI developer teams — OpenAI, Google, Anthropic, xAI, DeepSeek and Moonshot AI — alongside five specialized Longevity-LLMs developed by the team.
According to Zhavoronkov, the smaller models challenged the assumption that greater scale necessarily translates into better scientific performance.
“This is the finding I find most important, and it goes against the prevailing assumption in AI that bigger always wins,” he said.
The specialized models matched or exceeded most of the frontier systems evaluated on the benchmark. Zhavoronkov said that on DNA-methylation age comparison, a specialized model outperformed the best-performing frontier system tested, while a 0.6-billion-parameter model substantially reduced mean absolute error on a proteomic age-prediction task compared with the best general-purpose system.
“Scaling general intelligence gets you fluency. It does not automatically get you scientific competence in data-dense domain like aging biology,” he said. “The path forward is not simply bigger models. It is models that are properly grounded in the underlying biological measurements — and those models can be built, evaluated, and released from a research center in Abu Dhabi.”
The work was published in Cell in September and lists Insilico Medicine AI in Masdar City, Abu Dhabi, as the lead institutional affiliation.
But evaluating AI for longevity presents a problem that model size cannot solve: scientists do not have a single universally accepted definition or measurement of biological aging.
Rather than attempting to resolve that debate, LongevityBench anchors its tasks to measurable outcomes, including chronological age, mortality or survival and established biomarkers such as methylation and proteomic aging clocks.
“This is a real tension, and I will not pretend it is fully resolved,” Zhavoronov said.
“These are not perfect proxies for the full biological process of aging, but they are reproducible, they are tied to real cohorts such as NHANES and GTEx, and they let us measure whether a model’s biological reasoning is consistent with what we can observe and verify.”

The company is also testing how specialized AI can move beyond benchmark questions into generating scientific hypotheses.
Longevity Claw, an agentic research system released alongside the benchmark and models, combines the Longevity-LLMs with tools including aging clocks, gene-set enrichment, population profiling and evidence retrieval.
The aim is to allow researchers to move from biological data toward potential therapeutic targets without manually switching between multiple computational tools.
In an initial autonomous campaign, the system nominated 328 genes across the hallmarks of aging, according to Zhavoronkov. He said the results showed up to 5.6-fold enrichment against an independently published set of experimentally supported aging targets.
But identifying a promising target computationally and proving that it can safely affect human health are very different things.
“What it does not change is who is accountable for validating a hypothesis,” Zhavoronkov said. “Every target our systems propose still has to survive wet-lab validation. That is a non-negotiable step in our own pipeline.”
That distinction is particularly important in longevity research, where biomarkers can suggest a biological effect without proving that an intervention extends healthy human lifespan.
A separate Insilico Medicine study published in Nature Biotechnology in September applied six proteomic aging clocks to serum samples from a Phase 2a trial of rentosertib in patients with idiopathic pulmonary fibrosis.
According to Zhavoronkov, all six clocks detected the same direction of biological-age reduction among treated patients compared with placebo, although their absolute predictions differed.
However, he acknowledged that the result cannot establish that the treatment itself has an anti-aging effect independent of its effect on the underlying disease.
“We cannot fully separate a true anti-aging effect from the drug’s anti-fibrotic activity on the disease itself,” he said.
“Full disentanglement requires testing the same mechanism in healthy volunteers, not only a diseased population.”
The limitation illustrates a wider challenge as AI becomes more capable of finding patterns across biological datasets. A model may produce a compelling hypothesis, but its quality still depends on the data, how a task is framed and subsequent experimental validation.
Zhavoronov said LongevityBench itself showed that large language models can perform differently when the same biological information is presented as a comparison, regression or classification problem.
“If you are not aware of that sensitivity, you can convince yourself a model understands the biology when it is actually responding well to a particular prompt format,” he said.

The Abu Dhabi center behind the work has grown to around 60 AI engineers, computational biologists and chemists, according to Zhavoronov, and has also contributed to the company’s drug-discovery programs.
The company’s regional work extends into Saudi Arabia through a collaboration with Saudi Aramco applying generative AI methods to materials science, including the design of metal-organic frameworks for carbon capture and hydrogen storage.
Zhavoronov said the collaboration has produced open tools for validating metal-organic framework structures and a database of around 240,000 validated frameworks for direct-air-capture research, alongside academic partners including KAUST and the University of Montpellier.
The connection illustrates how AI architectures developed for one scientific field may increasingly be adapted to others, from biological molecules to industrial materials.
For longevity research, however, the immediate significance of AI may be less about replacing scientists than narrowing the enormous search space they face.
The technology can compare datasets, identify patterns and generate potential targets at a scale that would be difficult to reproduce manually. But it cannot determine by itself whether those targets will ultimately translate into safe and effective treatments.
“Aging clocks and AI target predictions are hypothesis-generation tools, not proof of efficacy or safety in humans,” Zhavoronkov said.
“The risk in this field is not that AI gets the biology wrong occasionally — every method does that. The risk is mistaking a well-calibrated hypothesis for a validated result.”
As specialized AI systems become more capable, that distinction may prove as important as the performance gains themselves. In scientific research, producing a better prediction is only the beginning of proving that it is true.






