Study

An open benchmark and language models for AI in aging biology

Alex Zhavoronkov, Vladimir Naumov, Denis Sidorenko, Alex Aliper, Vladimir Aladinskiy, Ramin Hasani, Alexander Amini, Katerina Nasto, Mathieu Reymond, Rim Shayakhmetov

BENCHMARK STUDY OF AI SYSTEMS ACROSS FIVE AGING-BIOLOGY BIODATA DOMAINS WITH DOMAIN-SPECIFIC MODEL FINE-TUNING 2026

LongevityBench found that compact aging-focused language models matched or exceeded larger systems on 17 aging-biology tasks, while omics-based age prediction remained difficult.

DesignBENCHMARK STUDY OF AI SYSTEMS ACROSS FIVE AGING-BIOLOGY BIODATA DOMAINS WITH DOMAIN-SPECIFIC MODEL FINE-TUNING
TierTier 2, Product RCT (not peer-reviewed)
Year2026
JournalCell
N17 tasks; 18 AI systems; 5 Longevity-LLMs
PublishedSep 17, 2026
Added to NO1GEVITYSep 20, 2026

Alex Zhavoronkov and colleagues introduced LongevityBench, an open benchmark with 17 tasks across five biodata domains, and tested 18 frontier AI systems from six developer teams. They also fine-tuned five aging-focused Longevity-LLMs ranging from 0.6 to 9 billion parameters. No single model dominated every task. Omics-based age prediction was the hardest category regardless of model scale. The compact Longevity-LLMs matched or exceeded much larger general-purpose systems on the benchmark. The work matters as infrastructure for longevity research: it creates a reproducible way to compare models on biological-age prediction and structured aging data rather than relying on general language benchmarks. It does not discover a longevity molecule, validate a biomarker clinically, or show that AI-generated hypotheses improve human health. Benchmark performance can also reflect training-data overlap, task construction, and the choice of metrics. The commercial interests are material. Alex Zhavoronkov founded Insilico Medicine, several authors work there, and the benchmark, model weights, and research interface are linked to commercial drug-discovery activities. The paper is best read as a tool and methods report. Independent replication, held-out datasets, prospective predictions, and experimental validation are needed before model performance is treated as evidence about aging biology.

The compact (0.6B-9B parameters) Longevity-LLMs matched or exceeded far larger frontier systems on LongevityBench, showing that general-purpose language models can be adapted to structured-omics tasks. We publicly release the benchmark, models, and Longevity Claw, an agentic research interface for aging researchers.
Critic notes

Methods and benchmark paper, not a human study. Results may depend on task design and training-data overlap. Multiple authors and released tools have commercial ties to Insilico Medicine or Liquid AI.

Read the source