Methodology
We evaluate AI scientist systems on their ability to run the full arc of research — autonomously, on real biomedical data. Our lens is Re-X: re-evaluate, reproduce, review, refine.
Evaluation Approach
Can an AI system carry out the full 10-stage scientific lifecycle — from literature review to published impact — without human intervention? That is the question we are answering, on real biomedical data. Our lens is Re-X: re-evaluate, reproduce, review, refine.
We surveyed 73+ platforms via PubMed and arXiv, scoring each on lifecycle coverage, autonomy, end-to-end support, multimodal capability, and tool integration.
We selected four systems spanning open-source and commercial: Sakana AI v2 (agentic tree search), freephdlabor (multi-agent pipeline), Edison Scientific Kosmos (world-model driven), and Novix / HKUDS AI-Researcher (statistical rigor focus).
Each system runs on the AI-READI multimodal diabetes dataset. We score across three dimensions: system-level (runtime, autonomy, workflow); technical (reproducibility, error handling, multimodal integration); and scientific rigor (hypothesis novelty, statistical soundness, manuscript quality).
These systems move fast. Findings reflect each platform's capabilities at the time of our landscape review.
Phase 2 — System Selection
We selected four systems spanning open-source and commercial: Sakana AI v2 (agentic tree search), freephdlabor (multi-agent pipeline), Edison Scientific Kosmos (world-model driven), and Novix / HKUDS AI-Researcher (statistical rigor focus).
Deep Dive
System Selection
Systems had to be end-to-end, fully autonomous, and multimodal — with a mix of open and commercial.
The Testbed
A multimodal diabetes dataset chosen for being real-world — unlabeled, with missing values — and domain-grounded for expert validation. Its complexity surfaces failure modes that synthetic benchmarks miss.
Evaluation Framework
Progress
Capabilities reflect the state at time of review — the field moves fast.
Re-X Framework