Research Papers

 Methodology

The Full Scientific
Discovery Loop.

We evaluate AI scientist systems on their ability to run the full arc of research — autonomously, on real biomedical data. Our lens is Re-X: re-evaluate, reproduce, review, refine.

Evaluation approach.

Can an AI system carry out the full 10-stage scientific lifecycle — from literature review to published impact — without human intervention? That is the question we are answering, on real biomedical data. Our lens is Re-X: re-evaluate, reproduce, review, refine.

Phase 1 — Landscape survey

We surveyed 73+ platforms via PubMed and arXiv, scoring each on lifecycle coverage, autonomy, end-to-end support, multimodal capability, and tool integration.

Phase 2 — System selection

We selected four systems spanning open-source and commercial: Sakana AI v2 (agentic tree search), freephdlabor (multi-agent pipeline), Edison Scientific Kosmos (world-model driven), and Novix / HKUDS AI-Researcher (statistical rigor focus).

Phase 3 — Evaluation on AI-READI

Each system runs on the AI-READI multimodal diabetes dataset. We score across three dimensions: system-level (runtime, autonomy, workflow); technical (reproducibility, error handling, multimodal integration); and scientific rigor (hypothesis novelty, statistical soundness, manuscript quality).

These systems move fast. Findings reflect each platform's capabilities at the time of our landscape review.

Phase 2 — System Selection

We selected four systems spanning open-source and commercial: Sakana AI v2 (agentic tree search), freephdlabor (multi-agent pipeline), Edison Scientific Kosmos (world-model driven), and Novix / HKUDS AI-Researcher (statistical rigor focus).

RE-X the loop Re-evaluate test the claims Reproduce on AI-READI Review score the framework Refine then repeat
Each stage is evaluated for AI capability level — from helper to fully autonomous scientist.
View capability framework →

Stage-by-stage breakdown.

Why these four.

Systems had to be end-to-end, fully autonomous, and multimodal — with a mix of open and commercial.

Open Source
Sakana AI v2
Agentic tree search — generates and iterates on research directions autonomously
Open Source
freephdlabor
Multi-agent — specialist sub-agents coordinate across the full research lifecycle
Commercial
Edison Scientific Kosmos
World-model driven — builds persistent knowledge representations across tasks
Open Source
HKUDS Novix
Statistical-first — emphasis on rigorous analysis and publication-grade manuscript quality

AI-READI.

A multimodal diabetes dataset chosen for being real-world — unlabeled, with missing values — and domain-grounded for expert validation. Its complexity surfaces failure modes that synthetic benchmarks miss.

2,280
Participants
6+
Modalities
4
T2DM Groups
NIH
Bridge2AI
CGM Retinal Sleep ECG

Three dimensions.

1
System-level
Runtime  ·  Cost  ·  Autonomy  ·  Completeness  ·  Iterative reasoning
2
Technical
Task completion  ·  Reproducibility  ·  Transparency  ·  Stability  ·  Error handling  ·  Multimodal integration
3
Scientific rigor
Hypothesis novelty  ·  Statistical rigor  ·  Domain correctness  ·  Interpretation quality  ·  Conclusion–evidence alignment  ·  Manuscript quality

Where we are.

Capabilities reflect the state at time of review — the field moves fast.

Re-X Framework

1
Re-evaluate
Test AI-generated claims on AI-READI data
2
Reproduce
Verify findings hold on real-world, unlabeled data
3
Review
Score across system, technical, and scientific rigor
4
Refine
Identify improvement paths and repeat the cycle