machine learning in biology intermediate

Accelerating scientific discovery with Co-Scientist

TL;DR

Co-Scientist, a multi-agent AI system built on Gemini 2.0, autonomously generates and refines novel scientific hypotheses that were successfully validated in wet-lab experiments for acute myeloid leukemia and liver fibrosis.

Problem / question

Generating novel scientific hypotheses requires mastering an overwhelming volume of specialized literature and making complex cross-disciplinary connections, a bottleneck that slows down the pace of scientific discovery and drug development.

Methods

Researchers developed Co-Scientist, a multi-agent AI system using Gemini 2.0 that mimics the scientific method through generation, reflection, and tournament-based ranking (using Elo ratings) of hypotheses. The system scales test-time compute to iteratively debate and refine ideas. It was computationally evaluated on 203 research goals and validated in wet-lab experiments using 4 acute myeloid leukemia (AML) cell lines (MOLM-13, KG-1a, HL-60, NOMO-1), a non-AML control (TK6), and human hepatic organoids.

Key findings

Co-Scientist outperformed models like OpenAI o1 and DeepSeek R1 in generating high-quality hypotheses across 15 expert-curated biomedical goals. In wet-lab validations for AML, it successfully identified Binimetinib (IC50 of 2 nM) and the novel IRE1-alpha inhibitor KIRA6 (IC50 of 10 nM in KG-1a cells versus 180 nM in healthy TK6 controls) as potent therapies. It also accurately predicted synergistic drug combinations like JNJ-64619178 with Selinexor, and independently deduced a novel bacterial gene transfer mechanism involving capsid-forming phage-inducible chromosomal islands.

Why it matters

By acting as an autonomous collaborator that can read vast amounts of literature and iteratively debate ideas, this AI system can significantly accelerate the early stages of research, such as drug repurposing and de novo target discovery.

Limitations

The system relies on open-access literature, meaning it misses paywalled research and unpublished negative results, and it inherits standard large language model flaws like potential hallucinations and the risk of propagating erroneous published findings.

Takeaway

An AI system using multi-agent debate and test-time compute scaling can generate novel, experimentally valid biological hypotheses, such as identifying KIRA6 as a highly selective drug candidate for acute myeloid leukemia.

Full paper

Your browser can’t display PDFs inline. Download the PDF instead.