SciUniverse
LEVEL 1Can frontier models carry out scientific work?
SciUniverse is a benchmark of scientific work that spans the entire research process across chemistry, biology, and materials science.
We start with SciUniverse Level 1: tasks that are straightforward for scientists and take no more than a few hours of work. Future releases will test more demanding scientific work.
Models direct work at Facility-0 using our facility harness by controlling machines and giving instructions to human operators. Working with real materials and instruments lets us test whether their plans produce the intended results and how they adapt when experiments fail.
The samples below show recorded footage (left), the equipment involved (center), and each action’s inputs and outputs (right).
GPT–6 AstraSynthesize N-benzyl-4-methylbenzamide
RESULTS
| Model | Pass@1 | Cost ($)Model/API inference cost only, averaged per task attempt. Lab and labor costs are not reported. Repeats are averaged using the same task and family weights as scores. | |
|---|---|---|---|
| 45.3% | $40.61 | ||
| 32.5% | $52.37 | ||
| 30.5% | $46.31 | ||
| 26.2% | $13.41 | ||
| 14.6% | $4.55 | ||
| 9.4% | $16.53 | ||
0%50%100% | |||
A task space for scientific work
We think most scientific benchmarks capture too narrow a slice of scientific work, starting after the experiment with clean data ready to analyze. That leaves out much of the work: choosing materials, operating instruments, running experiments, debugging failures, and adapting to real constraints.
To measure model performance across more of that work, we chose tasks covering:
- Basic sample preparation. We included simple sample-preparation tasks because early tests showed that models struggled with them.
- Instrument control. We compared how well models controlled instruments using vendor software, Python APIs, and direct firmware commands.
- Protocol adaptation. We chose standard processes and kits, then tested models’ ability to adapt to limited reagents, equipment, and time.
- Learning across experiments. We included optimization tasks to see whether models could learn from experiments and choose what to try next with limited materials and budgets.
- Facility management. We added facility management to see whether models could keep experiments on track through shortages, equipment failures, and deadlines.
- Interpreting real measurements. We used real NMR, XRD, and chromatography data to test whether models could identify products, quantify mixtures, and recognize unreliable signals.
SciUniverse Level 1 contains 92 tasks across 17 task families. Each task gives a model a specific scientific objective, such as synthesizing a target molecule. Tasks involving the same kind of scientific work form a family. Each family contributes equally to the overall benchmark score.
We map tasks by work horizon and scale of matter to compare chemistry, biology, and materials science in one frame. Select a task on the map to explore its results.
- Synthesize and detect an amide
- Amplify and recover DNA
- Synthesize alpha alumina
- Express sfGFP in a cell-free system
- Develop a two-minute LC–MS method
- Assign structures from NMR
- Analyze XRD patterns
- Press BaTiO₃ pellets
- Prepare PVA solution
- Maximize amide synthesis yield
- Identify unsuitable MRM peaks
- Optimize a single reaction
- Optimize reactions in parallel
- Optimize reactions with limited material
- Manage PCR lab operations
- Run sfGFP PCR · Hamilton
- Normalize and pool samples · Hamilton
SciUniverse
Synthesize an amide
Confirm product formation by LCMS when synthesizing N-benzyl-4-methylbenzamide from p-toluic acid and benzylamine.
Reasoning
Loading model transcript…
Video will appear when recorded activity begins.
Loading facility view…
Model activity
How models fail in the physical world
Models repeatedly miss critical details in laboratory work. We observed them trying to pipette samples that were still frozen, vortexing open well plates, reusing the same pipette tip across DNA-containing wells, and failing to account for evaporating solvents. They struggle with both physical reasoning and protocol design. The table below breaks down each model’s documented failures in physical tasks by category.
| Failure mode | Astra 6 | Fable 5.1 | Gemini 3.8 | Grok 4.6 | Opus 5 | Sol 5.6 |
|---|---|---|---|---|---|---|
Acknowledgments
We thank the scientists, operators, and engineers who made this benchmark and Facility-0 possible, and the researchers and maintainers who shared the experimental data and software behind some of the simulated tasks in SciUniverse Level 1.
Full credits
- Reaction optimization. The AIChemEco amide-coupling dataset, published by Zhang et al., supplies recorded reaction outcomes for our single-reaction, parallel-reaction, and material-limited optimization tasks.
- NMR structure assignment. Contributors to nmrXiv, the Biological Magnetic Resonance Data Bank (BMRB), and the Natural Products Magnetic Resonance Database (NP-MRD) shared the experimental proton and carbon NMR data used in our structure-assignment tasks.
- X-ray diffraction. The Dara dataset from Fei et al. provides measured powder mixtures and pure-phase reference patterns for our quantitative phase-analysis task.
- Peak quality review. The CPTAC Study 9 team (Abbatiello et al.) shared LC–MRM chromatograms and analyst peak annotations in Panorama Public. These recordings and annotations underpin our task for identifying peaks unsuitable for quantification.
We also thank the developers of PyLabRobot, one of our instrument-control interfaces, and the contributors to GSAS-II and the Crystallography Open Database, used to build our XRD reference baselines.
Contact us
Get in touch to evaluate your model on SciUniverse or work with us on new scientific tasks. We welcome collaborations with teams developing AI models, building evaluations, and doing experimental research.
Email contact@c5r.net