C5R CORP.

Manage PCR lab operations

Laboratory management
Task space

The agent is given charge of C5R’s automated lab for two simulated weeks and is tasked to run a series of PCR assays. As they do this, random and delayed effects emerge that can upset their plans: deliveries can be short, equipment can break or cause downstream quality problems, and both reagents and consumables can each come from defective batches. The agent is provided with a lab procedure document and eight weeks of recent records.

Agents need to buy requisite reagents, book the lab’s instruments and technician, run tests, deliver results by a deadline, sign for deliveries, and close the books at the end. Different assays can require different reagents, but sometimes can be combined (multiplexed) to save reagents and wells. Each sample in each batch of work only provides enough material for two runs, so agents cannot succeed by exhaustive repeats or random moves.

The task uses the same facility interface as agents doing physical runs. The facility here mirrors the one we have built physically, simulating inventory and experiments. Agents can set off experiments and see their results, track the availability and use of reagents, and how experiments occupy benches and equipment in the lab. Beyond this, agents can also leverage management tools for ordering reagents and consumables, tracking worker schedules, and handling budgeting and expenditure.

Performance is scored relative to a greedy earliest-deadline-first baseline.

Scoring

Deliver at least 39 of the 48 scientific outputs correctly and on time, reconcile deliveries on their arrival day, close orders and claim every eligible credit with no cash holds left, and accept or decline all four requests on their arrival day.

RESULTS

Manage an automated lab over two simulated weeks, completing PCR assays while handling supply shortages, equipment failures, and deadlines.
ModelPass@1Cost ($)Model/API inference cost only, averaged per task attempt. Lab and labor costs are not reported. Repeats are averaged using the same task and family weights as scores.
Claude Fable 5.1xhigh34.4%$27.47
GPT–6 Astraxhigh6.3%$17.93
Claude Opus 5xhigh6.3%$31.80
Gemini 3.8 Flashhigh0.0%$5.63
GPT–5.6 Solxhigh0.0%$17.13
Grok 4.6xhigh0.0%$32.78
0%50%100%

Whiskers show ±1 standard error from repeats.

Model trajectories

One example per model.

Run grading

Fail
MeasurementRecorded resultOutcome
Overall score6.8 %× Fail

Pass condition: Deliver at least 39 of the 48 scientific outputs correctly and on time, reconcile deliveries on their arrival day, close orders and claim every eligible credit with no cash holds left, and accept or decline all four requests on their arrival day.

Manage PCR lab operationsGemini 3.8 Flash high
Run gradingFail
00:0000:01

Model activity

Model transcript

Loading transcript…