Lab-in-the-Loop, and the Baselines It Has to Beat
Published:
New review and perspective: Lab-in-the-Loop, and the Baselines It Has to Beat — the 2019–2026 campaigns assembled into one ledger of what was measured, over how many rounds, at what sample cost, and at what level of evidence.
The short version. The real achievement of lab-in-the-loop is narrower and more useful than the headline fold-improvements suggest: it has moved protein engineering from screening millions of variants to characterising hundreds. Four rounds and fewer than 500 variants per enzyme; one round and roughly 120 variants for a seven-mutation multi-mutant; four rounds and 1,800 variants for sub-100 pM antibodies against four targets. That is a two- to four-order-of-magnitude cut in experimental cost, and it changes what a small lab can attempt.
Three things fall out of the ledger.
Sample efficiency, not raw performance, is the axis the field has actually moved along. The strongest recent results characterise hundreds of variants, not millions.
Autonomy and scientific value are close to orthogonal. The most fully autonomous platforms report the most modest gains; the largest reported gains come from loops that humans executed.
Every closed-loop result that has been independently re-examined has moved in the same direction. The A-Lab’s 41 novel compounds became 36 synthesised and none demonstrably novel. MULTI-evolve’s epistasis model was shown to track a plain additive baseline at r > 0.999. SAMPLE’s ≥12 °C stabilisation became about 10 °C on human re-measurement. A two-cycle metabolic campaign gained 21% titer while its models scored R² ≈ −0.29 out of sample. In each case the correction came from outsiders, not from the platforms — a structural point about incentives, not about anyone’s integrity.
Where the comparison has been run, the cheap alternative holds up better than expected: a BLOSUM substitution matrix beat four deep learning methods across 720 antibody designs measured by SPR, and the largest antibody campaign’s own control arm found classical repertoire mining produced the best variant for two of five seeds.
So the review ends with a reporting standard rather than a benchmark — report the additive baseline at the same budget, the assay’s measurement error, the full denominator of designs proposed versus built versus assayed, and which decisions were made by the model rather than by a person.
There is still no peer-reviewed cross-laboratory reproducibility study for self-driving labs in biology. That absence is the thing to fix.

