1Introduction
Why the loop, why now, and what the term has come to mean.
The idea that a machine might run its own experimental cycle is not new. In 2004 a robot built at Aberystwyth originated hypotheses about yeast gene function, designed auxotrophic growth experiments to test them, executed them physically, and repeated the cycle; its experiment-selection strategy was "competitive with human performance" at a threefold cost reduction over the cheapest alternative and a hundredfold over random selection.1 The system was named Adam only later, in the 2009 paper that reported its full 10,000-unit experimental tree.2 Its successor Eve applied closed-loop QSAR screening to drug repositioning and found that the anti-cancer compound TNP-470 potently inhibits dihydrofolate reductase from Plasmodium vivax — IC50 0.16 µM against the parasite enzyme versus more than 165 µM against the human one, roughly a thousandfold window, on top of the compound's known MetAP2 mechanism.3
What changed after 2019 was not the loop but what sits inside it. Protein language models supplied usable zero-shot priors over sequence space, so a campaign no longer had to begin from nothing.4 Bayesian optimisation and active learning gave a principled way to spend a small experimental budget. Biofoundries and cloud labs made the build-and-test arm fast enough that model retraining, not pipetting, could set the cycle time. The term lab-in-the-loop itself entered wide use through industry: Genentech's public framing from November 2023 onward describes "lab in a loop", where experimental data feed computational models that make new experimentally testable predictions,5 and the 2025 Prescient Design preprint that carries the phrase in its title is the largest published campaign of its kind.6
The vocabulary is crowded — self-driving labs, autonomous experimentation, design–build–test–learn, closed-loop optimisation — and the distinctions matter less than one structural question: what closes the loop? A useful scale already exists. Beal and Rogers proposed levels of autonomy for synthetic biology engineering in 2020;7 Garcia Martin and colleagues adapted it, added the self-driving-car analogy, and set the criterion that systems at level three or above count as self-driving labs, because at that point the DBTL loop is closed.8 (The levels are frequently misattributed to the latter paper; its own figure caption credits Beal and Rogers.) Under that criterion much of what is marketed as a self-driving lab is level one or two — robots doing repetitive work, humans making every decision — while several systems that make all in-loop decisions autonomously run on modest hardware. Chemistry and materials science, which are further along mechanically, have two comprehensive surveys worth reading before any biology-specific one — Tom and colleagues' review of the SDL literature through 2023,29 and Canty and Abolhasani's 2026 framing of the next phase around scalability, generalisability and provenance-complete experimentation.30
This article does three things. Section 2 assembles the campaigns with verified numbers, organised by domain, and then examines the corrections record as evidence in its own right. Section 3 argues that the field's central problem is the baseline and the assay rather than the model, and proposes what a LitL claim should be required to report. Section 4 states what I think the next three years should look like.
2Results
What the campaigns measured, at what sample cost, and what survived re-examination.
2.1 Antibodies and therapeutic proteins
The Prescient Design/Genentech campaign is the reference point for scale.6 Eleven seed antibodies raised by animal immunisation against four clinically relevant antigens — EGFR, IL-6, HER2 and OSM — entered a loop of generative design, multi-task property prediction, and tournament-style ranking by oracles selected on held-out data from the previous round. Over 1,800 unique variants were tested across four rounds, yielding "3–100× better binding variants for all targets and 10/11 seeds, with the best binders exceeding 100 pM affinity". Two points of precision are worth carrying, because secondary coverage routinely mangles them: exceeding 100 pM means affinity better than 100 pM, i.e. sub-100 pM KD; and the structural work comprises eight new apo crystal structures (four seeds, four designs) spanning three targets, plus two previously published — not co-crystal complexes with antigen. The campaign remains a preprint at version 3 (November 2025) with no journal version as of this writing, and all authors are Genentech or Roche employees.
Two further details from the same preprint matter more than the fold-changes, and both are to the authors' credit. First, the real closed-loop signal is the hit rate, not the maximum: round 1 produced 3×-improved binders for three of eleven seeds — about 2% of designs — while round 4 yielded a design library that was over 21% 3×-better binders, roughly a tenfold improvement in design quality across four cycles. That is what a loop learning looks like, and it is better evidence than the 100× headline. Second, the campaign ran a control arm and reported the result plainly: "Affinity measurements for the control samples had comparable distributions to our machine-learning-informed designs. For three of five seed antibodies, the highest affinity improvement was achieved via ML methods… For two out of five seed antibodies, the highest affinity improvement was achieved in the control antibodies." Against matched-budget classical repertoire mining, the models won three of five. The largest LitL campaign yet published contains, in its own text, one of the field's more candid statements about what the loop is worth.
Elsewhere in the antibody field the pattern is similar: strong numbers, thin peer review. Absci's zero-shot anti-HER2 work — frequently described in summaries as journal-published, which it is not — reports 421 diverse binders from a screening-heavy library, and is single-shot rather than closed-loop. Nabla Bio's GPCR campaign, which is genuinely closed-loop in that a validated agonist re-prompts the generative model, reports hundreds of VHH binders against CXCR4 and CXCR7 at picomolar to low-nanomolar affinity, and carries an explicit disclaimer that methodological details are withheld for commercial reasons. The exceptions that anchor the field are peer-reviewed but not closed-loop: A-Alpha Bio's AlphaBind reports a 766 fM scFv-Fc variant, 74-fold over parental, explicitly "using only a single round of data generation".9
The clinical end of this pipeline has now produced its first peer-reviewed readouts, and they are sobering in a useful way. Generate Biomedicines' GB-0669, an AI-designed monoclonal, completed a randomised placebo-controlled first-in-human trial (n = 51) reported in The Journal of Infectious Diseases: well tolerated, no dose-limiting toxicities, 54-day half-life.10 That is a safety result, not an efficacy one. Insilico Medicine's rentosertib (ISM001-055), the generative-AI-discovered TNIK inhibitor, published a randomised phase 2a trial in Nature Medicine whose primary endpoint was safety; the FVC change of +98.4 mL at 60 mg QD versus −20.3 mL for placebo was a secondary endpoint.11 A phase III trial began dosing on 9 September 2026. Any review that reports the FVC figure as the trial's primary result — and many do — has inverted the evidentiary structure.
2.2 Enzymes and protein fitness: the autonomous end
Two campaigns represent the field's genuine level-3-and-above demonstrations in biology. SAMPLE deployed four independent agents, each seeded with the same six natural GH1 glycoside hydrolase sequences, each choosing three sequences per round for twenty rounds, with no human in the decision path.12 All four converged on enzymes "at least 12 °C more stable than the six initial natural sequences" while searching less than 2% of a 1,352-member combinatorial landscape. Two details deserve to travel with that headline: when the designs were re-measured by humans under different expression and assay conditions, the margin over the best natural parent shrank to roughly 10 °C, which the authors attribute to assay conditions; and the campaign took just under six months in practice, against a nominal throughput of one to two weeks.
SAMPLE is also the only closed-loop biology campaign that reports its own economics honestly enough to be used as a planning number. A twenty-round campaign cost about $5,200 — $2,400 DNA, $1,300 reagents, $1,500 cloud-lab access — or roughly $87 per variant. Approximately 9% of experiments failed, "presumably due to liquid-handling errors", and in two cases faulty data passed the quality filters. The six-month elapsed time reflects system downtime, robotic malfunctions and reagent restocking; the authors' revised realistic estimate is about two months. Set against conventional directed evolution at six to twelve months, the honest speed-up is on the order of one- to twofold, not the order of magnitude the genre's rhetoric implies.
The UIUC/Zhao platform is the strongest generalist result.13 Built on the iBioFAB biofoundry with ESM-2 and EVmutation supplying priors, it engineered A. thaliana halide methyltransferase to a ~90-fold improvement in preference for ethyl iodide over methyl iodide and 16-fold higher ethyltransferase activity, and a Y. mollaretii phytase variant to 26.3-fold higher specific activity at pH 6.6 — "in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme" (482 and 448 respectively). An independent analogue from Zhang and colleagues, using ESM-2 to seed 96 variants per round through a biofoundry over four rounds in ten days, reports a 2.4-fold activity gain on a tRNA synthetase14 — a valuable robustness datapoint precisely because the number is modest.
2.3 Human-executed loops: where the largest gains are
ALDE applies uncertainty-aware active learning to five epistatic active-site residues and, "in three rounds of wet-lab experimentation", improves the yield of a non-native cyclopropanation product from 12% to 93%.15 EVOLVEpro, a few-shot active-learning framework over protein language model embeddings, reports "up to 100-fold improvements" across six proteins spanning RNA production, genome editing and antibody binding.16 That ceiling belongs to T7 RNA polymerase and concerns mRNA quality rather than catalytic rate; antibody gains were up to 30-fold, the prime editor about 2-fold. MULTI-evolve characterises all pairwise combinations of roughly fifteen top single mutations — on the order of 100–200 variants per campaign, not "~200 training examples" — and reports multi-mutants carrying up to seven mutations across three proteins, with the abstract's headline claim being "up to 10-fold improvements with a single round of machine learning-guided directed evolution".17 The widely quoted 256-fold figure is APEX-versus-wild-type specifically, and blending it with the headline is a misreading.
None of these three is autonomous. All are human-executed loops in which the model proposes and people build, assay and decide. They report the largest gains in this review.
2.4 Metabolic engineering: the honest failure
ART brought Bayesian ensemble recommendation to synthetic biology's small-data, recursive regime, demonstrated across biofuels, hopless "hoppy" beer, fatty alcohols and tryptophan.18 Only the tryptophan study was prospective; the rest are retrospective re-analyses showing how the tool could have been used.
The dodecanol campaign is the most instructive negative result in the literature, and it is negative in a specific place.19 Two DBTL cycles over 60 engineered E. coli strains produced a real 21% titer increase to 0.83 g/L — more than sixfold above previously reported batch values in minimal medium. The production improved. The learning did not: the authors' own later re-analysis records cross-validation R² of −0.29 and states plainly that "the machine learning algorithms were not able to produce accurate predictions with the low amount of data available for training". Worse, the recommendations were not physically actionable — a prescribed sixfold increase in protein expression could only be realised as about twofold, and some proposed strains could not be built at all owing to toxicity. A loop can deliver a gain while its model contributes nothing.
Simulation work points at the same bottleneck from the other direction. Comparing DBTL scenarios in silico, van Lent and colleagues find that "screening capacity is a dominant driver of optimization success, whereas DNA sequencing capacity has surprisingly little impact".20 The comparison is against other process parameters, not against the choice of ML method — that head-to-head lives in their earlier paper, which found gradient boosting and random forests best in the low-data regime and large first cycles preferable to uniform ones.21
2.5 Autonomous chemistry, as the adjacent case
Chemistry's closed loops are further along mechanically and worth reading for that reason. The Jensen group's autonomous platform "experimentally realized 294 unreported molecules across three automatic iterations" of design–make–test–analyze while exploring four rarely reported scaffolds, optimising absorption wavelength, lipophilicity and photo-oxidative stability in dye-like molecules.22 (The 312 figure in circulation is the ChemRxiv preprint's; peer review revised it to 294, and a second case study's six top performers to nine.) Its successor added multifidelity Bayesian optimisation over docking scores, single-point percent inhibition and dose–response IC50 — docking more than 3,500 molecules, synthesising and screening more than 120 against HDAC8, two full cycles in a month.23 Synthesis is the enabling step between fidelities, not a fidelity tier.
Coscientist showed that a GPT-4 planner with search, code execution, documentation retrieval and robotic execution could navigate hardware APIs and run cross-coupling reactions.24 It is routinely over-described. Its reaction "optimisation" ran over two fully enumerated published datasets — "any reaction proposed by Coscientist would be within these datasets and accessible as a lookup table" — so it was not wet-lab optimisation; the physical couplings gave qualitative GC-MS product detection, not yields; plates were moved by hand; and data, code and prompts were withheld pending regulation, with the public repository a simpler implementation that "may not produce the same results". The paper's own abstract says "(semi-)autonomous".
2.6 Agentic loops: the newest case, and the most candid
The 2026 entrant is the language-model agent as the loop's decision-maker. The most quantified example is a Ginkgo Bioworks–OpenAI study in which GPT-5 drove optimisation of cell-free protein synthesis.33 Over six iterative steps across six months, the model designed 480 384-well plates and received 29,527 unique reaction compositions, reaching sfGFP at $422/g specific cost against a prior state of the art of $698/g — a 40% reduction — with a simultaneous 27% titer increase to 3.04 g/L. Design-level hallucination was rare: two of 480 plates had fundamental flaws, one a volume-constraint override, one a nanolitre-to-microlitre unit error.
It is a company preprint, and it is unusually honest about its own confounds, which makes it more useful than a cleaner result would be. Four caveats come from the authors themselves. Ginkgo personnel wrote every execution protocol; the model operated inside a validated design envelope. The large performance jump at step 3 coincided with four simultaneous changes — tool access, a handout of the prior state-of-the-art preprint including its supplementary tables, a new plasmid design, and improved lysate preparation — so model capability cannot be separated from the literature handout and wet-lab process engineering. The assay had to be rescued by hand: early replicate coefficients of variation exceeded 40%, and human staff re-engineered reagent stocks to bring the median per-plate CV to 17%. And the winning composition generalised poorly — of twelve other proteins, only six produced at titer visible by SDS-PAGE.
One further result deserves more attention than it has received: the agent's own self-critique did not work. Across three design-scoring regimes, heuristic LLM scoring "did not obviously correlate with superior designs", top compositions came from both high- and low-scoring designs, and improvement continued in later steps with no scoring at all. An agent that cannot rank its own proposals is, in loop terms, an expensive sampler.
Where agentic systems are claimed to close the loop on discovery rather than optimisation, the validated output is thinner than the coverage suggests. The Virtual Lab, in which an LLM principal-investigator agent directs scientist agents through structured research meetings, designed 92 nanobodies against SARS-CoV-2.34 Two acquired binding to JN.1 while retaining ancestral-spike binding — a 2.2% yield — and the authors themselves describe the new binding as moderate. The readout was an indirect antigen-array ELISA: no SPR, no KD, no neutralisation, no structure. Only one of the two showed any KP.3 signal, at roughly 3% of its Wuhan signal, and the confirmatory twelve-point titration was run against Wuhan and JN.1 only. Secondary accounts that describe the Virtual Lab as having designed nanobodies that bind KP.3 are not supported by the deposited data.
Google's Co-Scientist reports three application areas, of which one is new agent-prompted wet-lab work: drug repurposing in acute myeloid leukaemia, assayed by MTS viability.35 Of three candidates proposed fully autonomously, one was active — KIRA6, with IC50 10 nM in KG-1a against 180 nM in the non-AML control line. That 18-fold window rests on a single AML line; in two of four AML lines KIRA6 was five- to tenfold less potent than in the normal control. The liver-fibrosis work is reported in a separate paper, in human hepatic organoids.51 And the antimicrobial-resistance result that dominated the coverage was not a prospective test at all: the experimental programme was complete and unpublished before the model was asked, and the paper reports that the top-ranked hypothesis matched an already experimentally confirmed mechanism.52 No AI-generated hypothesis was tested at a bench. In peer review the companion paper's title changed from "a novel mechanism" to "a mechanism".53 It is a retrodiction — a genuinely interesting one — and not a discovery.
FutureHouse's Robin, published in the same issue as Co-Scientist, ran a literature-and-analysis loop over retinal pigment epithelium phagocytosis and nominated ripasudil, an approved ROCK inhibitor.54 Humans ran every experiment. The detail worth the whole paper is this: on identical flow-cytometry data, the agent's automated analysis reported a 7.5-fold increase in phagocytosis, while human analysis of the same data gave 1.75-fold. The headline figure is the agent's. Nothing was fabricated; the analysis pipeline simply had latitude, and it used it in the flattering direction.
Two further agentic systems have genuine, modest wet-lab records worth stating precisely, because they bracket the range. CRISPR-GPT designed and analysed gene-editing experiments executed by junior researchers unfamiliar with the technique, reaching about 80% editing efficiency across four genes in A549 cells and 56.5–90.2% activation in a CRISPRa experiment55 — a real result, with all target choices made by humans. Biomni, an agent spanning 150 tools and 59 databases, reports strong benchmark numbers; its wet-lab contribution in the preprint is a single Golden Gate cloning of one sgRNA, executed by a human, with two of two colonies sequence-perfect.56 Benchmark performance and bench performance are not the same axis, and the gap between them is currently about three orders of magnitude in effort.
2.7 The ledger
Table 1 puts the campaigns on common axes. The column that matters most is not the gain but the pairing of sample cost with evidence tier.
| Campaign | Domain | Loop closed by | Rounds | Variants / experiments | Reported gain | Evidence |
|---|---|---|---|---|---|---|
| Prescient Design LitL6 | Antibodies, 4 targets | Human-executed | 4 | >1,800 | 3–100× binding; best sub-100 pM | preprint v3 |
| SAMPLE12 | GH1 thermostability | Agent, no human decisions | 20 | 60 per agent | ≥12 °C (~10 °C re-measured) | peer-reviewed |
| UIUC / iBioFAB13 | AtHMT; YmPhytase | Biofoundry, autonomous | 4 | 482; 448 | 90× & 16×; 26.3× | peer-reviewed |
| Zhang et al. biofoundry14 | tRNA synthetase | Biofoundry, automated | 4 | 384 | 2.4× | peer-reviewed |
| ALDE15 | Enzyme active site | Human-executed | 3 | — | 12% → 93% yield | peer-reviewed |
| EVOLVEpro16 | 6 proteins | Human-executed | few-shot | — | up to 100× (T7 RNAP) | peer-reviewed |
| MULTI-evolve17 | APEX, dCasRx, HuABC2 | Human-executed | 1 | ~100–200 | up to 10× (256× APEX vs WT) | peer-reviewed |
| Dodecanol DBTL19 | E. coli titer | Human-executed | 2 | 60 strains | +21% (models R² ≈ −0.29) | peer-reviewed |
| Jensen DMTA22 | Dye-like molecules | Robotic, autonomous | 3 | 294 realised | multi-property Pareto | peer-reviewed |
| MF-BO / HDAC823 | Small-molecule drug | Robotic, autonomous | 2 cycles | 3,500 docked; ~110 assayed | sub-µM non-hydroxamate hits | peer-reviewed |
| A-Lab25 | Inorganic synthesis | Robotic, autonomous | 17 days | 353 experiments | 36 of 57 targets (63%) | corrected 2026 |
| Amazon Bio Discovery | Nanobodies, DSRCT | Agent-orchestrated | — | ~288,000 designed | KD 0.66–305 nM (46 fits) | preprint |
| Ginkgo × OpenAI33 | Cell-free protein synthesis | GPT-5 agent, human protocols | 6 steps | 29,527 compositions | −40% cost; +27% titer | preprint |
| Virtual Lab34 | SARS-CoV-2 nanobodies | LLM agent team + human | — | 92 designed | 2 improved binders | peer-reviewed |
2.8 The corrections record
Treated as data rather than as gossip, the record of independent re-examination is the most informative result in this review, because it is unanimous in direction.
The A-Lab. The 2023 report claimed 41 novel compounds from 58 targets in 17 days. Leeman and colleagues re-analysed all 43 products — 36 successes plus seven partial successes — and "found significant issues with 42 of them", concluding that on a defensible reading "we could agree that three materials were correctly synthesized… the success rate would be 3/58, or 5%, which is far away from the claims in the paper."26 Two thirds of the claimed successes, they argue, were likely known compositionally disordered versions of the predicted ordered compounds; the systemic causes they name are that automated Rietveld analysis of powder diffraction "is not yet reliable" and that DFT-based prediction neglects disorder — which implicates the prediction databases, not only the robot. In January 2026 Nature published an Author Correction.25 The headline became 36 compounds from 57 targets; "discovery" became "synthesis"; one compound was removed as training-data contamination; and the authors acknowledged that "the original claims of material novelty were subject to misinterpretation — their intention was to indicate that the materials were new to the prediction platform, not necessarily new to science." Manual re-analysis confirmed the platform's call in 36 of 40 reported successes, with four inconclusive, and one compound was removed because it "was mistakenly included in the training data" — training-data contamination inside the flagship autonomous-discovery paper. The paper stands corrected, not retracted. The full arc is worth carrying as a unit: a claim of 41 novel compounds, a critique arguing for as few as three, and a correction that withdrew the novelty claim entirely.
MULTI-evolve. Visani, Verma and DeWitt report that the method's neural network predictions are almost perfectly correlated with an additive model's across all three engineering applications, that it does not outperform a simple additive baseline on held-out data, and that the apparent benefit of adding higher-order variants to training "emerges under null additive models" as an ordinary larger-training-set effect.27 Their conclusion is that the engineering reduces to combining the largest additive effects — "a standard protein engineering strategy for over four decades." As of this writing there is no published response from the original authors: no Science eLetter, no Technical Comment, no rebuttal preprint. It is worth being precise about the target: the critique disputes the mechanistic interpretation, not the wet-lab fold-improvements, which stand.
SAMPLE and the assay. The ≥12 °C figure comes from the automated pipeline's own measurements against the six natural parents; independent human re-measurement under different expression and assay conditions put the best agents at roughly 10 °C over the best natural parent.12 The authors report this themselves, which is to their credit, and it quantifies something important: the loop's own instrument is part of the claim.
Dodecanol. Discussed above: a real 21% gain, a model with negative out-of-sample R², and recommendations that could not be physically realised.19
2.9 When someone runs the baseline
The corrections above are retrospective. A separate literature runs the comparison prospectively, and it is the most decision-relevant body of work for anyone planning a campaign.
The one wet-lab-validated negative. Chinery and colleagues generated antibody library designs by five computational methods and validated 720 of them by surface plasmon resonance. "The BLOSUM substitution matrix outperformed all four deep learning design approaches tested, achieving an estimated minimum binder enrichment of 12.5% and producing nine sub-nanomolar binders. These results underscore the importance of comparing against simple baselines."36 A substitution matrix from 1992 beat four deep models on a trastuzumab–HER2 system, with the designs actually made and measured.
How much data before ML earns its place. The most careful evaluation of ML-assisted directed evolution — from the Arnold lab, and net-positive on the method — quantifies the crossover across sixteen exhaustively simulated landscapes.37 MLDE needed about 48 training samples to beat recombination DE and 96 to beat single-step DE, but "384 to achieve a comparable fraction reaching the global optimum as the most competitive DE strategy". Against a well-chosen classical baseline, several hundred labelled variants buy parity, not superiority. The same paper reports that learned PLM representations "showed minimal to no improvement over one-hot encoding" on landscapes with at least 1% active variants — a finding that should temper the assumption that a foundation model is the obvious front end.
Adjacent fields, same result. On gene-perturbation prediction, seven deep models including scGPT, Geneformer and GEARS failed to beat trivial baselines; on held-out double perturbations "all models had a prediction error substantially higher than the additive baseline", and for most genes several models' predictions "did not vary across perturbations" at all.38 A rebuttal argues the negative result stems partly from poorly calibrated metrics rather than the models, and is worth reading alongside it.39 In offline model-based optimisation, seven of ten sophisticated methods failed to beat the best design already present in the ChEMBL training set.40 And in a blinded industrial competition on antibody developability — 113 teams, 25 countries, 80 held-out clinical antibodies — the best Spearman correlations reached 0.708 for hydrophobicity but only 0.310–0.392 for titer, self-association, polyreactivity and thermostability, with cross-validation scores "consistently exceed[ing] held-out test performance, indicating overfitting and limited out-of-distribution generalization".41
What the acceleration numbers rest on. The only systematic attempt to quantify self-driving-lab speed-up reports a median acceleration factor of about 6 (range 1.3–100) across 33 cases.42 Read the table: 19 of the 33 are retrospective dataset replays and 7 computational, leaving 7 experimental; nearly every baseline is random search, grid or Latin hypercube rather than a competent human; and not one case is biological. There is, at present, no acceleration-factor meta-analysis for the life sciences at all. Where a human comparison does exist — 50 expert chemists against Bayesian optimisation on a 10-dimensional reaction — the humans were better for the first five experiments and the optimiser overtook them by about the fifteenth.43
Is the feedback doing anything? The sharpest test yet run on agentic loops asks whether the loop is load-bearing at all. Across seventeen closed-loop tasks derived from CRISPR screens, four LLM agents each beat the strongest non-agent baseline on at least fifteen — but controlled comparisons found no consistent advantage of true feedback over random or absent feedback, and of 576 round-to-round transitions only 43 (7.5%) completed a full feedback→state→action→outcome chain, 25 of them under random feedback.57 The authors' conclusion is the one that should be pinned above every agentic-science demo: "high final recall does not necessarily indicate effective feedback use." An agent can post good final numbers while the loop itself is decorative.
A related caution applies to the benchmarks these systems are scored on. An independent audit of one autonomous-scientist system tested three of its hypotheses against random-gene null models and found one well supported, one uncertain and one indistinguishable from random.58 And an expert review of chemistry and biology items in a widely used reasoning benchmark found roughly 29% had directly conflicting published evidence — which propagates silently into every agent score quoted from it.
The common structure
In each case the reported gain was real in the sense that something measurable changed. What failed was the attribution — of novelty to the platform, of improvement to epistasis learning, of a temperature margin to the designs rather than the assay, of a titer increase to the model. Closed loops are unusually good at producing numbers and unusually bad at telling you where the numbers came from, because the system that generates the hypothesis also generates the measurement that evaluates it.
3Perspective
Four arguments, and a reporting standard.
3.1 The measurement is the model
A closed loop optimises whatever its assay returns. This is obvious and routinely forgotten, because the assay is the least glamorous component and the one most likely to be treated as ground truth. SAMPLE's margin moved by 2 °C between its own pipeline and human re-measurement. The A-Lab's novelty claim collapsed not because its models were wrong about thermodynamics but because automated Rietveld refinement of powder diffraction was not good enough to support the phase identifications, and no human looked until outsiders did. The dodecanol campaign's models could not be evaluated properly because the build step could not hit its own targets.
The analysis layer belongs inside this argument too, and the agentic era sharpens it. When Robin's automated pipeline and a human analyst processed the same flow-cytometry data, they reported 7.5-fold and 1.75-fold respectively.54 Neither number is fabricated; an analysis path with latitude in gating, normalisation and exclusion produces a distribution of defensible answers, and an automated one embedded in a loop that rewards improvement has no particular reason to land at the conservative end of it. In a closed loop, the analysis code is as much a part of the instrument as the plate reader.
The implication for practice is unglamorous: in a closed loop, assay development is model development, and orthogonal validation — on a different instrument, and by a second analysis path — is not optional polish but part of the claim. The field's most sophisticated review of its own protein-engineering practice argues exactly this, for transparent, accessible and reproducible platforms with FAIR data.28
3.2 Report the cheap baseline, always
The MULTI-evolve dispute is the field's most useful argument because it is about an omission that is cheap to remedy. If a model's multi-mutant predictions correlate with an additive model at r > 0.999, then whatever the campaign achieved, it did not achieve it by learning epistasis — and the additive baseline costs nothing to run because it fits on data the campaign already collected.
Generalise it. For sequence design the baselines are the additive model, a substitution matrix, a greedy walk over single-mutant effects, and random mutagenesis at matched screening budget. For strain engineering it is best-of-library selection at equal throughput. For any autonomous platform it is a competent human given the same instrument-hours. None of these is hard. Their absence is what allows a 21% titer gain to be reported alongside a model with R² of −0.29 without the tension being visible. The right convention is the one Visani and colleagues propose: benchmark against additive baselines before attributing performance to anything more sophisticated.27
The empirical case for this is no longer speculative. When 720 antibody designs were actually made and measured, a BLOSUM substitution matrix beat four deep learning methods.36 When sixteen landscapes were exhaustively enumerated, ML-assisted evolution needed roughly 384 labelled variants merely to match the strongest classical strategy, and PLM embeddings added essentially nothing over one-hot encoding.37 When a blinded industrial split was used, the best of 113 teams reached Spearman ρ ≈ 0.31–0.39 on four of five developability properties.41 The pattern is consistent: performance estimated on the loop's own data is systematically optimistic relative to performance measured against a fair comparator.
A corollary applies to how the agentic literature evaluates itself. Expert Likert-scale ratings of hypothesis novelty — the standard evidence that an agent proposes better experiments — do not survive execution. In a controlled study, 43 researchers each spent over 100 hours executing an idea written either by an expert or by an LLM; on blind review after execution, the LLM ideas' scores dropped significantly more than the human ones on novelty, excitement, effectiveness and overall.44 The study was NLP rather than biology and was underpowered for a direct human-versus-model verdict, so it should not be over-read; but the differential drop is exactly the failure mode a lab-in-the-loop evaluation should fear, and it converges with the finding that an agent's own design scoring did not correlate with design quality.33
3.3 Autonomy is the wrong axis
Read Table 1 by the "loop closed by" column. The fully autonomous biological platforms — SAMPLE, the UIUC biofoundry, the Zhang biofoundry — report ≥12 °C, 90-fold and 2.4-fold. The human-executed loops — the Genentech antibody campaign, ALDE, EVOLVEpro — report 3–100×, 12%→93%, and up to 100×. The largest reported gains are in the loops with people in them.
This is not an argument against automation, which buys reproducibility, throughput and provenance. It is an argument that autonomy level is a measure of engineering maturity, not of scientific yield, and that the two are routinely conflated in both marketing and review articles. The useful question is not "did a human touch it" but "how many hypotheses per unit of experimental budget, and against what baseline". Note also that autonomy currently trades against verification: the platforms with no human in the decision path are exactly the ones whose claims went uninspected longest.
3.4 The evidence base is inverted
Of the flagship results that circulate as proof that lab-in-the-loop works, the most impressive are the least reviewed. The Genentech campaign — the paper that put the term into the literature — has been a preprint for nineteen months. Amazon's Bio Discovery numbers rest on a preprint and a workshop paper, with no confirmation from either named partner, and its "weeks versus up to a year" speed-up has no independent benchmark. Nabla's strongest figures are in a company PDF with no DOI, deposited nowhere. Genentech's DeepFitness and DyAb fold-improvements exist in investor decks. Meanwhile the peer-reviewed tier — SAMPLE, UIUC, ALDE, ART, the Jensen platform — reports smaller, better-characterised numbers.
A reader assembling a mental model of the field from citation counts will therefore systematically overestimate it. This is a predictable consequence of industrial labs publishing on their own timetable, and it is not fraud; but a review that repeats those numbers without their tier is doing the reader a disservice. Hence the pills in Table 1.
3.5 The protocol layer is the missing half of reproducibility
An argument for re-measurement is worth little if the protocol cannot travel. Here the infrastructure record is split cleanly in two, and the halves are moving in opposite directions.
The communication and data layer is healthy. SiLA 2 ships active implementations, OPC UA LADS has broad instrument-vendor backing, and Allotrope's simple model format is genuinely open and actively tooled. In July 2026 a group of pharmaceutical and agricultural buyers — Roche, Bayer, Takeda, Novo Nordisk, BioNTech, Syngenta, Lonza among them — signed a letter of intent demanding native, on-device support for these open standards, stating that reliance on driver PCs and middleware wrappers is "a legacy solution" and that recurring fees for instrument APIs "are not an option", with progress expected by 2027. That is procurement pressure of a kind the field has never had.
The protocol layer, which is what would actually let one lab re-run another's experiment, is close to dead. Autoprotocol is abandoned — its last release was April 2023 and its site still carries the copyright of a company that no longer operates. LabOP is dormant and its domain no longer resolves. AnIML remains a draft at version 0.90 after two decades. SBOL's specification has been static since 2023. The one thriving open project is PyLabRobot, a hardware-agnostic interface from the same MIT group that works on biosecurity, which is actively developed and increasingly cited by the agentic-science literature.59 The most substantial collective attempt to make biofoundry work interoperable is a 2025 abstraction hierarchy for workflows and unit operations.60
One consequence deserves to be stated plainly, because it is the field's sharpest reproducibility fact. The Strateos cloud lab folded in 2023.61 It is the facility on which SAMPLE — the cleanest autonomous protein-engineering campaign in this review — was run. Autoprotocol, the language that encoded such experiments, died with it. A landmark closed-loop result is therefore not re-runnable as executed, not because anyone withheld anything, but because the execution substrate was commercial and is gone. Synthetic biology's own assessment of its data practice is that assets remain "scattered across spreadsheets, proprietary exports, custom scripts".62 Chemistry has an open reaction database with schema, corpus and governance; biology has no equivalent.
3.6 Execution is running ahead of ideation, and governance is attached to the wrong end
The capability that has advanced fastest is not hypothesis generation. It is execution at the digital-to-physical boundary. On an agentic bio-capabilities benchmark, all eight tested models beat the median human expert at operating liquid handlers, designing DNA fragments and evading synthesis screening; one model's generated OpenTrons script performed Gibson assembly successfully in three of three wet-lab runs, verified by whole-plasmid sequencing.45 Separately, a study of 76,089 redesigned variants of 72 proteins of concern found that AI-redesigned sequences could evade nucleic-acid synthesis screening, with patches subsequently deployed to providers.46
Two things follow. First, the honest read of the risk literature is narrower than either enthusiasts or alarmists suggest: a randomised red-team study found no statistically significant difference in the viability of biological attack plans generated with or without LLM assistance, but it covered planning only, used 2023-era models, and was underpowered.47 The National Academies' assessment is similarly bounded — current tools cannot design a virus de novo, but can design simple biomolecules such as toxins that may evade security checkpoints.48 Neither finding is about the protocol-writing capability that benchmarks now show is ahead of expert humans. The most useful peer-reviewed treatment of this asymmetry argues for prioritising safeguarding over autonomy in AI-driven science32 — which, read alongside the execution benchmarks, is less a precautionary sentiment than a description of where the capability frontier actually is.
Second, and more prosaically, the field's legal and economic scaffolding is unsettled in ways that bear directly on who can run a loop at all. The USPTO rescinded its 2024 AI-inventorship guidance in November 2025, leaving the status of inventions from autonomously operating systems unresolved,49 which matters because, as one policy review puts it, "if the inventions they generate remain unpatentable, funding for SDLs may be constrained."31 Meanwhile access is priced out of reach of most academic groups: general cloud-lab access has been reported at over $250,000, or above $100,000 to automate and run a single method, on minimum one-year contracts.50 A field whose results cannot be independently re-measured because almost no one else can afford the instrument is a field that will keep discovering its errors late.
A minimal reporting standard for lab-in-the-loop claims
- The baseline. Report the additive (or best-of-library) model's performance on the same data, at the same budget. If the loop does not beat it, say so.
- The assay. State measurement error and report an orthogonal re-measurement of the top designs on a different instrument or in a different lab.
- The denominator. Report designs proposed, designs built, designs that failed to build, and designs assayed — not only the successes.
- The budget. Report wall-clock time and variants characterised per round, so sample efficiency is comparable across campaigns.
- The decision path. State explicitly which decisions were made by the model and which by people; give the autonomy level on a named scale.7
- The reference point. Compare against the best conventional result for the same target, not against the campaign's own starting seed.
4Conclusion
Lab-in-the-loop is real, and its genuine achievement is narrower and more valuable than the headline fold-improvements suggest: it has moved protein engineering from screening millions to characterising hundreds. Four rounds and fewer than 500 variants per enzyme, one round and ~120 variants for a seven-mutation multi-mutant, four rounds and 1,800 variants for sub-100 pM antibodies against four targets — that is a two-to-four-order-of-magnitude reduction in experimental cost, achieved by putting a decent prior in front of the assay. It is enough to change what a small lab can attempt.
What the field has not yet built is the habit of verification that would let those numbers be trusted at face value. Every closed-loop result that has been independently re-examined has come back smaller: 41 novel compounds became 36 synthesised and none demonstrably new; an epistasis model became an additive one; 12 °C became 10 °C. The corrections were, in each case, produced by outsiders — a critic's thread that became a PRX Energy paper, a baseline analysis from an unaffiliated lab — rather than by the platforms themselves. That is a structural finding about incentives, not a comment on any individual group's integrity. It is reinforced by an absence: there is still no peer-reviewed cross-laboratory reproducibility study for self-driving labs in biology, and cloud-lab access at six figures a year ensures that very few groups are positioned to produce one.
The timing matters, because the field is about to be scaled. A US executive order in November 2025 made autonomous laboratories a national challenge; by July 2026 more than $5 billion had been committed across fifteen federal agencies, the National Science Foundation had turned a $100 million solicitation into $380 million across twenty programmable cloud-lab nodes — several of them explicitly for protein engineering and biomanufacturing — and ARPA-H had launched a programme built around "a distributed marketplace of validated laboratories". The European Union's RAISE instrument has a forthcoming €29 million call for closed-loop experimentation, though the life sciences are not among its named domains. Standards that do not exist when money arrives at this scale tend not to be retrofitted.
It is worth noting that this argument is no longer contrarian. A July 2026 community roadmap on trustworthy autonomous science — written by the people building these systems, and citing the corrected flagship result discussed above — states the position more compactly than I have: "producing a candidate discovery is no longer the hard part, but verifying it is, and this asymmetry now limits autonomous science more than raw model capability."63 Between its 2025 and 2026 editions, that roadmap promoted trust, verification and reproducibility from a cross-cutting concern to a first-class dimension of the field. The diagnosis is converging faster than the practice.
So the next three years should be less about autonomy levels and more about the two things at the bottom of Figure 2. Report the baseline. Re-measure on an orthogonal instrument, and by a second analysis path. Publish the denominator, including what failed to build. Keep the protocol executable by someone who does not own your robot. The methods work — the fact that Nature corrected rather than retracted the A-Lab paper is evidence that the machinery of self-correction functions, slowly. A field that adopts the reporting standard above will still produce 90-fold enzymes and picomolar binders. It will simply be able to say how much of that came from the model, which is the claim everyone is actually interested in.
5References
- King RD, Whelan KE, Jones FM, et al. Functional genomic hypothesis generation and experimentation by a robot scientist. Nature 427(6971):247–252 (2004). doi:10.1038/nature02236
- King RD, Rowland J, Oliver SG, et al. The Automation of Science. Science 324(5923):85–89 (2009). doi:10.1126/science.1165620 — the paper in which the robot is named Adam.
- Williams K, Bilsland E, Sparkes A, et al. Cheaper faster drug development validated by the repositioning of drugs against neglected tropical diseases. J R Soc Interface 12(104):20141289 (2015). doi:10.1098/rsif.2014.1289
- Yang KK, Wu Z, Arnold FH. Machine-learning-guided directed evolution for protein engineering. Nat Methods 16(8):687–694 (2019). doi:10.1038/s41592-019-0496-6
- Genentech. Genentech and NVIDIA announce collaboration (21 Nov 2023), gene.com; and Roche Pharma Day, Lab-in-the-Loop: Embedding AI from target discovery to the clinic (22 Sep 2025). company
- Frey NC, Hötzel I, Stanton SD, et al.; Gligorijević V. Lab-in-the-loop therapeutic antibody design with deep learning. bioRxiv 2025.02.19.639050, v3 posted 8 Nov 2025. doi:10.1101/2025.02.19.639050 preprint — not peer-reviewed as of 22 Sep 2026
- Beal J, Rogers M. Levels of autonomy in synthetic biology engineering. Mol Syst Biol 16(12):e10019 (2020). doi:10.15252/msb.202010019
- Garcia Martin H, Radivojevic T, Zucker J, et al. Perspectives for self-driving labs in synthetic biology. Curr Opin Biotechnol 79:102881 (2023). doi:10.1016/j.copbio.2022.102881
- Agarwal AA, et al. AlphaBind, a domain-specific model to predict and optimize antibody–antigen binding affinity. mAbs 17(1):2534626 (2025). doi:10.1080/19420862.2025.2534626
- Borriello F, et al. Randomized, double-blind, placebo-controlled first-in-human trial of a first-in-class AI-designed monoclonal antibody (GB-0669). J Infect Dis jiag349 (2026). doi:10.1093/infdis/jiag349
- Xu Z, Ren F, Wang P, et al.; Zhavoronkov A. A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trial. Nat Med 31(8):2602–2610 (2025). doi:10.1038/s41591-025-03743-2 — primary endpoint was safety. Target discovery and phase 1: Ren F, et al. Nat Biotechnol 43(1):63–75 (2025). doi:10.1038/s41587-024-02143-0
- Rapp JT, Bremer BJ, Romero PA. Self-driving laboratories to autonomously navigate the protein fitness landscape. Nat Chem Eng 1(1):97–107 (2024). doi:10.1038/s44286-023-00002-4
- Singh N, Lane S, Yu T, et al.; Zhao H. A generalized platform for artificial intelligence-powered autonomous enzyme engineering. Nat Commun 16:5648 (2025). doi:10.1038/s41467-025-61209-y
- Zhang Q, Chen W, Qin M, et al. Integrating protein language models and automatic biofoundry for enhanced protein evolution. Nat Commun 16:1553 (2025). doi:10.1038/s41467-025-56751-8
- Yang J, Lal RG, Bowden JC, et al.; Arnold FH. Active learning-assisted directed evolution. Nat Commun 16:714 (2025). doi:10.1038/s41467-025-55987-8
- Jiang K, Yan Z, Di Bernardo M, et al.; Gootenberg JS, Abudayyeh OO. Rapid in silico directed evolution by a protein language model with EVOLVEpro. Science 387(6732):eadr6006 (2025). doi:10.1126/science.adr6006
- Tran VQ, Nemeth M, Bartie LJ, et al.; Konermann S, Hsu PD. Rapid directed evolution guided by protein language models and epistatic interactions. Science 392(6798):eaea1820 (7 May 2026). doi:10.1126/science.aea1820
- Radivojević T, Costello Z, Workman K, Garcia Martin H. A machine learning Automated Recommendation Tool for synthetic biology. Nat Commun 11:4879 (2020). doi:10.1038/s41467-020-18008-4
- Opgenorth P, Costello Z, Okada T, et al.; Garcia Martin H, Beller HR. Lessons from two design–build–test–learn cycles of dodecanol production in Escherichia coli aided by machine learning. ACS Synth Biol 8(6):1337–1351 (2019). doi:10.1021/acssynbio.9b00020. The R² = −0.29 re-analysis appears in ref. 18.
- van Lent P, Moreno Paz S, Schmitz J, Abeel T. Comparing metabolic engineering scenarios using simulated design-build-test-learn cycles. Front Bioeng Biotechnol 14:1802948 (2026). doi:10.3389/fbioe.2026.1802948
- van Lent P, Schmitz J, Abeel T. Simulated design–build–test–learn cycles for consistent comparison of machine learning methods in metabolic engineering. ACS Synth Biol 12(9):2588–2599 (2023). doi:10.1021/acssynbio.3c00186
- Koscher BA, Canty RB, McDonald MA, et al.; Jensen KF. Autonomous, multiproperty-driven molecular discovery: from predictions to measurements and back. Science 382(6677):eadi1407 (2023). doi:10.1126/science.adi1407
- McDonald MA, Koscher BA, Canty RB, et al.; Jensen KF. Bayesian optimization over multiple experimental fidelities accelerates automated discovery of drug molecules. ACS Cent Sci 11(2):346–356 (2025). doi:10.1021/acscentsci.4c01991
- Boiko DA, MacKnight R, Kline B, Gomes G. Autonomous chemical research with large language models. Nature 624(7992):570–578 (2023). doi:10.1038/s41586-023-06792-0
- Szymanski NJ, Rendy B, Fei Y, et al.; Ceder G. An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature 624(7990):86–91 (2023). doi:10.1038/s41586-023-06734-w. Author Correction: Nature 650(8100):E1 (2026). doi:10.1038/s41586-025-09992-y
- Leeman J, Liu Y, Stiles J, et al.; Schoop LM, Palgrave RG. Challenges in high-throughput inorganic materials prediction and autonomous synthesis. PRX Energy 3(1):011002 (2024). doi:10.1103/PRXEnergy.3.011002
- Visani GM, Verma A, DeWitt WS. Additive baselines furnish no evidence for epistasis learning by MULTI-evolve. bioRxiv 2026.04.23.719915 (24 Apr 2026). doi:10.64898/2026.04.23.719915 preprint
- Weigmann KFG, Bornscheuer UT, Doerr M. Advances and critical evaluation of autonomous protein engineering: towards transparent, accessible, and reproducible platforms. Curr Opin Biotechnol 97:103395 (2026). doi:10.1016/j.copbio.2025.103395
- Tom G, Schmid SP, Baird SG, et al.; Aspuru-Guzik A. Self-driving laboratories for chemistry and materials science. Chem Rev 124(16):9633–9732 (2024). doi:10.1021/acs.chemrev.4c00055
- Canty RB, Abolhasani M. The past, present and future of self-driving laboratories. Nat Rev Chem 10(8):523–537 (2026). doi:10.1038/s41570-026-00847-2
- Tobias AV, Wahab A. Autonomous "self-driving" laboratories: a review of technology and policy implications. R Soc Open Sci 12(7):250646 (2025). doi:10.1098/rsos.250646
- Tang X, et al. Risks of AI scientists: prioritizing safeguarding over autonomy. Nat Commun 16:8317 (2025). doi:10.1038/s41467-025-63913-1
- Smith AA, Wong EL, Donovan RC, et al. (Ginkgo Bioworks & OpenAI). Using a GPT-5-driven autonomous lab to optimize the cost and titer of cell-free protein synthesis. bioRxiv 2026.02.05.703998 (2026). doi:10.64898/2026.02.05.703998 preprint; all authors are company staff
- Swanson K, Wu W, Bulaong NL, Pak JE, Zou J. The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature 646:716–723 (2025). doi:10.1038/s41586-025-09442-9
- Gottweis J, Weng W-H, Daryin A, et al.; Natarajan V. Accelerating scientific discovery with Co-Scientist. Nature 655:487–496 (2026). doi:10.1038/s41586-026-10644-y
- Chinery L, Hummer AM, Mehta BB, et al.; Greiff V, Jeliazkov JR, Deane CM. Simple computational methods can outperform deep learning in designing diverse, binder-enriched antibody libraries. bioRxiv 2024.03.26.586756, v2 (2026). doi:10.1101/2024.03.26.586756 preprint
- Li F-Z, Yang J, Johnston KE, Gursoy E, Yue Y, Arnold FH. Evaluation of machine learning-assisted directed evolution across diverse combinatorial landscapes. Cell Syst 16(9):101387 (2025). PMID 40934912.
- Ahlmann-Eltze C, Huber W, Anders S. Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines. Nat Methods 22:1657–1661 (2025). doi:10.1038/s41592-025-02772-6
- Miller HE, Mejia GM, Leblanc FJA, Wang B, et al. Deep learning-based genetic perturbation models do outperform uninformative baselines on well-calibrated metrics. bioRxiv 2025.10.20.683304 (2025). doi:10.1101/2025.10.20.683304 preprint — rebuttal to ref. 38
- Trabucco B, Geng X, Kumar A, Levine S. Design-Bench: benchmarks for data-driven offline model-based optimization. Proc. ICML PMLR 162 (2022). arXiv:2202.08450
- van Niekerk L, Moller J, Ritter S, et al.; Deane CM, Tessier PM, Arsiwala A. Ginkgo Datapoints antibody developability competition outcomes: limited model performance and a call for data standardization. mAbs (2026). doi:10.1080/19420862.2026.2634216
- Adesiji AD, Wang J, Kuo C-S, Brown KA. Benchmarking self-driving labs. Digital Discovery (2026); arXiv:2508.06642. doi:10.1039/D5DD00337G — of 33 reported acceleration factors, 19 are retrospective replays, 7 computational and 7 experimental; none is biological.
- Shields BJ, Stevens J, Li J, et al.; Doyle AG. Bayesian reaction optimization as a tool for chemical synthesis. Nature 590:89–96 (2021). doi:10.1038/s41586-021-03213-y
- Si C, Hashimoto T, Yang D. The ideation-execution gap: execution outcomes of LLM-generated versus human research ideas. Proc. ICLR (2026); arXiv:2506.20803
- Liu AB, Nedungadi S, Cai B, et al.; Donoughe S. ABC-Bench: an agentic bio-capabilities benchmark for biosecurity. arXiv:2606.11150 (2026). preprint
- Wittmann BJ, Alexanian T, Bartling C, et al.; Wheeler NE, Horvitz E. Strengthening nucleic acid biosecurity screening against generative protein design tools. Science 390(6768):82–87 (2025). doi:10.1126/science.adu8578
- Mouton CA, Lucas C, Guest E. The operational risks of AI in large-scale biological attacks. RAND RR-A2977-2 (2024). doi:10.7249/RRA2977-2 — a null result on planning, with 2023-era models and limited power.
- National Academies of Sciences, Engineering, and Medicine. The age of AI in the life sciences: benefits and biosecurity considerations. National Academies Press (2025). doi:10.17226/28868
- USPTO. Revised inventorship guidance for AI-assisted inventions. Fed. Reg. doc. 2025-21457, 28 Nov 2025 (rescinding the Feb 2024 guidance). Analysis: Aboy M, Liddell K. J Intellect Prop Law Pract 21(5):272–275 (2026). doi:10.1093/jiplp/jpag021
- Armer C, Letronne F, DeBenedictis E. Support academic access to automated cloud labs to improve reproducibility. PLoS Biol 21(1):e3001919 (2023). doi:10.1371/journal.pbio.3001919
- Guan Y, Cui L, Inchai J, et al.; Peltz G. AI-assisted drug re-purposing for human liver fibrosis. Adv Sci 12(44):e08751 (2025). doi:10.1002/advs.202508751
- Penadés JR, Gottweis J, He L, et al.; Costa TRD. AI mirrors experimental science to uncover a mechanism of gene transfer crucial to bacterial evolution. Cell 188(23):6654–6665.e2 (2025). doi:10.1016/j.cell.2025.08.018 — the preprint title said "a novel mechanism"; peer review removed "novel".
- He L, Patkowski JB, Wang J, et al.; Penadés JR. Chimeric infective particles expand species boundaries in phage-inducible chromosomal island mobilization. Cell 188(23):6636–6653.e17 (2025). doi:10.1016/j.cell.2025.08.019 — the experimental work that preceded the AI hypothesis; preprint posted 11 Feb 2025.
- Ghareeb AE, Chang B, Mitchener L, et al.; Finnemann SC, Rodriques SG. A multi-agent system for automating scientific discovery. Nature 655:497–505 (2026). doi:10.1038/s41586-026-10652-y. The 7.5× vs 1.75× analysis discrepancy is reported in the preprint's supplementary material (arXiv:2505.13400).
- Qu Y, Huang K, Yin M, et al.; Cong L. CRISPR-GPT for agentic automation of gene-editing experiments. Nat Biomed Eng 10:245–258 (2026). doi:10.1038/s41551-025-01463-z
- Huang K, Zhang S, Wang H, et al.; Regev A, Leskovec J. Autonomous biomedical research with an artificial intelligence agent. Science 393:eadz4351 (2026). doi:10.1126/science.adz4351. Wet-lab content described here is from the open preprint (bioRxiv 2025.05.30.656746).
- Yu C, Liu S, Qiao G, Luo M, Xiang Y, Xu Z. PerturbTrace: evaluating feedback use by AI co-scientist agents in perturbation discovery. bioRxiv 2026.08.18.745260 (2026). doi:10.64898/2026.08.18.745260 preprint
- Nusrat H, Nusrat O. When AI does science: evaluating the autonomous AI scientist KOSMOS in radiation biology. arXiv:2511.13825 (2025). preprint
- Wierenga RP, Golas SM, Ho W, Coley CW, Esvelt KM. PyLabRobot: an open-source, hardware-agnostic interface for liquid-handling robots and accessories. Device 1(4):100111 (2023). doi:10.1016/j.device.2023.100111
- Kim H, Hillson NJ, Cho B-K, et al.; Freemont PS, Lee S-G. Abstraction hierarchy to define biofoundry workflows and operations for interoperable synthetic biology research and applications. Nat Commun 16:6056 (2025). doi:10.1038/s41467-025-61263-6
- Adam D. The automated lab of tomorrow. Proc Natl Acad Sci USA 121(17):e2406320121 (2024). doi:10.1073/pnas.2406320121 — reports that Strateos, the cloud lab used for the SAMPLE campaign, folded.
- Vitalis C, Vidal G, Samineni SP, Fontanarrosa P, Myers CJ. A framework for a standard-enabled FAIR data management workflow for synthetic biology. ACS Synth Biol 15(1):1–8 (2026). doi:10.1021/acssynbio.5c00813
- Ferreira da Silva R, Abolhasani M, Beaucage P, et al. Toward trustworthy autonomous science: a two-year community roadmap. arXiv:2607.12113 (2026); ORNL/TM-2026/4663. community roadmap