LatchBio released TxBench-Antibody Discovery, a test set for AI agents in therapeutic antibody research. It contains 100 tasks built from published experimental studies, each graded by fixed rules. Tasks span target assessment, assay design, cellular pharmacology, engineering and candidate de-risking. Twenty model and scaffold combinations were run three times per task. The best combination still failed close to half its attempts. Failures repeated across attempts rather than varying randomly. Spending more money, tokens or tool calls did not reliably raise accuracy. Most wrong answers came from framing the science, not from faulty calculation. Agents often computed correctly but answered a slightly different question. Every task was solved by at least one configuration.
What changed
Earlier antibody benchmarks scored predicted molecular properties or designed candidates.
What it unlocks
Comparing AI agents on whether they reach defensible antibody discovery decisions.
- 53.0% best pass rate
- 100 evaluations, 20 configurations
- 82.3% of failures were interpretation
- 34.7% scope or endpoint errors
Sources