Research2026-09-03

LatchBio released TxBench-Antibody Discovery, a test set for AI agents in therapeutic antibody research. It contains 100 tasks built from published experimental studies, each graded by fixed rules. Tasks span target assessment, assay design, cellular pharmacology, engineering and candidate de-risking. Twenty model and scaffold combinations were run three times per task. The best combination still failed close to half its attempts. Failures repeated across attempts rather than varying randomly. Spending more money, tokens or tool calls did not reliably raise accuracy. Most wrong answers came from framing the science, not from faulty calculation. Agents often computed correctly but answered a slightly different question. Every task was solved by at least one configuration.

What changed

Earlier antibody benchmarks scored predicted molecular properties or designed candidates.

What it unlocks

Comparing AI agents on whether they reach defensible antibody discovery decisions.

  • 53.0% best pass rate
  • 100 evaluations, 20 configurations
  • 82.3% of failures were interpretation
  • 34.7% scope or endpoint errors

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.