In plain English
The AIntibody challenge compared 511 AI-designed or predicted antibodies from 29 organizations under shared experimental affinity and developability assays. Several teams produced developable antibodies below 100 pM affinity, yet those wins were exceptions: performance often failed to transfer to other tasks, and nearly every model performed worse than random clone picking when ranking high-affinity antibodies within HCDR3 clusters.
How the study worked
A plain-language walk through the work behind the result.
Ran three prospective, blinded tasks covering affinity maturation, within-cluster affinity ranking, and out-of-library CDR design.
Synthesized submissions as full-length antibodies and measured affinity and developability under common experimental protocols.
What they found
- Several groups produced developable antibodies with affinities below 100 pM, especially in the affinity-maturation setting.
- Cross-task generalization was weak, and all but one model underperformed random clone picking on the within-cluster ranking task.
Why it matters
The study replaces retrospective benchmark claims with common wet-lab measurements, exposing both credible progress and unresolved gaps in computational antibody discovery.
The catch
- The first challenge iteration evaluated one antigen, SARS-CoV-2 receptor-binding domain.
- The organizing consortium also reports the results, and blinding relied on organizer integrity rather than informatic safeguards.
- The supplied datasets were unusually rich, so the authors describe the results as an upper bound for this target class rather than evidence of general-purpose performance.
