In plain English
A model cannot consistently predict experiments that do not consistently agree with one another.
How the study worked
A plain-language walk through the work behind the result.
Derived a metric-dependent ceiling from agreement among repeated biological measurements.
Measured concordance across MaveDB, ProteinGym, drug-screen, CRISPR-cell-line, and protease datasets.
Compared published predictor scores with the performance available below each estimated ceiling.
What they found
- Two assays of one target agreed at correlations of 0.56 to 0.68 in the analyzed registries.
- Published predictors reached 63% of achievable correlation performance and 18% of achievable top-1% selection performance in the reported analysis.
Why it matters
Biological-AI benchmarks may misstate model headroom when they ignore disagreement among the experiments used as ground truth.
The catch
- The record is a preprint and has not completed peer review.
- This summary is based on the abstract; methods and supplementary analyses were not independently rechecked.