In plain English
Protein models found useful candidates in unfamiliar sequence territory, but they did not reliably reproduce the detailed experimental fitness landscape.
How the study worked
A plain-language walk through the work behind the result.
Experimentally measured highly diverse natural and new-to-nature proteins across three protein families.
Compared experimental fitness with Potts-model and protein-language-model scores.
Examined both local landscape shape and broad trends across sequence space.
What they found
- Each family contained functional new-to-nature sequences with low identity to known orthologs.
- Sequence models were useful but inconsistent and did not reliably capture local or global experimental fitness patterns.
Why it matters
Protein design has a larger experimental search space than nature has sampled, but computational scores still need laboratory validation.
The catch
- The record is a preprint and has not completed peer review.
- This summary is based on the abstract; methods and supplementary analyses were not independently rechecked.