In plain English
The authors mutate active sites while keeping much of an enzyme sequence recognizable, then ask whether annotation models incorrectly assign the original Enzyme Commission label.
How the study worked
A plain-language walk through the work behind the result.
Generated putatively non-functional decoys with perturbations around catalytic residues.
Evaluated DIAMOND, CLEAN, and DeepEC against the benchmark.
What they found
- DIAMOND and CLEAN exceeded 90% false-positive rates for low-perturbation decoys.
- DeepEC improved at larger perturbations, but all tested approaches struggled with targeted active-site disruption.
Why it matters
Benchmarks that include convincing negative examples can expose shortcut learning hidden by tests made only from annotated functional proteins.
The catch
- This is a preprint and has not been peer reviewed.
- The decoys are expected to be inactive computationally; they were not experimentally validated as non-functional.