Skip to content
TechBio Today

The front page of AI-driven biology.

preprint · bioRxiv

A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models

Joao Sartori and colleagues report that enzymARC tests whether enzyme-function predictors reject structure-guided decoys whose catalytic machinery has been disrupted.

Author affiliations

  • Fiocruz
  • Instituto Oswaldo Cruz
  • Institute of Technology on Immunobiologicals (Bio-Manguinhos), Fiocruz

In plain English

The authors mutate active sites while keeping much of an enzyme sequence recognizable, then ask whether annotation models incorrectly assign the original Enzyme Commission label.

How the study worked

A plain-language walk through the work behind the result.

  1. Generated putatively non-functional decoys with perturbations around catalytic residues.

  2. Evaluated DIAMOND, CLEAN, and DeepEC against the benchmark.

What they found

  • DIAMOND and CLEAN exceeded 90% false-positive rates for low-perturbation decoys.
  • DeepEC improved at larger perturbations, but all tested approaches struggled with targeted active-site disruption.

Why it matters

Benchmarks that include convincing negative examples can expose shortcut learning hidden by tests made only from annotated functional proteins.

The catch

  • This is a preprint and has not been peer reviewed.
  • The decoys are expected to be inactive computationally; they were not experimentally validated as non-functional.

Evidence ledger

Sources behind this brief

  1. 01
    Primary source

    bioRxiv preprint version 1

    preprint · Accessed August 24, 2026

  2. 02
    Supporting context

    EnzymARC repository

    author reported result · Accessed August 24, 2026