In plain English
The preprint asks a broad first question before predicting a specific interaction: does the amino-acid sequence around a phosphotyrosine look compatible with the shared recognition rules of the SH2 family at all? The authors train a classifier from diverse binding datasets and use it to annotate more than 45,000 human phosphotyrosine sites, prioritize candidates for enrichment experiments, and analyze mutation effects.
How the study worked
A plain-language walk through the work behind the result.
Integrated phosphopeptide-binding data spanning 120 SH2 domains into a supervised sequence classifier.
Applied the classifier to the human phosphoproteome, individual experiments, super-SH2 enrichment, and mutation-effect analyses.
What they found
- A relatively small, representative binder set was sufficient for the reported shared-family classification task.
- The model produced a global annotation of potential SH2-binding participation across more than 45,000 human phosphotyrosine sites.
Why it matters
A family-level filter could reduce the search space before researchers attempt the harder task of assigning a phosphosite to a particular SH2 domain.
The catch
- The work is a preprint and has not completed peer review.
- The classifier predicts general SH2-binding potential, not a specific domain–phosphosite interaction.
- The abstract does not report full external-validation or subgroup-performance details.