In plain English
The study pairs case-control prediction with phenotype clustering, showing that a stable overall score can hide weaker detection for a clinically distinct, lightly documented group.
How the study worked
A plain-language walk through the work behind the result.
Matched 33,739 peripheral-artery-disease cases with 33,739 controls across five health systems.
Trained LightGBM on 14,023 EHR-derived features and assessed demographic and phenotype subgroups.
What they found
- Reported AUROC and AUPRC ranged from 0.76 to 0.79 across institutions.
- Sensitivity ranged from 0.87 in one phenotype cluster to 0.40 in the sparsely documented cluster.
Why it matters
Phenotype-level audits can reveal deployment failure modes that pooled fairness and performance summaries miss.
The catch
- The manuscript is a preprint.
- The private EHR data are unavailable publicly, limiting independent reproduction.