In plain English
Researchers prospectively evaluated SHAKED, a clinical decision-support system built on multiple large language models, across two parallel units of a tertiary emergency department. In a sample of 100 outputs, experts judged 99 clinically appropriate and detected no adverse events. Adoption nevertheless fell from 68% to 30%, and median emergency-department length of stay was 4.9 hours in both groups.
How the study worked
A plain-language walk through the work behind the result.
Ran a four-week DECIDE-AI stage 1 evaluation involving 1,138 patients in two parallel emergency-department units.
Measured clinician use, sampled output appropriateness, safety events, length of stay, and consultation timing.
What they found
- Expert review judged 99 of 100 sampled outputs clinically appropriate, with no detected adverse events.
- Adoption declined from 68% to 30%, and length of stay was unchanged at 4.9 hours.
Why it matters
The study adds prospective workflow evidence to a field dominated by retrospective accuracy tests and shows that technically acceptable output does not guarantee sustained clinical use.
The catch
- This was a stage 1 prospective evaluation, not a randomized efficacy trial.
- The detailed appropriateness review covered a sample of 100 outputs.
- A reported 9.4-minute consultation-cycle reduction was a nonsignificant trend, and the authors say the evidence does not justify deployment.