Not yet Quality score: 48/100

Effect of Time Since Baseline Data Collection, Mechanism of Injury, and Service Academy Sub-population on Concussion Risk Prediction in Military Service Academy Cadets and Midshipmen: An Observational Study

Sports medicine - open · 2026

Concussion baseline testing predicts risk with just 0.64-0.68 AUC, and that accuracy decays further the longer ago the test was done.

sports-medicineathletic-traininghead-neckconcussioninjury-preventioncohort-observational

The paper

Observational cohort study, n=16,642 cadets/midshipmen from 4 U.S. military service academies (2015-2020), using XGBoost machine learning on baseline assessment data to predict concussion risk.

What they found

Predictive accuracy (AUC) fell from 0.68 at baseline testing to under 0.64 as time since assessment increased, across the overall model and within every subgroup. Prediction was more accurate for intramural cadets and training/PE-related concussions than for varsity/club athletes and sports-related concussions.

The appraisal

This is a large, real-world dataset with a sound machine-learning approach, but an AUC of 0.64-0.68 is weak discrimination — only modestly better than chance — regardless of statistical rigor. No confidence intervals, calibration statistics, or external validation are reported, and the abstract gives no indication of held-out/cross-validation methodology in enough detail to judge overfitting risk. This is exploratory model-building, not a validated clinical prediction tool.

The gap

There's no external validation cohort and no calibration data, so it's unclear whether this model (or its finding about time-decay) generalizes beyond these four academies; the specific baseline variables driving prediction aren't detailed enough to act on.

Landmark context

This extends work from the NCAA-DoD CARE Consortium, the major multi-site initiative behind most modern baseline concussion testing research, and speaks to a long-running debate (echoed in ImPACT test-retest literature) about how much clinical value single-timepoint preseason baseline testing actually retains as time passes.

What to do Monday

No — an AUC below 0.70 isn't accurate enough to guide individual return-to-play or risk-stratification decisions. The actionable takeaway is institutional/policy-level: if baseline testing is used for risk prediction, it should be scheduled close to the season or activity of highest exposure rather than done once early and relied on for a full year.

Read the primary source ↗

The five papers that matter — every week, free.

We read the whole sports-medicine, rehab & performance literature and appraise the papers worth your time.

Get the free weekly issue

Related appraisals

Read in your language: English · Français · Deutsch · Español · Italiano · Nederlands · Dansk · Svenska