Reasoning or reciting? A temporal contamination audit of large language models in clinical medicine.
Sheppert Alexander P, Adams Brian, Sheppert Andrew D, +1 more·Journal of the American Medical Informatics Association : JAMIA
OBJECTIVE: Evaluate whether large language models reason or simply regurgitate training data in clinical diagnosis. MATERIALS AND METHODS: We audited 2000 clinical case reports from PubMed Central: 1000 from 2021 to 2022 (within training data) and 1000 from 2025 (after training cutoffs). Five frontier LLMs generated d…