Right Prediction, Wrong Reasoning
Uncovering LLM Misalignment in Rheumatoid Arthritis Disease Diagnosis: AI models can achieve high classification accuracy while generating clinically flawed medical explanations.
Right Prediction, Wrong Reasoning: Uncovering LLM Misalignment in RA Disease Diagnosis
Healthcare AI Safety Clinical Evaluation arXiv:2504.06581
Large Language Models (LLMs) are increasingly being explored for medical diagnostic assistance. However, high benchmark accuracy can be deceptively reassuring. In this clinical study, we demonstrate that LLMs frequently predict the correct Rheumatoid Arthritis (RA) diagnosis while relying on clinically dangerous, hallucinated, or specious justifications.
Direct Links & Resources
- đź“„ arXiv Paper: https://arxiv.org/abs/2504.06581
- đź“‘ PDF: https://arxiv.org/pdf/2504.06581.pdf
Authors & Collaborators
- Umakanta Maharana
- Sarthak Verma
- Avarna Agarwal
- Prakashini Mruthyunjaya
- Dwarikanath Mahapatra
- Sakir Ahmed
- Murari Mandal
Key Findings
- Accuracy vs. Explanation Disconnect: LLMs achieved >90% diagnostic classification accuracy on RA clinical notes, but blinded physician review revealed that over 68% of the generated reasoning steps contained critical factual errors or non-standard diagnostic criteria.
- The “Right for the Wrong Reasons” Risk: A model providing the right diagnosis for flawed reasons undermines clinical trust and poses severe safety risks if deployed without human-in-the-loop expert validation.
- Clinical Alignment Need: Highlights the urgent need for evaluation benchmarks that assess diagnostic justification faithfulness rather than just top-1 accuracy.