Right Prediction, Wrong Reasoning

Uncovering LLM Misalignment in Rheumatoid Arthritis Disease Diagnosis: AI models can achieve high classification accuracy while generating clinically flawed medical explanations.

Right Prediction, Wrong Reasoning: Uncovering LLM Misalignment in RA Disease Diagnosis

Healthcare AI Safety Clinical Evaluation arXiv:2504.06581

Large Language Models (LLMs) are increasingly being explored for medical diagnostic assistance. However, high benchmark accuracy can be deceptively reassuring. In this clinical study, we demonstrate that LLMs frequently predict the correct Rheumatoid Arthritis (RA) diagnosis while relying on clinically dangerous, hallucinated, or specious justifications.



Authors & Collaborators


Key Findings

  • Accuracy vs. Explanation Disconnect: LLMs achieved >90% diagnostic classification accuracy on RA clinical notes, but blinded physician review revealed that over 68% of the generated reasoning steps contained critical factual errors or non-standard diagnostic criteria.
  • The “Right for the Wrong Reasons” Risk: A model providing the right diagnosis for flawed reasons undermines clinical trust and poses severe safety risks if deployed without human-in-the-loop expert validation.
  • Clinical Alignment Need: Highlights the urgent need for evaluation benchmarks that assess diagnostic justification faithfulness rather than just top-1 accuracy.