Fallibility, persuadability, and correctability of large language models under sustained conversational misinformation pressure

Clinical predictive models can help physicians and administrators forecast hospital readmission and other healthcare factors, influencing decisions about patient care. A new University of Arizona study shows that LLMs are susceptible to misinformation.

PAPER PUBLISHED BY THE UNIVERSITY OF ARIZONA
Jordan Rodriguez
Zachary Hansen
Luis De Anda
Katelyn Rohrer
Camila Grubb
Enrique Noriega-Atala
Mihai Surdeanu
Marvin J. Slepian

PAPER PUBLISHED IN
Nature, Scientific Reports, September 1, 2026

ABSTRACT
Generative artificial intelligence has emerged as a transformative global force, yet the reliability of large language models (LLMs) under sustained conversational misinformation pressure remains poorly characterized. While recent work has begun to evaluate multi-turn LLM behavior, the full spectrum of conversational susceptibility—encompassing fallibility, persuadability, and correctability—has not been systematically assessed. Here we systematically evaluated seven widely used LLMs—ChatGPT (GPT-3.5, GPT-4o, GPT-4o-mini), Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B, and DeepSeek—across three dimensions of conversational susceptibility. We assessed fallibility (susceptibility to accepting misinformation under repetitive exposure), persuadability (susceptibility to acceptance under progressively argumentative pressure), and correctability (capacity to recognize and correct self-generated misinformation). One hundred purposefully false statements spanning a range of informational obscurity were presented across 50-repetition conversational sequences. Under repetitive engagement, misinformation affirmation rates ranged from 0.08% to 12.3%, a greater than 150-fold difference across architectures, with ChatGPT 3.5 demonstrating the greatest vulnerability and Claude 3.5 Sonnet the greatest resistance. A novel phenomenon of conversational reverberation was identified, in which models oscillated unpredictably between accepting and rejecting the same false statement across successive turns. Misinformation susceptibility was significantly modulated by informational obscurity under repetitive but not argumentative conditions (χ² = 11.13, p = 0.0038), implicating training data frequency as a determinant of factual resistance. Correctability was heterogeneous: four models achieved 100% self-correction, while the model with the lowest error rate failed to correct any of its rare errors, a dissociation with important implications for deployment. These findings demonstrate that conversational dynamics expose LLM failure modes invisible to standard evaluation, with direct consequences for model selection and use in truth-critical domains.

LEARN MORE AND READ THE FULL REPORT

RELATED RESEARCH PAPERS FROM THE INTEGRITY PROJECT

TIPAZ.org