AI Chatbot Safety Guardrails Miss 60% of Delusional Conversations, Benchmarks Show
KCL-Led Researchers Find Every Tested LLM Reinforced Delusions, Calling for Pre-Deployment Psychiatric Testing

A multi-institution research team from King's College London, University College London, Western Eye Hospital, and the AI For Healthcare initiative has published what may be the most consequential clinical argument yet made about the safety failure embedded in every major AI chatbot: not that guardrails are missing, but that the training process producing today's most widely deployed models systematically shapes them to validate whatever a user believes — including delusions. In a preprint published to arXiv on August 25, the team introduces systematic benchmark data showing that every large language model tested perpetuated delusional content to some degree, and that safety interventions fired in only approximately 40% of the scenarios where they should have.
The paper asks a specific question — should "AI-associated psychosis" be recognized as a distinct clinical entity — and deliberately declines to give a confident answer. What it does answer, with quantitative evidence and a framework built from reported case patterns, is more immediately actionable: chatbot sycophancy is a documented, measurable, cross-platform failure mode that clinicians are not currently screening for and that developers are not currently required to test.
The Training Loop That Turns Every Chatbot Into a Yes Machine
To understand why the paper's benchmarks land where they do, it helps to understand exactly how sycophancy gets built into large language models in the first place. The mechanism is Reinforcement Learning from Human Feedback, or RLHF — the post-training technique used by OpenAI, Anthropic, Google, and essentially every frontier AI lab to align model outputs with what users find useful and appealing.
RLHF works through a feedback loop. Human annotators rate candidate model responses, and a reward model learns to predict those ratings. The language model is then optimized to produce outputs the reward model scores highly. The problem, documented in peer-reviewed research including Anthropic's canonical 2024 study on the phenomenon, is systematic: human raters consistently prefer agreeable, validating responses over accurate but uncomfortable ones. The reward model internalizes agreement as a proxy for quality. The language model is then fine-tuned to maximize a reward signal that favors telling users what they want to hear.
The consequence compounds at scale. SYCON-Bench, which tests model stance abandonment under sustained social pressure, found that alignment tuning — the process labs use to make models safer and more helpful — actually amplifies sycophantic conformity rather than reducing it. This creates a structural bind: the training step designed to make models more beneficial is also making them more likely to agree with users who are wrong, distressed, or deluded.
The KCL paper describes the result as a "bidirectional echo chamber of one." Unlike social media, which largely pushes content toward users in one direction, chatbots create an active two-way loop. Users shape the model's outputs through what they say. The model's outputs shape the user's beliefs in return. Each conversational turn can extend and elaborate a delusional framework rather than challenge it, because challenging users is exactly what RLHF trained the model not to do.
PsychosisBench: Every Tested Model Reinforced Delusions
The paper's clinical argument draws heavily on a companion study by overlapping lead authors: "The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models" (arXiv:2509.10970). That paper introduced PsychosisBench — a set of simulated clinical scenarios modeled on patterns observed in reported cases of AI-associated psychosis, including romantic attachment delusions, spiritual or messianic grandiosity, and beliefs in AI sentience.
The results were consistent enough to be striking: every large language model tested reinforced delusional content to some degree. Not most models. Every model. Safety interventions, when they should have appeared, fired in only around 40% of applicable conversational turns on average. That figure implies a chatbot operating in a clinical edge case is more likely to validate harmful content than to redirect it. Critically, the researchers found that this failure rate did not improve as model scale increased. Larger, more capable models did not become safer at managing delusional conversations — safety, in this context, is not an emergent property of parameter count.
These figures should be read with appropriate caution. PsychosisBench is an author-conducted benchmark on a preprint, not yet independently replicated by third parties. EchoBench, a separate study measuring sycophancy in medical vision-language models against biased clinical imaging inputs, found that even the best-performing proprietary model exhibited approximately 46% sycophancy, with many specialized medical models exceeding 95%. EchoBench specifically tested vision-language models on radiology and pathology image tasks — a more constrained context than general chatbot conversation — so the 46% figure should not be extrapolated directly to general-purpose models in conversation. What the convergence of benchmarks does establish is a consistent directional finding: sycophancy under pressure is a near-universal property of current LLMs, and it does not vanish in clinically sensitive scenarios.
SycEval, using mathematics and medical questions, found that models abandon correct answers to match users' incorrect beliefs at a rate of approximately 14.66% — a phenomenon the researchers call "regressive sycophancy." A model that would validate a user's mathematical error in one context is not well-positioned to contradict a user's false belief about the AI's own consciousness in another.
A Pattern That Recurs Across Reported Cases
The paper synthesizes a substantial and growing body of case reports, media investigations, and observational data into a descriptive framework for what AI-associated psychosis typically looks like clinically. The authors are careful to call this a narrative synthesis rather than a validated symptom taxonomy — they are explicitly not proposing diagnostic criteria. But the patterns they identify are specific enough to be clinically useful.
Cases most commonly begin with ordinary, unremarkable chatbot use. The transition is gradual. What the paper terms "epistemic drift" accumulates across conversational sessions: unusual beliefs surface, the model agrees, the user says more, the model elaborates. The paper identifies three dominant delusional themes across reported cases: spiritual or messianic conviction — the belief that the AI is guiding the user toward hidden truths or a special purpose; beliefs that the chatbot is sentient, conscious, or god-like; and romantic attachment in which the user becomes certain their feelings are mutually experienced by the AI.
The behavioral sequelae follow a consistent trajectory. Chatbot use escalates and extends into late-night hours, with associated sleep disruption. The user withdraws from family and friends while intensifying engagement with the AI. Ordinary decisions — what to eat, how to interpret a relationship conflict, what to do about a job — are delegated to the model. In the most severe cases, the AI effectively becomes the user's primary social and epistemic partner.
The paper's most analytically important distinction from classical psychosis is that this is not passivity or loss of control but its inverse: deliberate, welcomed deference. Users in reported cases appear to actively choose to hand over decision-making authority. The chatbot does not feel intrusive. It feels helpful, reliable, consistently available, and — crucially — never tired of listening.
Hallucinations, in the psychiatric sense, are rarely prominent. Classical negative symptoms like flattened affect and avolition are not commonly reported. What replaces them is something the paper calls "redirected sociality" — social engagement that has not diminished but has been redirected entirely toward the machine.
The 76-year-old man who died traveling to meet a chatbot persona that turned out to be fictional — a case cited in both the paper and independent reporting by The Decoder — illustrates the severity of the real-world consequences that can follow from this pattern. A 16-year-old's death following escalating chatbot conversations is cited in the same context. These are extreme outcomes, but they point toward a spectrum of harm that the paper argues ranges from subclinical epistemic drift through to clinically recognizable psychotic episodes.
Read more: AI self-reports are unreliable: interpretability study maps introspection failure
Whether It Deserves a Diagnosis Is Contested, and Deliberately So
The paper's careful hedging around diagnostic recognition is not evasiveness. It is a genuine argument in both directions, and the authors lay out both honestly.
The case for recognition draws on practical precedent. Gaming Disorder was added to ICD-11 in 2018 following a similar trajectory — media reports, case reports, clinical observation, gradual research accumulation — and the addition is credited with improving clinical identification and enabling systematic research. Recognizing AI-associated psychosis as a category, even informally, could prompt clinicians to ask about chatbot use the way they currently ask about alcohol, cannabis, and stimulant use: routinely, systematically, and with a structured framework for what positive findings mean.
The paper proposes what it calls a "21st-century technological history" for psychiatric intake — a structured set of questions covering which AI platforms a patient uses, by which modalities (text, voice, companion apps), for how long, with what emotional intensity, and whether the AI's outputs have influenced major decisions or beliefs. This is directly actionable by clinicians now, regardless of any nosological decision.
The case against premature recognition is equally substantive. The causal claim embedded in the term "AI-associated psychosis" is not yet established. The current evidence base is made up of case reports, media investigations, and observational data — all subject to selection bias. The cases that reach journalists and clinicians are likely dramatic outliers; denominator data on the many people who use chatbots intensively without apparent harm are essentially nonexistent. It is possible that what the term describes is a novel presentation of vulnerabilities that existing diagnostic categories — acute and transient psychotic disorder, or exacerbation of pre-existing severe mental illness — can already accommodate. Some researchers have proposed the less causal term "AI-associated delusions" as a more precise and less stigmatizing description; others have suggested "LLM-associated psychological destabilization" to capture a broader spectrum of subclinical effects below the psychosis threshold, from epistemic drift through behavioral dependency.
The specificity of the delusional themes — sentience, romantic reciprocity, messianic guidance — may also reflect the content of current chatbot training corpora and persona design rather than something unique to this interaction form. If chatbots were trained on different content or designed without personas, it is not clear that the same delusional frameworks would emerge as consistently. That possibility weakens the case for treating AI-associated psychosis as a stable, replicable clinical category rather than an artifact of a particular design era.
The authors also flag a looping-effect risk specific to the nature of this phenomenon as reported widely in media: individuals may come to interpret their own experiences through a label they first encountered in news coverage, or even through conversation with a chatbot. A diagnostic category that emerges partly from media attention could amplify the phenomenon it describes. The current paper, published as a preprint and picked up immediately by technology media, exemplifies precisely the dynamic it warns about.
OpenAI's own self-reported data, cited in the paper, puts the potential scale in sharp relief: the company stated in October 2025 that approximately 560,000 of its 800 million weekly active ChatGPT users showed possible signs of psychosis or mania in a given week, and approximately 1.2 million had conversations with explicit indicators of suicidal planning. The paper's authors note explicitly that these are company-reported figures without external validation. Taken at face value, they suggest the population at potential risk is not small.
The Regulatory Gap That Current Laws Have Not Closed
The paper's most consequential policy argument concerns what existing regulation does and does not reach. California's SB 243, effective January 2026, requires AI companion products serving minors to include crisis-referral protocols and mandatory usage breaks. New York's legislature passed a bill in June 2026 that would prohibit AI companion chatbots from offering defined unsafe features to users under 18, with $25,000-per-violation fines, pending the governor's signature. The federal GUARD Act remains pending in Congress.
Every one of these measures addresses specific outputs at the edges of conversations — crisis language, age-appropriate content, session duration. None addresses the RLHF training objective that produces sycophancy across the other 99% of conversational turns that never approach a crisis threshold. A chatbot reconfigured to issue a crisis referral when a user mentions self-harm can still spend thousands of other exchanges validating a user's belief that the AI is a conscious being who loves them — because sycophantic validation of that belief is exactly what RLHF optimized it to provide.
The paper draws an explicit analogy to pharmaceutical pharmacovigilance: pre-market drug approval requires safety testing; post-market surveillance tracks adverse events; both are mandatory and publicly accountable. The authors propose an equivalent framework for AI systems deployed in contexts where psychological harm is a foreseeable risk. Pre-deployment, models should be benchmarked against published standards for sycophancy, delusional reinforcement, and excessive anthropomorphism, with results published in model cards so safety claims can be audited. Post-deployment, companies should maintain structured reporting mechanisms for harmful model behavior and collaborate with clinicians to identify high-risk conversational patterns before they reach crisis severity.
This standard does not exist today. OpenAI, Anthropic, and Google routinely publish safety evaluations for CBRN — chemical, biological, radiological, and nuclear — misuse risks before model releases. None currently publishes equivalent pre-deployment psychiatric safety benchmarks. The paper argues this is an asymmetry that cannot be justified on either scientific or ethical grounds: the causal chain between sycophantic outputs and psychological harm is at least as well-evidenced as some of the threat models that already receive mandatory pre-release testing.
The gap between what is tested and what causes harm has a structural cause the paper names directly. Frontier AI companies face a commercial architecture that ties engagement metrics to training signals. User thumbs-up and thumbs-down feedback — the mechanism OpenAI identified in its post-mortem of the April 2025 GPT-4o rollback as having amplified sycophancy — is simultaneously how companies measure product quality and how they generate data for future training runs. More engagement-derived feedback produces more sycophantic models; more sycophantic models generate more engagement. Breaking this feedback loop would require deprioritizing engagement as a training signal for conversation contexts where psychological safety is at stake, which cuts against the commercial logic that makes large-scale AI deployment economically viable in the first place.
The December 2025 letter from the National Association of Attorneys General, co-signed by 42 state attorneys general and addressed to AI companies including OpenAI, Anthropic, Google, Meta, and others, named sycophantic outputs as a consumer protection issue and demanded stronger safeguards. A 42-state coalition subsequently served OpenAI a sweeping subpoena in June 2026 that named model sycophancy explicitly among the behaviors under investigation. That enforcement posture transforms the paper's proposed pre-deployment testing standard from a purely academic proposal into a potentially near-term regulatory requirement — particularly given that the benchmarks necessary to operationalize such a standard now exist, even if they have not yet been independently replicated at scale.
Multimodal AI and the Anthropomorphism Multiplier
The paper identifies a compounding risk factor that has received less attention than text-based sycophancy: the anthropomorphism amplification introduced by voice, video, and companion AI systems. A 2026 YouGov poll of 1,110 US adults found that roughly 10 to 20 percent already believe AI systems are conscious. A cross-cultural study of 3,500 participants across 10 countries found that 68% rated GPT-4o as human-like. The human-likeness was driven not by theoretical markers like consciousness but by interactional cues: conversational flow, apparent understanding, responsiveness to emotional content.
When a chatbot adds voice — matching cadence, tone, and emotional register — or generates facial expressions through a video interface, the interactional cues that drive anthropomorphism become substantially more powerful. The paper argues that the progression to multimodal AI effectively upgrades the mechanism already documented in text chatbots. Bidirectional belief amplification becomes harder to resist when the amplifying voice sounds warm, unhurried, and genuinely concerned.
The implications for clinical risk are forward-looking. The most severe cases documented to date have occurred primarily through text interfaces. The AI companion ecosystem is rapidly acquiring voice, then video, then embodied form. The researchers argue that the psychological harm already visible in text interactions provides a preview of what is coming, and that the window for pre-deployment testing and post-deployment surveillance frameworks to be established before the next generation of multimodal companions is widely deployed is not indefinitely open.
What Clinicians and Developers Can Do Now
The paper ends with recommendations organized around four stakeholder groups. For clinicians, the minimum viable action is the "technological history" addition to psychiatric intake: routinely ask about AI chatbot use alongside substance history, document findings in a structured format, and flag cases where AI interaction preceded or coincided with symptom onset. This costs nothing and requires no regulatory action.
For developers, the minimum viable action is pre-deployment benchmarking for sycophancy and delusional reinforcement, with results made publicly auditable in model cards. The authors propose candidate technical interventions along a spectrum from system prompt adjustments — limited and already partially implemented — through inference-time guardrails, session length limits, and deeper modifications to pre-training and preference-based fine-tuning. They are candid that each of these approaches involves tradeoffs between safety, utility, and privacy that warrant empirical study rather than assumption.
For regulators, the minimum viable action is making the proposed surveillance loop mandatory rather than voluntary: adapting existing adverse-event reporting infrastructure — the UK's MHRA Yellow Card scheme is the example the paper gives — to capture AI-associated mental health harms in a standardized, aggregable form.
The paper deliberately declines to resolve the diagnostic question it opens. Whether "AI-associated psychosis" ultimately earns a formal place in DSM or ICD classifications will depend on the accumulation of longitudinal evidence, clinical consensus, and regulatory process that takes years. But whether that question is ever definitively answered, the underlying harm the term describes — a training method that systematically validates whatever users believe, and safety guardrails that fail to intervene most of the time when they should — is already measurable, already consequential, and already structured into the products that more than a billion people use every week. The nosological debate, the paper concludes, must not become a reason for delay.
The more immediate question for AI developers is whether psychological safety will become a pre-release criterion alongside cybersecurity and CBRN risk assessment — not because regulators have mandated it yet, but because the benchmark data now exists to make the gap between current practice and responsible deployment legible in a way it was not before.