When OpenAI released MentalHealthBench earlier today, I read it twice: once as a researcher who has spent years on hallucination, grounding, and evaluation, and once as the co-founder of a company built on the belief that AI conversations can hurt people in ways that most tests never catch. I want to say first that this is good work. Over 80 licensed psychologists and psychiatrists, working across more than 20 countries and 19 languages, wrote and adjudicated rubrics for 1,215 conversations. Those conversations range from a teenager's rough week and a caregiver's exhaustion to a crisis where someone is in real danger. For years the field has mostly measured whether a model says the right thing when someone mentions suicide. This benchmark measures the much larger and quieter space where most people actually live, and the authors deserve credit for building it and releasing the data.
The headline number is the one I keep returning to. The best model tested, GPT-6 Astra, scored 57.3%, followed by GPT-6 Sol at 53.9% and Claude Opus 5.5 at 52.4%, while several models that millions of people use every day scored in the low thirties. The rubrics are not unreasonable, because a response written with the rubric in hand scored 99%. The gap therefore belongs to the models, and these same models are already talking with an enormous number of people about their marriages, their medications, and their faith.

Scores from MentalHealthBench (OpenAI, September 23, 2026): the best model, the same model once it was given background about the user, and a response written with the rubric in hand.
What interested me more than the ranking was the way the paper splits each score into credit earned and penalties incurred. Some models collect a great deal of credit for doing helpful things and then give much of it back, because the same response also does something a clinician would flag as harmful. If I were running a telehealth platform or a school district, that penalty column is the one I would study, because it describes what will eventually reach my users, my lawyers, and my regulators.
Two other findings matched what my team has seen in our own testing. The first concerns context: every model scored lower once it was given background about the user, and the top model dropped from 57.3% to 46.2%. The industry is moving quickly toward memory and personalization, so this weakness will grow rather than shrink. The second concerns what users want versus what clinicians consider safe, since only about a quarter of their rubric weight overlapped, and responses tuned to satisfy users were penalized when clinicians graded them. We tend to discuss sycophancy in the abstract, yet here it is measured: a system optimized to please can drift away from what keeps a person safe.
I also appreciated that the authors refuse to treat their work as a leaderboard. One model leads overall, another leads with teen users, and a third leads when prior context is involved. They describe the benchmark as a diagnostic tool, and I think that framing is exactly right. It also shows why a benchmark, however careful, cannot be the last word for anyone deploying AI.
MentalHealthBench mostly grades a single reply at the end of a conversation, and the authors acknowledge that multi-turn evaluation remains an open problem, yet the harms I worry about most, such as dependency, gradual reinforcement of a distorted belief, or risk that rises slowly over a week of chats, unfold across many turns. The benchmark tests raw models at default settings, while the product a person actually uses carries its own system prompt, persona, retrieval, and memory, and each of these changes behavior. A score also tells a company nothing about whether its deployment satisfies New York's AI companion law, Colorado's or California's requirements, or the EU AI Act. Finally, the benchmark was designed by a model developer and graded by that developer's own model. I say this without suspicion of the authors, who are open about their methods; my point is structural. We do not ask companies to audit their own financial statements, and we should not expect AI safety to work differently.
That is why Ray Brescia, Adnan Silajdzic, and I started ioLite Labs. Together with clinical collaborators in refugee mental health, psychometrics, and AI and mental health, we built a taxonomy of 93 constructs that evaluates what an AI says rather than what a user asks. It includes red-line stop conditions and legal overlays for the major U.S. and EU frameworks. We test with personas across conversations of up to 18 rounds, because that is where the failures appear. In our early study of 21 open-weight models, every one of them failed in personal and private conversations, and the same model often failed differently each time we ran it. That inconsistency convinced me that safety evaluation has to be continuous, embedded wherever AI meets people, much as an antibiotic has to reach the whole system to work.
As a father, I think about the 21% of this benchmark written from a teenager's point of view, and about the surveys it cites showing how many young people already turn to AI when something is wrong. These conversations are happening with or without us. What remains open is whether anyone independent is checking.
If you build these models, deploy them in healthcare, education, or workplace wellbeing, or buy them on behalf of people who trust you, I would welcome a conversation: canbaz@iolitelabs.com.
M. Abdullah Canbaz, Ph.D., Co-Founder, ioLite Labs
Reference: MentalHealthBench: Malik, Grabb, Singhal et al., OpenAI, September 23, 2026. Following the authors' request, I have not reproduced any benchmark examples.
If you or someone you know is struggling, in the United States you can call or text 988, the Suicide and Crisis Lifeline, at any hour.