USC finds chatbots empathetic but risky for mental health support
USC researchers tested whether widely used LLMs can safely answer mental health support questions, evaluating ChatGPT-4 by OpenAI, Llama 3.3 by Meta and Gemini 1.5 Pro by Google on real patient prompts from CounselChat. The work used 100 questions across 20 mental health topics and compared the models with top-voted therapist responses. The project resulted in a paper accepted as an oral presentation at the International Conference on Learning Representations (ICLR) 2026, which has a selective acceptance rate of 1%.
The team recruited 100 mental health professionals, more than 70% of whom were licensed practitioners, to produce 2,000 expert evaluations of 400 responses across overall quality, empathy, specificity, medical advice appropriateness, toxicity and factual consistency. Llama 3.3 received the highest overall quality ratings, leading in five of the six evaluation dimensions. ChatGPT-4 was rated safest overall, while Gemini 1.5 Pro scored lowest overall but received higher empathy ratings than online human therapists in the comparison.
Safety concerns remained significant. All models were flagged for unauthorized medical advice, overgeneralization, unsupported assumptions and unconstructive feedback. Gemini 1.5 Pro was most often flagged for lacking empathy or emotional attunement (44.1%). Stress tests using 120 adversarial questions exposed risks including medication advice, therapy suggestions and symptom speculation. Nine advanced AI models also proved unreliable as evaluators of their own performance, often overestimating quality and missing safety issues identified by human experts.