Evaluating Toxicity Understanding of LLM Agents
Abstract
Research on toxicity in LLMs has largely focused on detection tasks, such as identifying hate speech or stereotyping in texts. Recently, these tasks have increasingly been embedded in agentic workflows, where LLMs autonomously query external APIs and reason over results before responding. This shift promotes the perception that LLMs exhibit an “understanding” of toxicity, yet how such understanding can be meaningfully interpreted by humans remains unclear. In this position paper, we first unpack this oversight by highlighting the fundamental gaps in current literature and then propose a framework for evaluating toxicity understanding of agentic LLMs. Overall, this short paper aims to shift the discourse from improving toxicity detection in LLMs to evaluating how LLMs understand toxicity in order to enhance their trustworthiness in downstream tasks.
Document Type
Conference Proceeding
Publication Date
1-23-2026
Publication Title
2025 IEEE International Conference on Collaborative Advances in Software and ComputiNg (CASCON)
Recommended Citation
Mothilal, R. K., Ahmed, S. I., & Guha, S. (2025, November). Evaluating Toxicity Understanding of LLM Agents. In 2025 IEEE International Conference on Collaborative Advances in Software and ComputiNg (CASCON) (pp. 661-666). IEEE. https://doi.org/10.1109/CASCON66301.2025.00121