Evaluating Toxicity Understanding of LLM Agents

Abstract

Research on toxicity in LLMs has largely focused on detection tasks, such as identifying hate speech or stereotyping in texts. Recently, these tasks have increasingly been embedded in agentic workflows, where LLMs autonomously query external APIs and reason over results before responding. This shift promotes the perception that LLMs exhibit an “understanding” of toxicity, yet how such understanding can be meaningfully interpreted by humans remains unclear. In this position paper, we first unpack this oversight by highlighting the fundamental gaps in current literature and then propose a framework for evaluating toxicity understanding of agentic LLMs. Overall, this short paper aims to shift the discourse from improving toxicity detection in LLMs to evaluating how LLMs understand toxicity in order to enhance their trustworthiness in downstream tasks.

Document Type

Conference Proceeding

Publication Date

1-23-2026

Publication Title

2025 IEEE International Conference on Collaborative Advances in Software and ComputiNg (CASCON)

Share

COinS