By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
AIModelKitAIModelKitAIModelKit
  • Home
  • News
    NewsShow More
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    4 Min Read
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    4 Min Read
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    5 Min Read
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    5 Min Read
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    4 Min Read
  • Open-Source Models
    Open-Source ModelsShow More
    Exploring How Mobility Enhances Language Models’ Understanding of Location
    Exploring How Mobility Enhances Language Models’ Understanding of Location
    5 Min Read
    Optimize Candidate Biomarkers with Our AI Tool for Wearable Sensor Data Analysis
    Optimize Candidate Biomarkers with Our AI Tool for Wearable Sensor Data Analysis
    4 Min Read
    Beyond BMI: Assessing Cardiometabolic Risk Using Smartphone Images
    Beyond BMI: Assessing Cardiometabolic Risk Using Smartphone Images
    5 Min Read
    Overcoming Recall Challenges: The Impact of Empty Shelves and Lost Keys on Parametric Factuality
    Overcoming Recall Challenges: The Impact of Empty Shelves and Lost Keys on Parametric Factuality
    6 Min Read
    Enhancing AMIE for Expert-Level Audio-Visual Clinical Consultations
    Enhancing AMIE for Expert-Level Audio-Visual Clinical Consultations
    5 Min Read
  • Guides
    GuidesShow More
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    6 Min Read
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    4 Min Read
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    6 Min Read
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    6 Min Read
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    6 Min Read
  • Tools
    ToolsShow More
    Optimizing LFM2.5 Q4_0 Checkpoints through Quantization-Aware Distillation Techniques
    Optimizing LFM2.5 Q4_0 Checkpoints through Quantization-Aware Distillation Techniques
    4 Min Read
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    6 Min Read
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    5 Min Read
    July 2026 Security Incident Disclosure: Key Insights and Updates
    July 2026 Security Incident Disclosure: Key Insights and Updates
    6 Min Read
    Boosting Performance with Native-Speed vLLM Transformers for Enhanced Modeling Backend
    Boosting Performance with Native-Speed vLLM Transformers for Enhanced Modeling Backend
    5 Min Read
  • Events
    EventsShow More
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    4 Min Read
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    5 Min Read
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    5 Min Read
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    5 Min Read
    Unlocking the Power of Open Models at Nemotron Labs: Discover the Advantage
    Unlocking the Power of Open Models at Nemotron Labs: Discover the Advantage
    7 Min Read
  • Ethics
    EthicsShow More
    Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
    Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
    6 Min Read
    Exploring Space Threats from Mirrors and Recognizing AI Drug Innovations: The Download
    Exploring Space Threats from Mirrors and Recognizing AI Drug Innovations: The Download
    5 Min Read
    Understanding AI Bias: How Human Decisions Shape Algorithmic Errors
    Understanding AI Bias: How Human Decisions Shape Algorithmic Errors
    5 Min Read
    How This Company’s Space Mirror Plans Could Threaten the Night Sky for Everyone
    How This Company’s Space Mirror Plans Could Threaten the Night Sky for Everyone
    5 Min Read
    Understanding Orphan Risks in Artificial Intelligence: Insights from Diverging Safety and Compliance Frameworks on AI Companies’ Risk Prioritization
    Understanding Orphan Risks in Artificial Intelligence: Insights from Diverging Safety and Compliance Frameworks on AI Companies’ Risk Prioritization
    5 Min Read
  • Comparisons
    ComparisonsShow More
    Exploring DuckDB v2.0: Transforming Architecture for Enhanced Distributed Network Capabilities
    Exploring DuckDB v2.0: Transforming Architecture for Enhanced Distributed Network Capabilities
    6 Min Read
    Unlocking Self-Knowledge: SKILL-RAG for Enhanced Learning and Filtering in Retrieval-Augmented Generation
    Unlocking Self-Knowledge: SKILL-RAG for Enhanced Learning and Filtering in Retrieval-Augmented Generation
    4 Min Read
    Understanding Decentralization: An Ontological Exploration and Definition
    Understanding Decentralization: An Ontological Exploration and Definition
    5 Min Read
    Microsoft Transitions AI Governance from Policy Frameworks to Real-time Enforcement
    Microsoft Transitions AI Governance from Policy Frameworks to Real-time Enforcement
    6 Min Read
    Optimizing Multi-Turn Reasoning in LLM Agents with Fine-Grained Reward Structures and Effective Credit Assignment Strategies
    Optimizing Multi-Turn Reasoning in LLM Agents with Fine-Grained Reward Structures and Effective Credit Assignment Strategies
    6 Min Read
Search
  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
Reading: Recognizing Toxicity: Understanding Span and Target in Chemical Safety
Share
Notification Show More
Font ResizerAa
AIModelKitAIModelKit
Font ResizerAa
  • 🏠
  • 🚀
  • 📰
  • 💡
  • 📚
  • ⭐
Search
  • Home
  • News
  • Models
  • Guides
  • Tools
  • Ethics
  • Events
  • Comparisons
Follow US
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
AIModelKit > Comparisons > Recognizing Toxicity: Understanding Span and Target in Chemical Safety
Comparisons

Recognizing Toxicity: Understanding Span and Target in Chemical Safety

aimodelkit
Last updated: January 7, 2026 8:15 pm
aimodelkit
Share
Recognizing Toxicity: Understanding Span and Target in Chemical Safety
SHARE
[Submitted on 2 Jun 2025 (v1), last revised 5 Jan 2026 (this version, v2)]

View a PDF of the paper titled Something Just Like TRuST: Toxicity Recognition of Span and Target, by Berk Atil and two other authors.

View PDF | HTML (experimental)

Abstract: Toxic language includes content that is offensive, abusive, or that promotes harm. Progress in preventing toxic output from large language models (LLMs) is hampered by inconsistent definitions of toxicity. We introduce TRuST, a large-scale dataset that unifies and expands prior resources through a carefully synthesized definition of toxicity and corresponding annotation scheme. It consists of ~300k annotations, with high-quality human annotation on ~11k. To ensure high quality, we designed a rigorous, multi-stage human annotation process and evaluated the diversity of the annotators. Then we benchmarked state-of-the-art LLMs and pre-trained models on three tasks: toxicity detection, identification of the target group, and of toxic words. Our results indicate that fine-tuned PLMs outperform LLMs on the three tasks, and that current reasoning models do not reliably improve performance. TRuST constitutes one of the most comprehensive resources for evaluating and mitigating LLM toxicity and other research in socially-aware and safer language technologies.

Submission History

From: Berk Atil [view email]

[v1] Mon, 2 Jun 2025 23:48:16 UTC (1,094 KB)
[v2] Mon, 5 Jan 2026 21:38:57 UTC (1,098 KB)

—

### Understanding Toxicity in Language Models

Toxic language can take many forms, often presenting as offensive, abusive, or harmful content. In today’s digital landscape, large language models (LLMs) are becoming increasingly integrated into everyday applications, from customer service to content creation. However, the challenge of ensuring these models do not propagate toxic language remains a pressing issue. This calls for a comprehensive understanding of toxicity and how to effectively manage it within LLMs.

### The TRuST Dataset: A New Benchmark

The introduction of the TRuST dataset marks a significant advancement in the field of toxicity recognition. This large-scale dataset consolidates and expands upon previous resources, creating a comprehensive framework for defining toxicity. With approximately 300,000 annotations, of which around 11,000 have undergone high-quality human annotation, TRuST aims to provide a robust foundation for future research. The dataset is instrumental for developers and researchers seeking to refine the capabilities of LLMs by providing a clear understanding of what constitutes toxic language.

More Read

Using Sentence Space Embedding for Enhanced Classification of Fake News Data Streams
Using Sentence Space Embedding for Enhanced Classification of Fake News Data Streams
Enhancing Mathematical Reasoning in Smaller Models Through Arithmetic Learning Integration: A Study
Enhancing Large-Scale Mixture of Experts Training with Piper: Resource Modeling and Pipelined Hybrid Parallelism Solutions
Empower Your Creativity: Agentic Crafting in Rock and Roll and the ROME Model in an Open Agentic Learning Ecosystem
Comparing Generation vs. QA-Based Evaluations: Which Method Reigns Supreme?

### Multi-Stage Annotation Process

The quality of any dataset is paramount, and TRuST has been developed through a meticulous multi-stage human annotation process. This was designed to not only ensure the accuracy of the toxicity classifications but also to evaluate the diversity of the annotators. Diversity among annotators helps to reduce bias in the dataset, ensuring it reflects a variety of perspectives and cultural contexts. This thoughtful approach to annotation signifies a leap forward in addressing the complexities surrounding toxic language.

### Benchmarking LLMs and Pre-trained Models

The efficacy of any dataset can be evaluated through benchmarking against existing models. In the case of TRuST, state-of-the-art LLMs and pre-trained models were assessed on three primary tasks: toxicity detection, identification of the target group, and pinpointing toxic words. The findings revealed that fine-tuned pre-trained language models (PLMs) significantly outperform LLMs in these tasks. This insight is crucial for developers aiming to build safer language technologies, as it informs the selection of models based on specific functionalities.

### Addressing Current Limitations

One of the key discoveries in the TRuST study is that the current reasoning models do not consistently enhance performance in toxicity detection. This indicates a critical area for future research, highlighting the need for ongoing development to improve the capabilities of these models. It brings to light the importance of understanding the limitations of current technologies and the necessity for continual innovation in the field of natural language processing.

### Implications for Safer Language Technologies

TRuST stands out as one of the most comprehensive resources available for evaluating and mitigating toxicity in large language models. It paves the way for further research into socially-aware language technologies, enabling developers and researchers to create applications that are not only effective but also responsible. By utilizing the TRuST dataset, stakeholders can contribute to building a digital landscape that prioritizes safety, inclusivity, and respect, making strides towards a more harmonious online environment.

### Future Directions in Toxicity Research

As the discourse surrounding toxic language continues to evolve, the need for innovative solutions remains significant. Future research should focus on enhancing the methodologies used to identify and mitigate toxicity, as well as expanding datasets like TRuST to encapsulate a broader spectrum of language nuances. By fostering collaboration within the research community, we can collectively work towards developing more sophisticated tools that effectively curb the spread of toxic language in digital communication.

—

This article serves as a detailed exploration of the TRuST dataset and its implications in the fight against toxic language in AI. By focusing on critical aspects such as annotation quality, benchmarking, and future directions for research, it provides a well-rounded understanding of the ongoing challenges and advancements in this essential area of study.

Inspired by: Source

Exploring the Architectures Driving Modern AI Systems: Insights from QCon San Francisco 2025
Enhanced Context-Aware Dense Retrieval Techniques for Better Semantic Associations and Comprehensive Long Story Understanding
Optimizing Map Question Answering with Multimodal Large Language Models: An Evaluation Study
Splits! A Comprehensive Dataset and Evaluation Framework for Sociocultural Linguistic Research
Enhancing LLM Evaluation with Adaptive Testing: A Superior Psychometric Approach to Static Benchmarks

Sign Up For Daily Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

By signing up, you agree to our Terms of Use and acknowledge the data practices in our Privacy Policy. You may unsubscribe at any time.
Share This Article
Facebook Copy Link Print
Previous Article Ultimate Guide to Converting Bytes to Strings in Python: Take the Quiz – Real Python Ultimate Guide to Converting Bytes to Strings in Python: Take the Quiz – Real Python
Next Article Dell Acknowledges Consumer Disinterest in AI-Driven PCs Dell Acknowledges Consumer Disinterest in AI-Driven PCs

Stay Connected

XFollow
PinterestPin
TelegramFollow
LinkedInFollow

							banner							
							banner
Explore Top AI Tools Instantly
Discover, compare, and choose the best AI tools in one place. Easy search, real-time updates, and expert-picked solutions.
Browse AI Tools

Latest News

Exploring DuckDB v2.0: Transforming Architecture for Enhanced Distributed Network Capabilities
Exploring DuckDB v2.0: Transforming Architecture for Enhanced Distributed Network Capabilities
Comparisons
Unlocking Self-Knowledge: SKILL-RAG for Enhanced Learning and Filtering in Retrieval-Augmented Generation
Unlocking Self-Knowledge: SKILL-RAG for Enhanced Learning and Filtering in Retrieval-Augmented Generation
Comparisons
Understanding Decentralization: An Ontological Exploration and Definition
Understanding Decentralization: An Ontological Exploration and Definition
Comparisons
Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
Ethics
//

Leading global tech insights for 20M+ innovators

Quick Link

  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events

Support

  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us

Sign Up for Our Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

AIModelKitAIModelKit
Follow US
© 2025 AI Model Kit. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?