By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
AIModelKitAIModelKitAIModelKit
  • Home
  • News
    NewsShow More
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    4 Min Read
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    4 Min Read
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    5 Min Read
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    5 Min Read
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    4 Min Read
  • Open-Source Models
    Open-Source ModelsShow More
    AgentHands: Creating Interactive Hand Gestures for Enhanced Conversations with Spatially Grounded Agents in XR
    AgentHands: Creating Interactive Hand Gestures for Enhanced Conversations with Spatially Grounded Agents in XR
    5 Min Read
    Exploring How Mobility Enhances Language Models’ Understanding of Location
    Exploring How Mobility Enhances Language Models’ Understanding of Location
    5 Min Read
    Optimize Candidate Biomarkers with Our AI Tool for Wearable Sensor Data Analysis
    Optimize Candidate Biomarkers with Our AI Tool for Wearable Sensor Data Analysis
    4 Min Read
    Beyond BMI: Assessing Cardiometabolic Risk Using Smartphone Images
    Beyond BMI: Assessing Cardiometabolic Risk Using Smartphone Images
    5 Min Read
    Overcoming Recall Challenges: The Impact of Empty Shelves and Lost Keys on Parametric Factuality
    Overcoming Recall Challenges: The Impact of Empty Shelves and Lost Keys on Parametric Factuality
    6 Min Read
  • Guides
    GuidesShow More
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    6 Min Read
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    4 Min Read
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    6 Min Read
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    6 Min Read
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    6 Min Read
  • Tools
    ToolsShow More
    Unlock Agentic Coding: Experimenting with Qwen 3.8 Flash-Next 176B Model on NVIDIA GB300 NVL72
    Unlock Agentic Coding: Experimenting with Qwen 3.8 Flash-Next 176B Model on NVIDIA GB300 NVL72
    5 Min Read
    Optimizing LFM2.5 Q4_0 Checkpoints through Quantization-Aware Distillation Techniques
    Optimizing LFM2.5 Q4_0 Checkpoints through Quantization-Aware Distillation Techniques
    4 Min Read
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    6 Min Read
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    5 Min Read
    July 2026 Security Incident Disclosure: Key Insights and Updates
    July 2026 Security Incident Disclosure: Key Insights and Updates
    6 Min Read
  • Events
    EventsShow More
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    4 Min Read
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    5 Min Read
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    5 Min Read
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    5 Min Read
    Unlocking the Power of Open Models at Nemotron Labs: Discover the Advantage
    Unlocking the Power of Open Models at Nemotron Labs: Discover the Advantage
    7 Min Read
  • Ethics
    EthicsShow More
    Taiwan Prosecutes Nine Individuals for Smuggling Advanced AI Servers to China: A Tech Industry Update
    Taiwan Prosecutes Nine Individuals for Smuggling Advanced AI Servers to China: A Tech Industry Update
    4 Min Read
    Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
    Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
    6 Min Read
    Exploring Space Threats from Mirrors and Recognizing AI Drug Innovations: The Download
    Exploring Space Threats from Mirrors and Recognizing AI Drug Innovations: The Download
    5 Min Read
    Understanding AI Bias: How Human Decisions Shape Algorithmic Errors
    Understanding AI Bias: How Human Decisions Shape Algorithmic Errors
    5 Min Read
    How This Company’s Space Mirror Plans Could Threaten the Night Sky for Everyone
    How This Company’s Space Mirror Plans Could Threaten the Night Sky for Everyone
    5 Min Read
  • Comparisons
    ComparisonsShow More
    Optimizing Nonconvex-Nonconcave Min-Max Problems with a Limited Maximization Domain: Insights from [2110.03950]
    Optimizing Nonconvex-Nonconcave Min-Max Problems with a Limited Maximization Domain: Insights from [2110.03950]
    5 Min Read
    Optimizing Social Media Safety: Scalable Few-Shot Harmful Content Moderation with Large Language Models
    Optimizing Social Media Safety: Scalable Few-Shot Harmful Content Moderation with Large Language Models
    5 Min Read
    Exploring Infinite-Dimensional Generative Diffusions through Doob’s h-Transform: A 2602.06621 Study
    Exploring Infinite-Dimensional Generative Diffusions through Doob’s h-Transform: A 2602.06621 Study
    4 Min Read
    DynHD: Detecting Hallucinations in Diffusion Large Language Models through Denoising Dynamics Deviation Learning
    DynHD: Detecting Hallucinations in Diffusion Large Language Models through Denoising Dynamics Deviation Learning
    5 Min Read
    Enhancing Web Content with GEO-Flag: Detecting and Measuring GEO-Optimized Content for Improved SEO
    Enhancing Web Content with GEO-Flag: Detecting and Measuring GEO-Optimized Content for Improved SEO
    4 Min Read
Search
  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
Reading: SmellBench: Assessing LLM Agents for Repairing Architectural Code Smells
Share
Notification Show More
Font ResizerAa
AIModelKitAIModelKit
Font ResizerAa
  • 🏠
  • 🚀
  • 📰
  • 💡
  • 📚
  • ⭐
Search
  • Home
  • News
  • Models
  • Guides
  • Tools
  • Ethics
  • Events
  • Comparisons
Follow US
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
AIModelKit > Comparisons > SmellBench: Assessing LLM Agents for Repairing Architectural Code Smells
Comparisons

SmellBench: Assessing LLM Agents for Repairing Architectural Code Smells

aimodelkit
Last updated: May 14, 2026 2:00 pm
aimodelkit
Share
SmellBench: Assessing LLM Agents for Repairing Architectural Code Smells
SHARE

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair

In a rapidly evolving software development landscape, maintaining code quality is more important than ever. One of the prevailing challenges facing developers is the presence of architectural code smells—issues in the code design that can erode maintainability and are often costly to repair manually. While localized bugs are typically easier to fix due to their more straightforward nature, architectural code smells necessitate a higher level of cross-module reasoning about design intent. This complexity complicates the repair process, making automated tools less effective.

Contents
  • Understanding Architectural Code Smells
  • The Role of LLM Agents
    • Task Orchestration Framework
    • Evaluation Methodology
  • Empirical Findings
  • Implications for Automated Software Engineering

In this article, we explore SmellBench, an innovative framework designed by Ion George Dinu and his collaborators. This framework serves to evaluate the ability of large language model (LLM) agents to repair architectural code smells effectively.

Understanding Architectural Code Smells

Architectural code smells suggest deficiencies in code structure and design that can hinder long-term maintainability. Unlike simple bugs, these issues require an understanding of inter-module relationships and overall design principles. This need for broader architectural insights presents a significant challenge for both developers and automated tools. Some common types of architectural code smells include:

  • God Objects: Classes that control too much behavior, leading to high coupling and low cohesion.
  • Spaghetti Code: Code that is tangled and difficult to follow, making it hard to manage and maintain.
  • Feature Envy: Situations where one class is overly interested in another’s data or functionality, indicating a potential design flaw.

Addressing these smells is critical for creating maintainable and scalable software systems, but the challenge lies in their complexity.

The Role of LLM Agents

Large language model agents have demonstrated remarkable capabilities in code-level tasks, particularly in bug fixing and localized refactoring. However, their potential for repairing architectural code smells remains an underexplored area. SmellBench sets out to fill this gap by providing a structured evaluation of various agent configurations from four prominent model families: GPT, Claude, Gemini, and Mistral.

More Read

Universal Multi-Agent Framework for Time-Persistent Cipher-Based Jailbreak Attacks on Language Models
Universal Multi-Agent Framework for Time-Persistent Cipher-Based Jailbreak Attacks on Language Models
Enhancing Robustness in Vision-Language Models with Partially Recentralization Softmax Loss
Overcoming Limitations of Discrete Neuronal Attribution in Neuroscience
Scalable Rapid Attention Distillation for Enhanced Linear Attention Decoders
Stripe Engineers Unleash Minions: How Autonomous Agents Generate Thousands of Weekly Pull Requests

Task Orchestration Framework

At the heart of SmellBench is its task orchestration framework. This framework incorporates smell-type-specific optimized prompts, which help guide the LLM agents in their attempts to repair detected smells. Additionally, the framework supports iterative multi-step execution, allowing agents to refine their approaches based on outcomes.

Evaluation Methodology

The evaluation methodologies employed by SmellBench are comprehensive. They include a scoring system that measures:

  • Repair Effectiveness: How well the agents manage to fix the identified architectural smells.
  • False Positive Identification: The ability of agents to discern between actual smells and those erroneously flagged.
  • Net Codebase Impact: The broader effects of the repairs on the overall codebase quality.

By using these criteria, SmellBench can paint a more nuanced picture of LLM agent performance in relation to architectural code smell repair.

Empirical Findings

The empirical evaluation conducted on 11 agent configurations revealed some enlightening insights into the current capabilities of LLM agents. The study focused on 65 hard-severity architectural smells detected by PyExamine in the widely used Python project, scikit-learn, and compared the results with expert judgments for validation.

Notably, the expert validation process indicated that a staggering 63.1% of detected smells were false positives. Despite this high false-positive rate, the best-performing LLM agent achieved a commendable 47.7% resolution rate for actual architectural code smells. This shows that while LLMs are making strides, there remains a critical need for development in their architectural understanding.

Moreover, an intriguing relationship was uncovered between repair aggressiveness and net codebase quality. While some agents exhibited high repair rates, they inadvertently introduced up to 140 new smells—a clear indicator that aggressive repairs do not always lead to improved quality.

Implications for Automated Software Engineering

The findings from SmellBench underscore a significant gap between the current capabilities of LLMs in performing localized code transformations and the architectural awareness essential for effective cross-module refactoring. As developers increasingly rely on automated tools to maintain code, understanding these limitations becomes crucial for informed decision-making.

Beyond individual agent performance, SmellBench is positioned to serve as a reusable infrastructure that tracks progress in this critical yet underexplored domain of automated software engineering. By focusing on architectural code smells, it opens avenues for further research and development aimed at enhancing LLM capabilities.

This framework not only aims to improve LLM behavior but also enriches the discussions around automated software engineering practices, helping to shape the future of code maintenance and quality assurance in tech development.

To explore the comprehensive findings and methodologies, researchers and developers can access the paper, “SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair,” available in PDF format from the authors.

Inspired by: Source

Microsoft Expands Azure AI Foundry Agent Service with Advanced Research Features
Exploring the Geometry of Sentiment: Are Sentiment Vectors Shaped Like Bananas?
MOIS-SAM2: A Cutting-Edge Exemplar-Based Model for Interactive Multilesion Segmentation of Neurobromas in Whole-Body MRI
Optimizing Structural Pruning with Connectivity-Based Regularization Techniques
Enhancing Bifidelity Parameter Estimation with Conditional Diffusion Models: A Comprehensive Study

Sign Up For Daily Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

By signing up, you agree to our Terms of Use and acknowledge the data practices in our Privacy Policy. You may unsubscribe at any time.
Share This Article
Facebook Copy Link Print
Previous Article NVIDIA and Ineffable Intelligence Join Forces to Revolutionize Reinforcement Learning Infrastructure NVIDIA and Ineffable Intelligence Join Forces to Revolutionize Reinforcement Learning Infrastructure
Next Article Humanoid Robots: The Future of Physical AI in Manufacturing Facilities Humanoid Robots: The Future of Physical AI in Manufacturing Facilities

Stay Connected

XFollow
PinterestPin
TelegramFollow
LinkedInFollow

							banner							
							banner
Explore Top AI Tools Instantly
Discover, compare, and choose the best AI tools in one place. Easy search, real-time updates, and expert-picked solutions.
Browse AI Tools

Latest News

Optimizing Nonconvex-Nonconcave Min-Max Problems with a Limited Maximization Domain: Insights from [2110.03950]
Optimizing Nonconvex-Nonconcave Min-Max Problems with a Limited Maximization Domain: Insights from [2110.03950]
Comparisons
Unlock Agentic Coding: Experimenting with Qwen 3.8 Flash-Next 176B Model on NVIDIA GB300 NVL72
Unlock Agentic Coding: Experimenting with Qwen 3.8 Flash-Next 176B Model on NVIDIA GB300 NVL72
Tools
Optimizing Social Media Safety: Scalable Few-Shot Harmful Content Moderation with Large Language Models
Optimizing Social Media Safety: Scalable Few-Shot Harmful Content Moderation with Large Language Models
Comparisons
Exploring Infinite-Dimensional Generative Diffusions through Doob’s h-Transform: A 2602.06621 Study
Exploring Infinite-Dimensional Generative Diffusions through Doob’s h-Transform: A 2602.06621 Study
Comparisons
//

Leading global tech insights for 20M+ innovators

Quick Link

  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events

Support

  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us

Sign Up for Our Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

AIModelKitAIModelKit
Follow US
© 2025 AI Model Kit. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?