By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
AIModelKitAIModelKitAIModelKit
  • Home
  • News
    NewsShow More
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    4 Min Read
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    4 Min Read
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    5 Min Read
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    5 Min Read
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    4 Min Read
  • Open-Source Models
    Open-Source ModelsShow More
    GlucoFM: Advanced Foundation Model for Continuous Glucose Monitoring Insights
    GlucoFM: Advanced Foundation Model for Continuous Glucose Monitoring Insights
    5 Min Read
    AgentHands: Creating Interactive Hand Gestures for Enhanced Conversations with Spatially Grounded Agents in XR
    AgentHands: Creating Interactive Hand Gestures for Enhanced Conversations with Spatially Grounded Agents in XR
    5 Min Read
    Exploring How Mobility Enhances Language Models’ Understanding of Location
    Exploring How Mobility Enhances Language Models’ Understanding of Location
    5 Min Read
    Optimize Candidate Biomarkers with Our AI Tool for Wearable Sensor Data Analysis
    Optimize Candidate Biomarkers with Our AI Tool for Wearable Sensor Data Analysis
    4 Min Read
    Beyond BMI: Assessing Cardiometabolic Risk Using Smartphone Images
    Beyond BMI: Assessing Cardiometabolic Risk Using Smartphone Images
    5 Min Read
  • Guides
    GuidesShow More
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    6 Min Read
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    4 Min Read
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    6 Min Read
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    6 Min Read
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    6 Min Read
  • Tools
    ToolsShow More
    Unlock Agentic Coding: Experimenting with Qwen 3.8-Flash-Next on NVIDIA GB300 NVL72
    Unlock Agentic Coding: Experimenting with Qwen 3.8-Flash-Next on NVIDIA GB300 NVL72
    6 Min Read
    Unlock Agentic Coding: Experimenting with Qwen 3.8 Flash-Next 176B Model on NVIDIA GB300 NVL72
    Unlock Agentic Coding: Experimenting with Qwen 3.8 Flash-Next 176B Model on NVIDIA GB300 NVL72
    5 Min Read
    Optimizing LFM2.5 Q4_0 Checkpoints through Quantization-Aware Distillation Techniques
    Optimizing LFM2.5 Q4_0 Checkpoints through Quantization-Aware Distillation Techniques
    4 Min Read
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    6 Min Read
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    5 Min Read
  • Events
    EventsShow More
    Exploring the Future of EdTech: Highlights from the ‘Best of ISTE’ Virtual Playground
    Exploring the Future of EdTech: Highlights from the ‘Best of ISTE’ Virtual Playground
    4 Min Read
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    4 Min Read
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    5 Min Read
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    5 Min Read
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    5 Min Read
  • Ethics
    EthicsShow More
    Assessing the Environmental Impact of Data Centres: Are We Finally Acknowledging the Consequences?
    Assessing the Environmental Impact of Data Centres: Are We Finally Acknowledging the Consequences?
    5 Min Read
    Survey Reveals Surprising Impact of AI on Job Losses: Insights from Workers
    Survey Reveals Surprising Impact of AI on Job Losses: Insights from Workers
    6 Min Read
    Understanding DAO-to-DAO Voting: On-Chain and Off-Chain Mechanisms Explored
    Understanding DAO-to-DAO Voting: On-Chain and Off-Chain Mechanisms Explored
    5 Min Read
    Taiwan Prosecutes Nine Individuals for Smuggling Advanced AI Servers to China: A Tech Industry Update
    Taiwan Prosecutes Nine Individuals for Smuggling Advanced AI Servers to China: A Tech Industry Update
    4 Min Read
    Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
    Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
    6 Min Read
  • Comparisons
    ComparisonsShow More
    InternBootcamp: Enhancing LLM Reasoning Through Verifiable Task Scaling Techniques
    InternBootcamp: Enhancing LLM Reasoning Through Verifiable Task Scaling Techniques
    4 Min Read
    Enhancing Anomaly Detection in Collider Experiments through Contrastive Learning for Better Interpretability
    Enhancing Anomaly Detection in Collider Experiments through Contrastive Learning for Better Interpretability
    6 Min Read
    Exploring the Impact of Quantization on Self-Explanations in Large Language Models: Can LLMs Explain Themselves?
    Exploring the Impact of Quantization on Self-Explanations in Large Language Models: Can LLMs Explain Themselves?
    5 Min Read
    CytoNet: A Foundation Model for Understanding the Human Cerebral Cortex at Cellular Resolution
    CytoNet: A Foundation Model for Understanding the Human Cerebral Cortex at Cellular Resolution
    5 Min Read
    Optimizing Nonconvex-Nonconcave Min-Max Problems with a Limited Maximization Domain: Insights from [2110.03950]
    Optimizing Nonconvex-Nonconcave Min-Max Problems with a Limited Maximization Domain: Insights from [2110.03950]
    5 Min Read
Search
  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
Reading: OpenAI at QCon AI NYC: Mastering Enterprise Fine-Tuning Strategies
Share
Notification Show More
Font ResizerAa
AIModelKitAIModelKit
Font ResizerAa
  • 🏠
  • 🚀
  • 📰
  • 💡
  • 📚
  • ⭐
Search
  • Home
  • News
  • Models
  • Guides
  • Tools
  • Ethics
  • Events
  • Comparisons
Follow US
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
AIModelKit > Comparisons > OpenAI at QCon AI NYC: Mastering Enterprise Fine-Tuning Strategies
Comparisons

OpenAI at QCon AI NYC: Mastering Enterprise Fine-Tuning Strategies

aimodelkit
Last updated: December 18, 2025 2:00 am
aimodelkit
Share
SHARE

At QCon AI NYC 2025, a fascinating presentation by Will Hang from OpenAI unveiled the innovative approach of Agent RFT—a reinforcement fine-tuning method specifically designed to enhance the performance of tool-using agents.

Hang laid out a tactical roadmap that emphasizes initial improvements in prompt and task optimization before diving into adjustments in model weights. This pragmatic path includes simplifying task requirements, implementing guardrails to mitigate tool misuse, enhancing tool descriptions, and refining tool outputs to empower agents in making more informed downstream decisions. Notably, he observed that while these modifications can yield significant benefits, they may reach a plateau for tasks demanding consistent multi-step reasoning during tool interactions.

QCon AI Presentation

In his discourse, Hang positioned fine-tuning strategies along a spectrum. He highlighted supervised fine-tuning as particularly effective for tasks with predictable input-output mappings, aiming to replicate a specific style or structure. On the other hand, preference optimization was discussed as a technique for steering outputs towards desired responses through paired comparisons. According to OpenAI’s Direct Preference Optimization guide, this method currently focuses on text inputs and outputs. Hang argued that reinforcement fine-tuning is a more suitable choice for scenarios requiring the discovery of strategies over extended sequences rather than merely replicating a singular completion pattern.

Beware of reward hacking! Resolve any edge cases in your grader. Continuous rewards work better than binary rewards. – Will Hang, OpenAI

Hang introduced Agent RFT as a reinforcement fine-tuning approach tailored for tool-using agents, where models explore diverse strategies during training rollouts and receive feedback from a defined grader. OpenAI’s documentation illustrates this process as one that involves sampling potential responses, evaluating them with a custom grader, and updating the model based on those evaluations. He stressed the importance of credit assignment along the entire trajectory, so earlier choices—like tool selection and the structure of tool calls—can be reinforced or discouraged based on the outcomes downstream. In this context, an agent is not limited to responding to user prompts but is capable of interacting proactively with external tools.

Agent RFT in Action

Examples of tools discussed included coding terminals for agents, internal business systems for customer support, and search or retrieval endpoints for documents. Hang emphasized that tool outputs are integrated back into the context window, meaning that tool calls, outputs, reasoning tokens, and final responses collectively form a single, comprehensive multi-step trajectory. He noted that graders have become essential artifacts within this workflow, employing various grading techniques like simple matchers, model-based judges, code-based graders, endpoint graders, and hybrid systems optimizing both accuracy and latency.

The presentation also addressed operational characteristics beyond simple answer accuracy. Hang demonstrated how Agent RFT can help minimize unnecessary tool calls, uphold tool-call budgets, and reduce lengthy trajectories that might lead to unpredictable latency and a diminished user experience. Training traces illustrated a decrease in both reasoning tokens and tool calls, aligning with the concept that agents can learn to achieve equivalent or superior outcomes with fewer steps.

Following Hang’s insights, Wenjie Zi expanded on real-world applications, sharing platform setup details and use cases, including a finance-oriented scenario. In this example, an agent must sift through a large document corpus to locate relevant information while operating under a constrained tool-call budget. Here, the agent employed a series of tools—searching, listing, and reading files—with a grader evaluating the final output. Zi emphasized the advantage of a model-based grader, even for numeric responses, as it helps minimize false negatives caused by minor formatting differences or variations in units.

Use Cases of Agent RFT

Further illustrating the adaptability of Agent RFT, Zi covered broader applications in coding and other sectors, highlighting environments rich in tools, isolated execution contexts, and reward designs that strike a balance between correctness and efficiency. The reported benefits included enhanced planning abilities, shorter long trajectory tails, and in some instances, a pivot towards parallel tool calls, thereby accelerating responsiveness.

Developers eager to delve deeper into this cutting-edge approach can consult OpenAI’s Reinforcement Fine-Tuning and Model Optimization documentation. Future broadcasts on infoq.com will feature a video recording of Hang and Zi’s engaging presentation, making this invaluable content readily accessible.

Inspired by: Source

Self-Supervised Learning Techniques for Enhanced Social Recommendations: Insights from Paper 2412.18735
Revolutionizing Gesture Recognition: Redefining Domains for Enhanced WiFi-Based Interaction Beyond Physical Labels
Enhancing LLM Comprehension: Effective Step-by-Step Reading Strategies
Enhancing NLG Evaluation Prompts with Inversion Learning Techniques
Enhancing Insights into Reasoning Abilities of Large Language Models

Sign Up For Daily Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

By signing up, you agree to our Terms of Use and acknowledge the data practices in our Privacy Policy. You may unsubscribe at any time.
Share This Article
Facebook Copy Link Print
Previous Article Amazon Appoints New Leader for ‘AGI’ Group Amidst Competitive AI Landscape Amazon Appoints New Leader for ‘AGI’ Group Amidst Competitive AI Landscape
Next Article Senators Outraged as AI Toys Guide Kids on Knife Discovery Senators Outraged as AI Toys Guide Kids on Knife Discovery

Stay Connected

XFollow
PinterestPin
TelegramFollow
LinkedInFollow

							banner							
							banner
Explore Top AI Tools Instantly
Discover, compare, and choose the best AI tools in one place. Easy search, real-time updates, and expert-picked solutions.
Browse AI Tools

Latest News

Assessing the Environmental Impact of Data Centres: Are We Finally Acknowledging the Consequences?
Assessing the Environmental Impact of Data Centres: Are We Finally Acknowledging the Consequences?
Ethics
Exploring the Future of EdTech: Highlights from the ‘Best of ISTE’ Virtual Playground
Exploring the Future of EdTech: Highlights from the ‘Best of ISTE’ Virtual Playground
Events
Survey Reveals Surprising Impact of AI on Job Losses: Insights from Workers
Survey Reveals Surprising Impact of AI on Job Losses: Insights from Workers
Ethics
InternBootcamp: Enhancing LLM Reasoning Through Verifiable Task Scaling Techniques
InternBootcamp: Enhancing LLM Reasoning Through Verifiable Task Scaling Techniques
Comparisons
//

Leading global tech insights for 20M+ innovators

Quick Link

  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events

Support

  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us

Sign Up for Our Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

AIModelKitAIModelKit
Follow US
© 2025 AI Model Kit. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?