By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
AIModelKitAIModelKitAIModelKit
  • Home
  • News
    NewsShow More
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    4 Min Read
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    4 Min Read
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    5 Min Read
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    5 Min Read
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    4 Min Read
  • Open-Source Models
    Open-Source ModelsShow More
    Overcoming Recall Challenges: The Impact of Empty Shelves and Lost Keys on Parametric Factuality
    Overcoming Recall Challenges: The Impact of Empty Shelves and Lost Keys on Parametric Factuality
    6 Min Read
    Enhancing AMIE for Expert-Level Audio-Visual Clinical Consultations
    Enhancing AMIE for Expert-Level Audio-Visual Clinical Consultations
    5 Min Read
    Unlocking the Secrets of Diffusion Models: Understanding Their Creative Potential
    Unlocking the Secrets of Diffusion Models: Understanding Their Creative Potential
    5 Min Read
    Discover TabFM: A Zero-Shot Foundation Model Optimized for Tabular Data Analysis
    Discover TabFM: A Zero-Shot Foundation Model Optimized for Tabular Data Analysis
    5 Min Read
    Maximizing Cloud Cost Efficiency Through Linear Elastic Caching Strategies
    Maximizing Cloud Cost Efficiency Through Linear Elastic Caching Strategies
    5 Min Read
  • Guides
    GuidesShow More
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    6 Min Read
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    4 Min Read
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    6 Min Read
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    6 Min Read
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    6 Min Read
  • Tools
    ToolsShow More
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    6 Min Read
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    5 Min Read
    July 2026 Security Incident Disclosure: Key Insights and Updates
    July 2026 Security Incident Disclosure: Key Insights and Updates
    6 Min Read
    Boosting Performance with Native-Speed vLLM Transformers for Enhanced Modeling Backend
    Boosting Performance with Native-Speed vLLM Transformers for Enhanced Modeling Backend
    5 Min Read
    Hugging Face and Cerebras Launch Gemma 4 for Advanced Real-Time Voice AI Solutions
    Hugging Face and Cerebras Launch Gemma 4 for Advanced Real-Time Voice AI Solutions
    4 Min Read
  • Events
    EventsShow More
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    4 Min Read
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    5 Min Read
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    5 Min Read
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    5 Min Read
    Unlocking the Power of Open Models at Nemotron Labs: Discover the Advantage
    Unlocking the Power of Open Models at Nemotron Labs: Discover the Advantage
    7 Min Read
  • Ethics
    EthicsShow More
    Flock Strengthens Regulations to Address Rising Backlash Against Surveillance
    Flock Strengthens Regulations to Address Rising Backlash Against Surveillance
    5 Min Read
    How Brazil’s Child Online Safety Law Provides an Alternative to Social Media Bans
    How Brazil’s Child Online Safety Law Provides an Alternative to Social Media Bans
    6 Min Read
    Study Reveals AI’s Climate Benefits Diminished by Increased Fossil Fuel Support
    Study Reveals AI’s Climate Benefits Diminished by Increased Fossil Fuel Support
    6 Min Read
    Why No Degree is AI-Proof: How Delaying Specialization Can Give Students a Competitive Advantage
    Why No Degree is AI-Proof: How Delaying Specialization Can Give Students a Competitive Advantage
    6 Min Read
    Unveiling ‘The Download’: Exploring a Censorship Conspiracy Theory and the First AI-Created Virus
    Unveiling ‘The Download’: Exploring a Censorship Conspiracy Theory and the First AI-Created Virus
    6 Min Read
  • Comparisons
    ComparisonsShow More
    Enhancing LLM Robustness: A Comprehensive Diagnostic Stress Test for Decoding-Level Taboo
    Enhancing LLM Robustness: A Comprehensive Diagnostic Stress Test for Decoding-Level Taboo
    5 Min Read
    Enhanced Trust Region Constrained Bayesian Optimization: Effective Penalized Constraint Handling Techniques
    Enhanced Trust Region Constrained Bayesian Optimization: Effective Penalized Constraint Handling Techniques
    5 Min Read
    Enhancing Credit Risk Detection in Weixin Pay Using Billion-Scale Deep Graph Learning Techniques
    Enhancing Credit Risk Detection in Weixin Pay Using Billion-Scale Deep Graph Learning Techniques
    6 Min Read
    Spotify Develops External Index for Fast Point Queries on Its Data Lake
    5 Min Read
    Challenges of Multilingual Embedding Probes: Lack of Generalization Across Diverse Learner Corpora
    Challenges of Multilingual Embedding Probes: Lack of Generalization Across Diverse Learner Corpora
    5 Min Read
Search
  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
Reading: Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
Share
Notification Show More
Font ResizerAa
AIModelKitAIModelKit
Font ResizerAa
  • 🏠
  • 🚀
  • 📰
  • 💡
  • 📚
  • ⭐
Search
  • Home
  • News
  • Models
  • Guides
  • Tools
  • Ethics
  • Events
  • Comparisons
Follow US
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
AIModelKit > Tools > Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
Tools

Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning

aimodelkit
Last updated: August 13, 2026 4:00 am
aimodelkit
Share
Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
SHARE

Unveiling Alibaba’s Qwen3.8-2.4T-A95B: The Future of AI Model Deployment

Alibaba has made a significant leap in the world of artificial intelligence by unveiling the open weights for Qwen3.8-2.4T-A95B (often referred to as Qwen3.8-Max). This robust model boasts an extraordinary 2.4 trillion total parameters, with 95 billion activated per token, making it one of the largest open-weight models available today. What makes Qwen3.8-Max particularly noteworthy is its deployment versatility and cutting-edge architecture designed to tackle demanding reasoning and agentic workloads.

Contents
  • Breaking Down the Architecture
  • Compute Demands and Deployment Challenges
  • Architectural Innovations for Long-Context Inference
  • Efficient Use of Parameters with MoE
  • Tailored Reasoning Controls
  • Optimized Performance on the GB300 NVL72
  • Multiple Inference Options for Developers
  • Start Your Journey with Qwen3.8-2.4T-A95B

Breaking Down the Architecture

At its core, Qwen3.8-2.4T-A95B is based on a fine-grained mixture of experts (MoE) architecture. This hybrid design utilizes both full and linear attention, optimizing the model’s ability to handle tasks that require extensive context. With a context window of up to one million tokens and an output length of up to 128K, it offers a unique capability to manage increasingly complex workflows and data sets without losing efficiency.

Compute Demands and Deployment Challenges

Deploying a model of this scale is not trivial. It requires data-center-scale accelerated compute, underscoring the essential nature of extreme co-design across chips, system architecture, and software. Companies like NVIDIA collaborate closely with the open-source community to deliver the necessary infrastructure for multinode deployments. This includes optimized kernels, inference runtimes, and distributed serving recipes.

Out of the box, Qwen3.8-Max achieves impressive throughput metrics, exceeding 4,000 tokens per second per GPU, while delivering over 350 tokens per second per user on NVIDIA GB300 NVL72 using FP8 precision. Further advancements, including potential offerings like NVFP4 precision, are expected to enhance performance and efficiency even more.

Architectural Innovations for Long-Context Inference

Qwen3.8-2.4T-A95B excels in handling the most complex agentic workloads, which often involve:

More Read

Step-by-Step Guide: Hosting a Unity Game in a Virtual Space
Step-by-Step Guide: Hosting a Unity Game in a Virtual Space
Creating Native Multimodal Agents with Qwen 3.5 VLM on NVIDIA GPU-Accelerated Endpoints
DeepSpeed Joins PyTorch Foundation as a New Hosted Project: Enhancing AI Development
Optimizing olmOCR: Enhancing Accuracy for a Reliable OCR Engine
End of Password-Based Git Authentication: What You Need to Know
  • Coding scenarios
  • Large-scale document analysis
  • Long-running multi-step workflows

Unlike models optimized for instant chat responses, this architecture accumulates various system instructions, tool outputs, and reasoning traces over time. The challenge lies in maintaining efficiency as context expands—even up to one million tokens.

To address these challenges, the model alternates between full-attention and linear-attention layers. In the full-attention configuration, every token interacts with every other token, ensuring rich contextual understanding. Conversely, the linear-attention layers present a bounded recurrent state, effectively replacing the growing key-value (KV) cache and keeping both compute and memory requirements manageable.

Efficient Use of Parameters with MoE

One of the standout features of Qwen3.8-2.4T-A95B is its combination of fine-grained MoE, which makes the vast parameter count more practical and efficient to serve. Rather than relying on a handful of large experts, this model distributes its capacity across numerous smaller experts. This strategy not only enhances specialization but also boasts improved routing efficiency per unit of activated compute.

A learned router activates only the necessary experts for each token, thereby ensuring that the model’s operational costs align with active parameters, offering frontier-scale performance at a fraction of the traditional cost associated with dense models.

Tailored Reasoning Controls

Another noteworthy feature is the built-in reasoning controls (low/high/xhigh), which empower developers to configure inference depth based on specific task requirements. This adaptability allows for a trade-off between compute resources and reasoning quality—dial up for intricate, multi-step reasoning tasks or dial down for high-throughput document processing.

Optimized Performance on the GB300 NVL72

The NVIDIA GB300 NVL72 platform plays a crucial role in the high-performance capabilities of Qwen3.8-2.4T-A95B. This robust architecture integrates 72 NVIDIA Blackwell Ultra GPUs into one cohesive platform, facilitating efficient all-to-all communication at an impressive 130 TB/s. This eradicates the standard bottlenecks that typically hinder expert traffic across conventional networks.

With Qwen3.8-Max mounted on the NVIDIA GB300 NVL72, the model transcends expectations by delivering phenomenal throughput, allowing AI factories to operate large-parameter models efficiently in production environments.

Multiple Inference Options for Developers

NVIDIA has recognized the diverse needs of developers and offers various inference stacks tailored for enhanced control over performance. Solutions like SGLang, vLLM, and NVIDIA Dynamo equip developers with open-source inference recipes, ensuring they can leverage the NVIDIA-accelerated platform to their advantage.

Alternatively, a model-free NVIDIA NIM can facilitate deployment, allowing developers to serve any supported model effortlessly. This facilitates immediate post-training, enabling customized use cases with ease.

Start Your Journey with Qwen3.8-2.4T-A95B

For developers looking to get started with Qwen3.8-2.4T-A95B, model weights are available for download from Hugging Face or ModelScope. By deploying through a model-free NVIDIA NIM from NVIDIA NGC, developers can seamlessly integrate this advanced AI model into their applications, opening the doors to unprecedented capabilities across various verticals.

With its robust features and impressive flexibility, the Qwen3.8-2.4T-A95B is paving the way for future advancements in artificial intelligence, particularly in complex reasoning and multi-step workflows.

Inspired by: Source

Submit Your Proposals for PyTorch Day China 2025: Call for Contributions Now Open!
Hugging Face and Cloudflare Collaborate to Enhance Real-Time Speech and Video with FastRTC Integration
How to Enable Cluster Launch Control with TLX in PyTorch: A Step-by-Step Guide
NVIDIA Unveils 3 Million Sample Dataset for Enhanced OCR, Visual Question Answering, and Image Captioning Applications
Collaborating for a Brighter Future: Introducing OpenEnv and the Open Agent Ecosystem

Sign Up For Daily Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

By signing up, you agree to our Terms of Use and acknowledge the data practices in our Privacy Policy. You may unsubscribe at any time.
Share This Article
Facebook Copy Link Print
Previous Article Spotify Develops External Index for Fast Point Queries on Its Data Lake
Next Article Enhancing Credit Risk Detection in Weixin Pay Using Billion-Scale Deep Graph Learning Techniques Enhancing Credit Risk Detection in Weixin Pay Using Billion-Scale Deep Graph Learning Techniques

Stay Connected

XFollow
PinterestPin
TelegramFollow
LinkedInFollow

							banner							
							banner
Explore Top AI Tools Instantly
Discover, compare, and choose the best AI tools in one place. Easy search, real-time updates, and expert-picked solutions.
Browse AI Tools

Latest News

Enhancing LLM Robustness: A Comprehensive Diagnostic Stress Test for Decoding-Level Taboo
Enhancing LLM Robustness: A Comprehensive Diagnostic Stress Test for Decoding-Level Taboo
Comparisons
Flock Strengthens Regulations to Address Rising Backlash Against Surveillance
Flock Strengthens Regulations to Address Rising Backlash Against Surveillance
Ethics
Enhanced Trust Region Constrained Bayesian Optimization: Effective Penalized Constraint Handling Techniques
Enhanced Trust Region Constrained Bayesian Optimization: Effective Penalized Constraint Handling Techniques
Comparisons
Enhancing Credit Risk Detection in Weixin Pay Using Billion-Scale Deep Graph Learning Techniques
Enhancing Credit Risk Detection in Weixin Pay Using Billion-Scale Deep Graph Learning Techniques
Comparisons
//

Leading global tech insights for 20M+ innovators

Quick Link

  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events

Support

  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us

Sign Up for Our Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

AIModelKitAIModelKit
Follow US
© 2025 AI Model Kit. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?