By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
AIModelKitAIModelKitAIModelKit
  • Home
  • News
    NewsShow More
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    4 Min Read
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    4 Min Read
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    5 Min Read
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    5 Min Read
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    4 Min Read
  • Open-Source Models
    Open-Source ModelsShow More
    Enhancing AI Image Generation with Diffusion Controller: A Simplified Unified Approach
    Enhancing AI Image Generation with Diffusion Controller: A Simplified Unified Approach
    5 Min Read
    Effortless Long-Form Video Creation: Automating Coherent Content Generation
    Effortless Long-Form Video Creation: Automating Coherent Content Generation
    5 Min Read
    Overcoming Inference Bottlenecks: Speeding Up Complex AI Search with Retrieve-for-Train
    Overcoming Inference Bottlenecks: Speeding Up Complex AI Search with Retrieve-for-Train
    5 Min Read
    ToolGrad: Generate Efficient Tool-Use Datasets Using Textual Gradients
    ToolGrad: Generate Efficient Tool-Use Datasets Using Textual Gradients
    5 Min Read
    Enhancing Genomic Prediction in Underserved Populations through Transfer Learning
    Enhancing Genomic Prediction in Underserved Populations through Transfer Learning
    5 Min Read
  • Guides
    GuidesShow More
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    6 Min Read
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    4 Min Read
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    6 Min Read
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    6 Min Read
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    6 Min Read
  • Tools
    ToolsShow More
    Create Local AI Applications Using C++ and NVIDIA TensorRT RTX Samples
    Create Local AI Applications Using C++ and NVIDIA TensorRT RTX Samples
    5 Min Read
    Unlock Near-Astra Intelligence in Your Daily Work with GPT-6.1 Sol on Amazon Bedrock
    Unlock Near-Astra Intelligence in Your Daily Work with GPT-6.1 Sol on Amazon Bedrock
    6 Min Read
    Reproducible Benchmark Results: How UK AISI and EvalEval Are Leading the Way
    Reproducible Benchmark Results: How UK AISI and EvalEval Are Leading the Way
    6 Min Read
    Hugging Face Welcomes Jun Kim, oMLX Creator and Maintainer, to Boost the MLX Community
    Hugging Face Welcomes Jun Kim, oMLX Creator and Maintainer, to Boost the MLX Community
    4 Min Read
    AWS Crowned Leader in The Forrester Wave: AI Infrastructure Solutions, Q4 2025 Report
    AWS Crowned Leader in The Forrester Wave: AI Infrastructure Solutions, Q4 2025 Report
    5 Min Read
  • Events
    EventsShow More
    Jensen Huang at Dreamforce: ‘Now We Can Know Everything and Achieve Anything’
    Jensen Huang at Dreamforce: ‘Now We Can Know Everything and Achieve Anything’
    5 Min Read
    Essential Strategies for Preparing Students for a Career in Quantum Computing
    Essential Strategies for Preparing Students for a Career in Quantum Computing
    5 Min Read
    Skild AI Leverages NVIDIA’s Physical AI to Enable Robots to Learn New Tasks from Just One Video
    Skild AI Leverages NVIDIA’s Physical AI to Enable Robots to Learn New Tasks from Just One Video
    6 Min Read
    Top 4 Mistakes New Teachers Make and Proven Strategies to Overcome Them
    Top 4 Mistakes New Teachers Make and Proven Strategies to Overcome Them
    5 Min Read
    NVIDIA Set to Acquire Hugging Face: What This Means for AI Development
    NVIDIA Set to Acquire Hugging Face: What This Means for AI Development
    5 Min Read
  • Ethics
    EthicsShow More
    Trump’s AI Safety Accord: A Closer Look at Its True Significance
    Trump’s AI Safety Accord: A Closer Look at Its True Significance
    6 Min Read
    Trump Unveils Unclear ‘Morally Binding’ AI Agreement with Tech CEOs for Enhanced Self-Policing
    Trump Unveils Unclear ‘Morally Binding’ AI Agreement with Tech CEOs for Enhanced Self-Policing
    5 Min Read
    Enhancing Resilience: How AI Agents Can Strengthen New Zealand’s Vulnerable Supply Chains
    Enhancing Resilience: How AI Agents Can Strengthen New Zealand’s Vulnerable Supply Chains
    7 Min Read
    How Meta’s Settlement Won’t Restore the Time Lost to Social Media
    How Meta’s Settlement Won’t Restore the Time Lost to Social Media
    6 Min Read
    Can an ‘Australian AI’ Safeguard Us from Hacks? Understanding the Complexity of Cybersecurity
    Can an ‘Australian AI’ Safeguard Us from Hacks? Understanding the Complexity of Cybersecurity
    5 Min Read
  • Comparisons
    ComparisonsShow More
    InternBootcamp: Enhancing LLM Reasoning Through Verifiable Task Scaling Techniques
    InternBootcamp: Enhancing LLM Reasoning Through Verifiable Task Scaling Techniques
    4 Min Read
    Enhancing Anomaly Detection in Collider Experiments through Contrastive Learning for Better Interpretability
    Enhancing Anomaly Detection in Collider Experiments through Contrastive Learning for Better Interpretability
    6 Min Read
    Exploring the Impact of Quantization on Self-Explanations in Large Language Models: Can LLMs Explain Themselves?
    Exploring the Impact of Quantization on Self-Explanations in Large Language Models: Can LLMs Explain Themselves?
    5 Min Read
    CytoNet: A Foundation Model for Understanding the Human Cerebral Cortex at Cellular Resolution
    CytoNet: A Foundation Model for Understanding the Human Cerebral Cortex at Cellular Resolution
    5 Min Read
    Optimizing Nonconvex-Nonconcave Min-Max Problems with a Limited Maximization Domain: Insights from [2110.03950]
    Optimizing Nonconvex-Nonconcave Min-Max Problems with a Limited Maximization Domain: Insights from [2110.03950]
    5 Min Read
Search
  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
Reading: Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
Share
Notification Show More
Font ResizerAa
AIModelKitAIModelKit
Font ResizerAa
  • 🏠
  • 🚀
  • 📰
  • 💡
  • 📚
  • ⭐
Search
  • Home
  • News
  • Models
  • Guides
  • Tools
  • Ethics
  • Events
  • Comparisons
Follow US
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
AIModelKit > Tools > Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
Tools

Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning

aimodelkit
Last updated: August 13, 2026 4:00 am
aimodelkit
Share
Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
SHARE

Unveiling Alibaba’s Qwen3.8-2.4T-A95B: The Future of AI Model Deployment

Alibaba has made a significant leap in the world of artificial intelligence by unveiling the open weights for Qwen3.8-2.4T-A95B (often referred to as Qwen3.8-Max). This robust model boasts an extraordinary 2.4 trillion total parameters, with 95 billion activated per token, making it one of the largest open-weight models available today. What makes Qwen3.8-Max particularly noteworthy is its deployment versatility and cutting-edge architecture designed to tackle demanding reasoning and agentic workloads.

Contents
  • Breaking Down the Architecture
  • Compute Demands and Deployment Challenges
  • Architectural Innovations for Long-Context Inference
  • Efficient Use of Parameters with MoE
  • Tailored Reasoning Controls
  • Optimized Performance on the GB300 NVL72
  • Multiple Inference Options for Developers
  • Start Your Journey with Qwen3.8-2.4T-A95B

Breaking Down the Architecture

At its core, Qwen3.8-2.4T-A95B is based on a fine-grained mixture of experts (MoE) architecture. This hybrid design utilizes both full and linear attention, optimizing the model’s ability to handle tasks that require extensive context. With a context window of up to one million tokens and an output length of up to 128K, it offers a unique capability to manage increasingly complex workflows and data sets without losing efficiency.

Compute Demands and Deployment Challenges

Deploying a model of this scale is not trivial. It requires data-center-scale accelerated compute, underscoring the essential nature of extreme co-design across chips, system architecture, and software. Companies like NVIDIA collaborate closely with the open-source community to deliver the necessary infrastructure for multinode deployments. This includes optimized kernels, inference runtimes, and distributed serving recipes.

Out of the box, Qwen3.8-Max achieves impressive throughput metrics, exceeding 4,000 tokens per second per GPU, while delivering over 350 tokens per second per user on NVIDIA GB300 NVL72 using FP8 precision. Further advancements, including potential offerings like NVFP4 precision, are expected to enhance performance and efficiency even more.

Architectural Innovations for Long-Context Inference

Qwen3.8-2.4T-A95B excels in handling the most complex agentic workloads, which often involve:

More Read

Optimizing Dynamic Kernel Selection with KleidiAI and Quantized Tied Embeddings in PyTorch
Optimizing Dynamic Kernel Selection with KleidiAI and Quantized Tied Embeddings in PyTorch
Stanford Das Lab Boosts RNA Folding Research Efficiency Using NVIDIA DGX Cloud Technology
Step-by-Step Guide: Hosting a Unity Game in a Virtual Space
Exciting News: XetHub Joins Forces with Hugging Face!
DeepSpeed Joins PyTorch Foundation as a New Hosted Project: Enhancing AI Development
  • Coding scenarios
  • Large-scale document analysis
  • Long-running multi-step workflows

Unlike models optimized for instant chat responses, this architecture accumulates various system instructions, tool outputs, and reasoning traces over time. The challenge lies in maintaining efficiency as context expands—even up to one million tokens.

To address these challenges, the model alternates between full-attention and linear-attention layers. In the full-attention configuration, every token interacts with every other token, ensuring rich contextual understanding. Conversely, the linear-attention layers present a bounded recurrent state, effectively replacing the growing key-value (KV) cache and keeping both compute and memory requirements manageable.

Efficient Use of Parameters with MoE

One of the standout features of Qwen3.8-2.4T-A95B is its combination of fine-grained MoE, which makes the vast parameter count more practical and efficient to serve. Rather than relying on a handful of large experts, this model distributes its capacity across numerous smaller experts. This strategy not only enhances specialization but also boasts improved routing efficiency per unit of activated compute.

A learned router activates only the necessary experts for each token, thereby ensuring that the model’s operational costs align with active parameters, offering frontier-scale performance at a fraction of the traditional cost associated with dense models.

Tailored Reasoning Controls

Another noteworthy feature is the built-in reasoning controls (low/high/xhigh), which empower developers to configure inference depth based on specific task requirements. This adaptability allows for a trade-off between compute resources and reasoning quality—dial up for intricate, multi-step reasoning tasks or dial down for high-throughput document processing.

Optimized Performance on the GB300 NVL72

The NVIDIA GB300 NVL72 platform plays a crucial role in the high-performance capabilities of Qwen3.8-2.4T-A95B. This robust architecture integrates 72 NVIDIA Blackwell Ultra GPUs into one cohesive platform, facilitating efficient all-to-all communication at an impressive 130 TB/s. This eradicates the standard bottlenecks that typically hinder expert traffic across conventional networks.

With Qwen3.8-Max mounted on the NVIDIA GB300 NVL72, the model transcends expectations by delivering phenomenal throughput, allowing AI factories to operate large-parameter models efficiently in production environments.

Multiple Inference Options for Developers

NVIDIA has recognized the diverse needs of developers and offers various inference stacks tailored for enhanced control over performance. Solutions like SGLang, vLLM, and NVIDIA Dynamo equip developers with open-source inference recipes, ensuring they can leverage the NVIDIA-accelerated platform to their advantage.

Alternatively, a model-free NVIDIA NIM can facilitate deployment, allowing developers to serve any supported model effortlessly. This facilitates immediate post-training, enabling customized use cases with ease.

Start Your Journey with Qwen3.8-2.4T-A95B

For developers looking to get started with Qwen3.8-2.4T-A95B, model weights are available for download from Hugging Face or ModelScope. By deploying through a model-free NVIDIA NIM from NVIDIA NGC, developers can seamlessly integrate this advanced AI model into their applications, opening the doors to unprecedented capabilities across various verticals.

With its robust features and impressive flexibility, the Qwen3.8-2.4T-A95B is paving the way for future advancements in artificial intelligence, particularly in complex reasoning and multi-step workflows.

Inspired by: Source

Triton-Powered Operator Library for Accelerating Universal AI with PyTorch
NVIDIA cuQuantum Enhances Simulation Speed with Dynamic Gradients and DMRG Features
Discover Snowball Fight ☃️: Our First ML-Agents Environment for Exciting Gameplay
How AI Technology Safeguards Marine Life by Locating Abandoned Fishing Nets in Oceans
Enhance Your LLMs Using Gradio MCP Servers for Effective Upskilling

Sign Up For Daily Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

By signing up, you agree to our Terms of Use and acknowledge the data practices in our Privacy Policy. You may unsubscribe at any time.
Share This Article
Facebook Copy Link Print
Previous Article Spotify Develops External Index for Fast Point Queries on Its Data Lake
Next Article Enhancing Credit Risk Detection in Weixin Pay Using Billion-Scale Deep Graph Learning Techniques Enhancing Credit Risk Detection in Weixin Pay Using Billion-Scale Deep Graph Learning Techniques

Stay Connected

XFollow
PinterestPin
TelegramFollow
LinkedInFollow

							banner							
							banner
Explore Top AI Tools Instantly
Discover, compare, and choose the best AI tools in one place. Easy search, real-time updates, and expert-picked solutions.
Browse AI Tools

Latest News

Create Local AI Applications Using C++ and NVIDIA TensorRT RTX Samples
Create Local AI Applications Using C++ and NVIDIA TensorRT RTX Samples
Tools
Trump’s AI Safety Accord: A Closer Look at Its True Significance
Trump’s AI Safety Accord: A Closer Look at Its True Significance
Ethics
Enhancing AI Image Generation with Diffusion Controller: A Simplified Unified Approach
Enhancing AI Image Generation with Diffusion Controller: A Simplified Unified Approach
Open-Source Models
Trump Unveils Unclear ‘Morally Binding’ AI Agreement with Tech CEOs for Enhanced Self-Policing
Trump Unveils Unclear ‘Morally Binding’ AI Agreement with Tech CEOs for Enhanced Self-Policing
Ethics
//

Leading global tech insights for 20M+ innovators

Quick Link

  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events

Support

  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us

Sign Up for Our Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

AIModelKitAIModelKit
Follow US
© 2025 AI Model Kit. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?