Enhancing Local Applications with AI: The Power of DIN Deploy
Incorporating AI models into local applications is essential for developers seeking to harness advanced functionalities. However, this integration requires a compatible model format, a robust runtime, and acceleration capabilities that work seamlessly across different systems. Enter Do Inference Now (DIN) Deploy, the open-source solution simplifying AI model deployment for applications on Windows and Linux.
Bridging the AI Gap
DIN Deploy is designed to facilitate the transition from model checkpoints to fully operational, hardware-accelerated applications. By utilizing ONNX Runtime combined with NVIDIA TensorRT RTX execution provider, DIN Deploy offers developers a streamlined pathway to optimize their applications. This toolkit leverages a common ONNX Runtime API accessible through WinML 2.0, providing flexibility in deployment.
From ONNX Export to C++ Implementation
Every DIN Deploy sample begins with an intuitive Python exporter. This tool downloads model checkpoints from Hugging Face and transforms them into ONNX artifacts—a critical step for developers wanting to ensure their models are ready for deployment. Meanwhile, the application layer revolves around a native C++ CLI crafted using ONNX Runtime (ORT). This division allows developers to manage model conversion independently from deployment logic, reducing complexity.
Most sample codes rely on ONNX Runtime session and tensor APIs in C++. If necessary, vendor-specific codes, such as CUDA APIs, are infused only into optional acceleration paths. Thanks to ORT’s copy tensor API, data locality remains manageable without relying heavily on vendor-specific APIs in shared code.
Furthermore, for preprocessing and postprocessing surrounding the exported model inference, the FLUX.2 sample employs ONNX Runtime’s graphics interoperability feature, introduced in version 1.25. This capability allows developers to work with Vulkan and DirectX, thereby enriching data sampling.
A Multitude of AI Tasks Supported by DIN Deploy
DIN Deploy opens up a world of AI capabilities, supporting various tasks essential for modern applications.
-
Automatic Speech Recognition (ASR): Developers can implement both offline and streaming pipelines for ASR. OpenAI’s Whisper model facilitates offline transcription across diverse model sizes, while NVIDIA’s Parakeet TDT and Nemotron ASR Streaming cater to real-time needs.
-
Interactive Masking: The samples from Meta SAM 2.1 enable interactive masking for images and videos, generating segmentation masks that are practical for selection, tracking, and enhancing computer vision workflows.
Comparative analyses reveal notable performance distinctions between GPU and CPU implementations within DIN Deploy. For instance, during evaluations on DGX Spark, speed metrics showcase how GPU accelerations dramatically surpass CPU capabilities, as illustrated in Table 1.
Performance Insights
| Model | GPU (DGX Spark) | CPU (DGX Spark) |
|---|---|---|
openai/whisper-large-v3-turbo |
58.5x | 3.8x |
nvidia/nemotron-3.5-asr-streaming-0.6b |
39.01x | 3.24x |
nvidia/parakeet-tdt-0.6b-v3 |
206.41x | 14.44x |
facebook/sam2.1-hiera-base-plus |
38.3 FPS | 0.5 FPS |
Table 1 highlights the substantial benefits of utilizing GPU acceleration, demonstrating that higher efficiency leads to enhanced application performance in AI tasks.
Additionally, the FLUX.2-klein-4B sample illustrates prompt-driven image generation, revealing how cross-vendor interoperability with graphics APIs like Vulkan and DirectX can integrate GPU-resident resources into applications effortlessly. NVIDIA’s Model Optimizer facilitates post-training quantization (PTQ), which enhances model performance significantly. Despite hardware dependencies, the ONNX interfaces remain unchanged, allowing quantized models to serve as drop-in replacements without necessitating code alterations.
Getting Started with DIN Deploy
To jumpstart your experience with DIN Deploy, you can utilize the repository’s CMake presets supporting Windows, Linux, x86-64, and Arm64 architectures. After configuring and building your project, the next steps involve exporting a model to ONNX and running the command-line interface with TensorRT RTX. The CMake configuration streamlines this process by downloading both ONNX Runtime and TensorRT RTX automatically.
Further Learning
For a deeper dive into the realms of TensorRT for RTX, NVIDIA Local AI, and the comprehensive DIN Deploy repository, explore the extensive resources and tutorials available on the NVIDIA blog series, particularly regarding model quantization.
By adopting DIN Deploy, developers are empowered to tackle the complexities of integrating AI into local applications effectively, enhancing user experiences and creating innovative solutions across various platforms.
Inspired by: Source

