Nvidia Unveils NIM Microservices to Accelerate AI Model Deployment

Nvidia introduces NIM, a new software platform that packages AI models with optimized inferencing engines into easy-to-use microservices. The tool aims to drastically reduce the time it takes for businesses to deploy AI applications in production.

Nvidia unveils NIM, a new software platform designed to streamline the deployment of custom and pre-trained AI models into production environments. The system combines a given model with an optimized inferencing engine and packages it into a container that functions as a microservice, drastically reducing a process that typically takes developers weeks or months to complete.

The platform supports models from Nvidia, Adept, Cohere, Google, Meta, Microsoft, and Mistral AI, and is already partnering with Amazon, Google, and Microsoft to bring these microservices to major cloud services like SageMaker, Kubernetes Engine, and Azure AI. NIM utilizes Nvidia's Triton Inference Server, TensorRT, and TensorRT-LLM as its core inferencing engines to ensure maximum efficiency.

Specific Nvidia microservices available through NIM include Riva for speech and translation customization, cuOpt for routing optimizations, and Earth-2 for climate simulations. The company plans to expand these capabilities over time, with upcoming additions like the Nvidia RAG LLM operator that makes building custom data-driven generative AI chatbots significantly easier for enterprise developers.

Read More at the original source →