Serverless Inference with Hugging Face and NVIDIA NIM

Practical AI: Tools, Models & Frameworksinference

What changed

Update: This service is deprecated and no longer available as of April 10th, 2025. For an alternative, you should consider Inference Providers Today, we are thrilled to announce the launch of Hugging Face NVIDIA NIM API (serverless), a new service on the Hugging Face Hub, available to Enterprise Hub organizations. This new service makes it easy to use open models with the accelerated compute platform, of NVIDIA DGX Cloud accelerated compute platform for inference serving.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Bringing serverless GPU inference to Hugging Face users
  • Public AI on Hugging Face Inference Providers ๐Ÿ”ฅ
  • Groq on Hugging Face Inference Providers ๐Ÿ”ฅ

Sources