Fast Inference on Large Language Models: BLOOMZ on Habana Gaudi2 Accelerator

Practical AI: Tools, Models & Frameworksinference

What changed

As models get bigger and bigger, deploying them into production to run inference has become increasingly challenging. More recently, another model with the exact same architecture was released: BLOOMZ, which is a fine-tuned version of BLOOM on several tasks leading to better generalization and zero-shot[^1] capabilities. We will update these numbers as new versions of SynapseAI are released and integrated within Optimum Habana.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • The Open Medical-LLM Leaderboard: Benchmarking Large Language Models in Healthcare
  • A hazard analysis framework for code synthesis large language models
  • Faster Training and Inference: Habana Gaudi®2 vs Nvidia A100 80GB

Sources