Optimization story: Bloom inference

Practical AI: Tools, Models & Frameworksinference

What changed

We achieved a 5x latency reduction over several weeks (and 50x more throughput). We wanted to share all the struggles and epic wins we went through to achieve such speed improvements. If your favorite flavor of optimizations is not discussed or improperly represented, we're sorry, please share it with us we're more than happy to try out new stuff and correct our mistakes.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Incredibly Fast BLOOM Inference with DeepSpeed and Accelerate
  • Introducing The World's Largest Open Multilingual Language Model: BLOOM
  • Introducing Optimum: The Optimization Toolkit for Transformers at Scale

Sources