Scaling up BERT-like model Inference on modern CPU - Part 2

Practical AI: Tools, Models & Frameworksinference

What changed

In this blog post, we will focus on software optimizations and give you a sense of the performances of the new Ice Lake generation of Xeon CPUs from Intel. As in the previous blog post, we show the performance with benchmark results and charts, along with new tools to make all these knobs and features easy to use. Back in April, Intel launched its latest generation of Intel Xeon processors, codename Ice Lake, targeting more efficient and performant AI workloads.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Scaling-up BERT Inference on CPU (Part 1)
  • LFM2.5-Encoders for Fast Long-Context Inference on CPU
  • Accelerate BERT inference with Hugging Face Transformers and AWS Inferentia

Sources