Scaling-up BERT Inference on CPU (Part 1)

Practical AI: Tools, Models & Frameworksinference

What changed

Back in October 2019, my colleague Lysandre Debut published a comprehensive (at the time) inference performance benchmarking blog (1). Since then, ๐Ÿค— transformers (2) welcomed a tremendous number of new architectures and thousands of new models were added to the ๐Ÿค— hub (3) which now counts more than 9,000 of them as of first quarter of 2021. As the NLP landscape keeps trending towards more and more BERT-like models being used in production, it remains challenging to efficiently deploy and run these architectures at scale.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Scaling up BERT-like model Inference on modern CPU - Part 2
  • LFM2.5-Encoders for Fast Long-Context Inference on CPU
  • Accelerate BERT inference with Hugging Face Transformers and AWS Inferentia

Sources