Scaling-up BERT Inference on CPU (Part 1)
What changed
Back in October 2019, my colleague Lysandre Debut published a comprehensive (at the time) inference performance benchmarking blog (1). Since then, ๐ค transformers (2) welcomed a tremendous number of new architectures and thousands of new models were added to the ๐ค hub (3) which now counts more than 9,000 of them as of first quarter of 2021. As the NLP landscape keeps trending towards more and more BERT-like models being used in production, it remains challenging to efficiently deploy and run these architectures at scale.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Scaling up BERT-like model Inference on modern CPU - Part 2
- LFM2.5-Encoders for Fast Long-Context Inference on CPU
- Accelerate BERT inference with Hugging Face Transformers and AWS Inferentia
Sources
- Scaling-up BERT Inference on CPU (Part 1) (huggingface-blog)primary