Accelerate BERT inference with Hugging Face Transformers and AWS Inferentia
What changed
The adoption of BERT and Transformers continues to grow. Transformer-based models are now not only achieving state-of-the-art performance in Natural Language Processing but also for Computer Vision, Speech, and Time-Series. ๐ฌ ๐ผ ๐ค โณ Companies are now slowly moving from the experimentation and research phase to the production phase in order to use transformer models for large-scale workloads.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Introducing Decision Transformers on Hugging Face ๐ค
- Hugging Face Text Generation Inference available for AWS Inferentia2
- Introducing ๐ค Accelerate
Sources
- Accelerate BERT inference with Hugging Face Transformers and AWS Inferentia (huggingface-blog)primary