Overview of natively supported quantization schemes in ๐Ÿค— Transformers

Practical AI: Tools, Models & Frameworksquantization

What changed

Currently, quantizing models are used for two main purposes: - Running inference of a large model on a smaller device - Fine-tune adapters on top of quantized models So far, two integration efforts have been made and are natively supported in transformers : bitsandbytes and auto-gptq. Note that some additional quantization schemes are also supported in the ๐Ÿค— optimum library, but this is out of scope for this blogpost. To learn more about each of the supported schemes, please have a look at one of the resources shared below.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • An overview of inference solutions on Hugging Face
  • Introducing Decision Transformers on Hugging Face ๐Ÿค—
  • Exploring Quantization Backends in Diffusers

Sources