Accelerating SD Turbo and SDXL Turbo Inference with ONNX Runtime and Olive
What changed
SD Turbo is a distilled version of Stable Diffusion 2.1, and SDXL Turbo is a distilled version of SDXL 1.0. In this post, we will introduce optimizations in the ONNX Runtime CUDA and TensorRT execution providers that speed up inference of SD Turbo and SDXL Turbo on NVIDIA GPUs significantly. Additionally, Oracle has released a Stable Diffusion sample with Java that runs inference on top of ONNX Runtime.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- GPT-3.5 Turbo fine-tuning and API updates
- New models and developer products announced at DevDay
- Accelerating Stable Diffusion Inference on Intel CPUs
Sources
- Accelerating SD Turbo and SDXL Turbo Inference with ONNX Runtime and Olive (huggingface-blog)primary