Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with ๐ค Transformers
What changed
Wav2Vec2 is a pretrained model for Automatic Speech Recognition (ASR) and was released in September 2020 by Alexei Baevski, Michael Auli, and Alex Conneau. XLSR's successor, simply called XLS-R (refering to the ''XLM-R for Speech''), was released in November 2021 by Arun Babu, Changhan Wang, Andros Tjandra, et al. This is thanks to the new "Audio" feature introduced in datasets == 1.18.3, which loads and resamples audio files on-the-fly upon calling.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Fine-Tune W2V2-Bert for low-resource ASR with ๐ค Transformers
- Fine-Tune Wav2Vec2 for English ASR in Hugging Face with ๐ค Transformers
- Fine-Tune MMS Adapter Models for low-resource ASR
Sources
- Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with ๐ค Transformers (huggingface-blog)primary