Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with ๐Ÿค— Transformers

Practical AI: Tools, Models & Frameworksfine-tune

What changed

Wav2Vec2 is a pretrained model for Automatic Speech Recognition (ASR) and was released in September 2020 by Alexei Baevski, Michael Auli, and Alex Conneau. XLSR's successor, simply called XLS-R (refering to the ''XLM-R for Speech''), was released in November 2021 by Arun Babu, Changhan Wang, Andros Tjandra, et al. This is thanks to the new "Audio" feature introduced in datasets == 1.18.3, which loads and resamples audio files on-the-fly upon calling.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Fine-Tune W2V2-Bert for low-resource ASR with ๐Ÿค— Transformers
  • Fine-Tune Wav2Vec2 for English ASR in Hugging Face with ๐Ÿค— Transformers
  • Fine-Tune MMS Adapter Models for low-resource ASR

Sources