Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers
What changed
As a practical example, I'll walk through finetuning Qwen/Qwen3-VL-Embedding-2B for Visual Document Retrieval (VDR), the task of retrieving relevant document pages (as images, with charts, tables, and layout intact) for a given text query. The resulting tomaarsen/Qwen3-VL-Embedding-2B-vdr demonstrates how much performance you can gain by finetuning on your own domain. If you're new to multimodal models in Sentence Transformers, I recommend reading Multimodal Embedding & Reranker Models with Sentence Transformers first.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Multimodal Embedding & Reranker Models with Sentence Transformers
- Train and Fine-Tune Sentence Transformers Models
- New embedding models and API updates
Sources
- Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers (huggingface-blog)primary