Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers

Practical AI: Tools, Models & Frameworksmultimodal

What changed

As a practical example, I'll walk through finetuning Qwen/Qwen3-VL-Embedding-2B for Visual Document Retrieval (VDR), the task of retrieving relevant document pages (as images, with charts, tables, and layout intact) for a given text query. The resulting tomaarsen/Qwen3-VL-Embedding-2B-vdr demonstrates how much performance you can gain by finetuning on your own domain. If you're new to multimodal models in Sentence Transformers, I recommend reading Multimodal Embedding & Reranker Models with Sentence Transformers first.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Multimodal Embedding & Reranker Models with Sentence Transformers
  • Train and Fine-Tune Sentence Transformers Models
  • New embedding models and API updates

Sources