Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models

Practical AI: Tools, Models & Frameworksfine-tuning

What changed

Florence supports many tasks out of the box: captioning, object detection, OCR, and more. The authors report that Florence 2 can perform visual question answering (VQA), but the released models don't include VQA capability. We encourage the open-source community to leverage this fine-tuning tutorial and explore the remarkable potential of Florence-2 for a wide range of new tasks!

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Introducing vision to the fine-tuning API
  • Fine-Tuning Gemma Models in Hugging Face
  • Building smarter maps with GPT-4o vision fine-tuning

Sources