Introducing ConTextual: How well can your Multimodal model jointly reason over text and image in text-rich scenes?

Practical AI: Tools, Models & Frameworksmultimodal

What changed

Introducing ConTextual: How well can your Multimodal model jointly reason over text and image in text-rich scenes? We refer to these tasks as "context-sensitive text-rich visual reasoning tasks". We also released a leaderboard, so that the community can see for themselves which models are the best at this task.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Introducing SynthID Text
  • Benchmarking Text Generation Inference
  • GPT-4

Sources