Introducing the Enterprise Scenarios Leaderboard: a Leaderboard for Real World Use Cases

Practical AI: Tools, Models & Frameworks

What changed

The leaderboard aims to evaluate the performance of language models on real-world enterprise use cases. We currently support 6 diverse tasks - FinanceBench, Legal Confidentiality, Creative Writing, Customer Support Dialogue, Toxicity, and Enterprise PII. We measure the performance of models on metrics like accuracy, engagingness, toxicity, relevance, and Enterprise PII.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
  • Introducing the Open Arabic LLM Leaderboard
  • Introducing the Open FinLLM Leaderboard

Sources