Introducing the Enterprise Scenarios Leaderboard: a Leaderboard for Real World Use Cases
What changed
The leaderboard aims to evaluate the performance of language models on real-world enterprise use cases. We currently support 6 diverse tasks - FinanceBench, Legal Confidentiality, Creative Writing, Customer Support Dialogue, Toxicity, and Enterprise PII. We measure the performance of models on metrics like accuracy, engagingness, toxicity, relevance, and Enterprise PII.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
- Introducing the Open Arabic LLM Leaderboard
- Introducing the Open FinLLM Leaderboard
Sources
- Introducing the Enterprise Scenarios Leaderboard: a Leaderboard for Real World Use Cases (huggingface-blog)primary