Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More
What changed
As part of our ongoing efforts, we are excited to share the following updates: - Arabic-Leaderboards Space, launched in collaboration with Mohammed bin Zayed University of Artificial Intelligence (MBZUAI) to consolidate Arabic AI evaluations in one place. - AraGen 03-25 release with improvements and updated benchmark. AraGen-03-25 (SP2) Rankings As part of our December release, we introduced 3C3H as a new evaluation measure of the chat capability of models, aimed at assessing both the factuality and usability of LLMs’ answers.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Introducing the Open Arabic LLM Leaderboard
- The Open Arabic LLM Leaderboard 2
- Introducing Falcon-H1-Arabic: Pushing the Boundaries of Arabic Language AI with Hybrid Architecture
Sources
- Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More (huggingface-blog)primary