๐ 3LM: A Benchmark for Arabic LLMs in STEM and Code
What changed
Arabic Large Language Models (LLMs) have seen notable progress in recent years, yet existing benchmarks fall short when it comes to evaluating performance in high-value technical domains. Most evaluations to date have focused on general-purpose tasks like summarization, sentiment analysis, or generic question answering. To address this gap, we introduce 3LM (ุนูู ), a multi-component benchmark tailored to evaluate Arabic LLMs on STEM (Science, Technology, Engineering, and Mathematics) subjects and code generation.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- The Open Arabic LLM Leaderboard 2
- Introducing the Open Arabic LLM Leaderboard
- Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More
Sources
- ๐ 3LM: A Benchmark for Arabic LLMs in STEM and Code (huggingface-blog)primary