๐Ÿ“š 3LM: A Benchmark for Arabic LLMs in STEM and Code

Practical AI: Tools, Models & Frameworksbenchmark

What changed

Arabic Large Language Models (LLMs) have seen notable progress in recent years, yet existing benchmarks fall short when it comes to evaluating performance in high-value technical domains. Most evaluations to date have focused on general-purpose tasks like summarization, sentiment analysis, or generic question answering. To address this gap, we introduce 3LM (ุนู„ู…), a multi-component benchmark tailored to evaluate Arabic LLMs on STEM (Science, Technology, Engineering, and Mathematics) subjects and code generation.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • The Open Arabic LLM Leaderboard 2
  • Introducing the Open Arabic LLM Leaderboard
  • Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More

Sources