Open Arabic LLM Leaderboard

non-profit

Activity Feed

AI & ML interests

Arabic LLM Evaluation

Recent Activity

amztheory updated a dataset 39 minutes ago

OALL/details_QCRI__Fanar-1-9B-Instruct_v2_alrage

amztheory published a dataset 39 minutes ago

OALL/details_QCRI__Fanar-1-9B-Instruct_v2_alrage

alielfilali01 updated a dataset about 1 hour ago

OALL/requests_v2

View all activity

Articles

The Open Arabic LLM Leaderboard 2

Feb 10

• 32

OALL's activity

amztheory

updated a dataset 39 minutes ago

OALL/details_QCRI__Fanar-1-9B-Instruct_v2_alrage

Viewer • Updated 39 minutes ago • 4.21k

amztheory

published a dataset 39 minutes ago

OALL/details_QCRI__Fanar-1-9B-Instruct_v2_alrage

Viewer • Updated 39 minutes ago • 4.21k

alielfilali01

updated a dataset about 1 hour ago

OALL/requests_v2

Preview • Updated about 1 hour ago • 2.08k

amztheory

updated a dataset about 6 hours ago

OALL/requests_v2

Preview • Updated about 1 hour ago • 2.08k

amztheory

updated a dataset about 24 hours ago

OALL/details_google__gemma-3-27b-pt_v2_alrage

Viewer • Updated about 24 hours ago • 4.21k • 8

amztheory

published a dataset about 24 hours ago

OALL/details_google__gemma-3-27b-pt_v2_alrage

Viewer • Updated about 24 hours ago • 4.21k • 8

amztheory

updated a dataset 1 day ago

OALL/details_microsoft__Phi-4-mini-instruct_v2_alrage

Viewer • Updated 1 day ago • 4.21k • 11

amztheory

published a dataset 1 day ago

OALL/details_microsoft__Phi-4-mini-instruct_v2_alrage

Viewer • Updated 1 day ago • 4.21k • 11

amztheory

updated a dataset 1 day ago

OALL/details_microsoft__phi-4_v2_alrage

Viewer • Updated 1 day ago • 4.21k • 17

amztheory

published a dataset 1 day ago

OALL/details_microsoft__phi-4_v2_alrage

Viewer • Updated 1 day ago • 4.21k • 17

amztheory

updated a dataset 2 days ago

OALL/details_google__gemma-3-27b-it_v2_alrage

Viewer • Updated 2 days ago • 4.21k • 24

alielfilali01

authored a paper 9 days ago

Llama-3-Nanda-10B-Chat: An Open Generative Large Language Model for Hindi

Paper • 2504.06011 • Published Apr 8 • 1

clefourrier

posted an update 19 days ago

Post

611

Always surprised that so few people actually read the FineTasks blog, on
✨how to select training evals with the highest signal✨

If you're serious about training models without wasting compute on shitty runs, you absolutely should read it!!

An high signal eval actually tells you precisely, during training, how wel & what your model is learning, allowing you to discard the bad runs/bad samplings/...!

The blog covers in depth prompt choice, metrics, dataset, across languages/capabilities, and my fave section is "which properties should evals have"👌
(to know on your use case how to select the best evals for you)

Blog: HuggingFaceFW/blogpost-fine-tasks

2 replies

alielfilali01

posted an update about 1 month ago

Post

614

Great efforts from @AtlasIA folks to adapt text2image models (ghibli style) for Moroccan Context

Read the blog is here : https://huggingface.co/blog/atlasia/creating-your-custom-ghibli-text-to-image-model

clefourrier

posted an update 3 months ago

Post

2485

Gemma3 family is out! Reading the tech report, and this section was really interesting to me from a methods/scientific fairness pov.

Instead of doing over-hyped comparisons, they clearly state that **results are reported in a setup which is advantageous to their models**.
(Which everybody does, but people usually don't say)

For a tech report, it makes a lot of sense to report model performance when used optimally!
On leaderboards on the other hand, comparison will be apples to apples, but in a potentially unoptimal way for a given model family (like some user interact sub-optimally with models)

Also contains a cool section (6) on training data memorization rate too! Important to see if your model will output the training data it has seen as such: always an issue for privacy/copyright/... but also very much for evaluation!

Because if your model knows its evals by heart, you're not testing for generalization.

alielfilali01

posted an update 4 months ago

Post

1034

🚨 Arabic LLM Evaluation 🚨

Few models join the ranking of https://huggingface.co/spaces/inceptionai/AraGen-Leaderboard Today.

The new MistralAI model, Saba, is quite impressive, Top10 ! Well done @arthurmensch and team.

Sadly Mistral did not follow its strategy about public weights this time, we hope this changes soon and we get the model with a permissive license.

We added other Mistral models and apparently, we have been sleeping on mistralai/Mistral-Large-Instruct-2411 !

Another impressive model that joined the ranking today is ALLaM-AI/ALLaM-7B-Instruct-preview. After a long wait finally ALLaM is here and it is IMPRESSIVE given its size !

ALLaM is ranked on OALL/Open-Arabic-LLM-Leaderboard as well.

clefourrier

authored a paper 4 months ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published Feb 4 • 232

alielfilali01

posted an update 5 months ago

Post

2126

3C3H AraGen Leaderboard welcomes today deepseek-ai/DeepSeek-V3 and 12 other models (including the late gpt-3.5 💀) to the ranking of best LLMs in Arabic !

Observations:
- DeepSeek-v3 ranked 3rd and only Open model among the top 5 !

- A 14B open model ( Qwen/Qwen2.5-14B-Instruct) outperforms gpt-3.5-turbo-0125 (from last year). This shows how much we came in advancing and supporting Arabic presence within the LLM ecosystem !

- Contrary to what observed in likelihood-acc leaderboards (like OALL/Open-Arabic-LLM-Leaderboard) further finetuned models like maldv/Qwentile2.5-32B-Instruct actually decreased the performance compared to the original model Qwen/Qwen2.5-32B-Instruct.
It's worth to note that the decrease is statiscally insignificant which imply that at best, the out-domain finetuning do not really hurts the model original capabilities acquired during pretraining.
Previous work addressed this (finetuning VS pretraining) but more investigation in this regard is required (any PhDs here ? This could be your question ...)

Check out the latest rankings: https://huggingface.co/spaces/inceptionai/AraGen-Leaderboard

alielfilali01

posted an update 5 months ago

Post

2047

~75% on the challenging GPQA with only 40M parameters 🔥🥳

GREAT ACHIEVEMENT ! Or is it ?

This new Work, "Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation", take out the mystery about many models i personally suspected their results. Speacially on leaderboards other than the english one, Like the Open Arabic LLM Leaderbaord OALL/Open-Arabic-LLM-Leaderboard.

The authors of this work, first started by training a model on the GPQA data, which, unsurprisingly, led to the model achieving 100% performance.

Afterward, they trained what they referred to as a 'legitimate' model on legitimate data (MedMCQA). However, they introduced a distillation loss from the earlier, 'cheated' model.

What they discovered was fascinating: the knowledge of GPQA leaked through this distillation loss, even though the legitimate model was never explicitly trained on GPQA during this stage.

This raises important questions about the careful use of distillation in model training, especially when the training data is opaque. As they demonstrated, it’s apparently possible to (intentionally or unintentionally) leak test data through this method.

Find out more: Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation (2412.15255)

1 reply

alielfilali01

posted an update 6 months ago

Post

3538

Unpopular opinion: Open Source takes courage to do !

Not everyone is brave enough to release what they have done (the way they've done it) to the wild to be judged !
It really requires a high level of "knowing wth are you doing" ! It's kind of a super power !

Cheers to the heroes here who see this!

5 replies

AI & ML interests

Recent Activity

Articles

The Open Arabic LLM Leaderboard 2

Team members 11

OALL's activity