Instructions to use reemalyami/AraRoBERTa-LB with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use reemalyami/AraRoBERTa-LB with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="reemalyami/AraRoBERTa-LB")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("reemalyami/AraRoBERTa-LB") model = AutoModelForMaskedLM.from_pretrained("reemalyami/AraRoBERTa-LB", device_map="auto") - Notebooks
- Google Colab
- Kaggle
AraRoBERTa-LB
AraRoBERTa-LB is a RoBERTa-base model pre-trained from scratch on Lebanese Arabic text. It is one of seven mono-dialectal models introduced in Weakly and Semi-Supervised Learning for Arabic Text Classification using Monodialectal Language Models (WANLP 2022).
| Architecture | RoBERTa-base (12 layers, 768 hidden, 126M params), trained from scratch with MLM |
| Tokenizer | Byte-level BPE, 52K vocabulary |
| Pre-training data | 204,430 geolocated tweets from Lebanon (3.6M tokens) |
| Training | Batch size 32, 10 epochs, 1× Tesla P100 |
Usage
from transformers import pipeline
fill_mask = pipeline("fill-mask", model="reemalyami/AraRoBERTa-LB")
fill_mask("الجو اليوم كتير <mask>")
For classification, load the model with AutoModelForSequenceClassification and fine-tune it as usual. During pre-training, text were normalized (Arabic letter normalization, with digits and character elongation removed), so applying similar preprocessing to your input may help.
Results
Binary dialect identification (F1) on manually annotated text plus NADI 2020 data:
| AraRoBERTa-LB | AraBERT | mBERT | XLM-R | LR (TF-IDF) |
|---|---|---|---|---|
| 0.849 | 0.849 | 0.879 | 0.866 | 0.892 |
With the paper's semi-supervised method, F1 is 0.88. With weak supervision from dialect dictionaries, it is 0.78.
The AraRoBERTa Family
| Model | Dialect | Tokens | Supervised F1 |
|---|---|---|---|
| AraRoBERTa-SA | Saudi Arabia | 45.4M | 0.836 |
| AraRoBERTa-EGY | Egypt | 37.2M | 0.934 |
| AraRoBERTa-KU | Kuwait | 8.9M | 0.916 |
| AraRoBERTa-OM | Oman | 3.8M | 0.718 |
| AraRoBERTa-LB (this model) | Lebanon | 3.6M | 0.849 |
| AraRoBERTa-JO | Jordan | 2.6M | 0.848 |
| AraRoBERTa-DZ | Algeria | 1.9M | 0.859 |
Citation
@inproceedings{alyami-al-zaidy-2022-weakly,
title = "Weakly and Semi-Supervised Learning for {A}rabic Text Classification using Monodialectal Language Models",
author = "AlYami, Reem and Al-Zaidy, Rabah",
booktitle = "Proceedings of the The Seventh Arabic Natural Language Processing Workshop (WANLP)",
month = dec,
year = "2022",
address = "Abu Dhabi, United Arab Emirates (Hybrid)",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.wanlp-1.24",
pages = "260--272",
}
Contact: Reem AlYami · LinkedIn · reem.yami@kfupm.edu.sa
- Downloads last month
- 20