Lmaana v0
Lmaana v0 is a 1B-parameter-class research checkpoint for automatic speech recognition in Moroccan Darija. It combines a Dataset13-adapted OmniASR model with mixed replay over Dataset13 and MoulSot. The objective is to learn the MoulSot domain without losing the speech representations already acquired on Dataset13.
Model Summary
| Property | Value |
|---|---|
| Task | Automatic speech recognition |
| Language | Moroccan Darija (ary) |
| Architecture | OmniASR CTC 1b_v2 (wav2vec2_asr) |
| Framework | fairseq2 |
| Tokenization | Character tokenizer (omniASR_tokenizer_written_v2) |
| Decoding used for evaluation | Greedy CTC, no external language model |
| Checkpoint type | Native resumable training checkpoint |
Why Mixed Replay?
Adapting the Dataset13 model only on MoulSot improved MoulSot WER from 46.0435% to 43.9947%, but Dataset13 WER deteriorated from 44.1165% to 50.0524%. This is the catastrophic-forgetting problem: specialization on the new corpus damages performance on the original domain.
Lmaana v0 instead replays Dataset13 examples while learning from MoulSot. The result is a general-purpose checkpoint that improves MoulSot while preserving Dataset13 almost exactly.
Training Approach
- Starting point: OmniASR CTC
1b_v2after 6,000 Dataset13 adaptation steps. - Data construction: Dataset13 and MoulSot were exposed as one mixture-Parquet dataset. The local preparation uses hard links, avoiding a second physical copy of the audio payload.
- Replay sampling: corpus and language sampling coefficients were both set
to
0.5(beta_corpusandbeta_language). - Optimization: all encoder parameters remained trainable, with a
5e-7learning rate and gradient accumulation over 8 batches. - Selection: validation and checkpointing ran every 250 steps. The retained checkpoint is step 1,000, selected using validation WER.
Main Training Parameters
| Parameter | Value |
|---|---|
| Training steps | 1,000 |
| Learning rate | 5e-7 |
| Frozen encoder steps | 0 |
| Gradient accumulation | 8 batches |
| Maximum audio length | 320,000 elements |
| Maximum batch size | 1,280,000 elements |
| Example shuffle window | 1,000 |
| Validation interval | 250 steps |
| Checkpoint interval | 250 steps |
| Audio normalization | Enabled |
Evaluation Protocol
The selected checkpoint was evaluated separately on the held-out Dataset13 and MoulSot test partitions with the same fairseq2 evaluation pipeline. Reported WER and UER use greedy CTC decoding without an external language model. The test partitions were not used for checkpoint selection.
Final Test Results
| Test set | CTC loss | UER | WER | Evaluated examples |
|---|---|---|---|---|
| Dataset13 | 183.9310 | 17.7329% | 44.1211% | 6,576 |
| MoulSot | 76.4528 | 14.8204% | 45.4075% | 1,960 |
Comparison With Previous Checkpoints
| Checkpoint | Dataset13 WER | MoulSot WER | Interpretation |
|---|---|---|---|
| Dataset13 step 6,000 | 44.1165% | 46.0435% | Strong starting baseline |
| MoulSot-only step 1,000 | 50.0524% | 43.9947% | Better specialist, substantial forgetting |
| Lmaana v0 | 44.1211% | 45.4075% | Best general-purpose trade-off |
Against the starting checkpoint, Lmaana v0 changes Dataset13 WER by only
+0.0046 point while improving MoulSot WER by 0.6360 point. The MoulSot-only
checkpoint remains preferable only when MoulSot specialization matters more
than cross-domain retention.
Intended Use
- Research and evaluation of Moroccan Darija speech recognition.
- Continued fairseq2/OmniASR training from the native checkpoint.
- Development of Darija transcription systems with domain-specific testing.
This checkpoint has not been validated for high-stakes or fully automated decision-making.
Checkpoint Format
This repository stores a complete native fairseq2/OmniASR training checkpoint,
including optimizer and trainer state. It is not a Transformers
from_pretrained() export.
from huggingface_hub import snapshot_download
local_path = snapshot_download(repo_id="sailu4/lmaana-v0")
print(local_path)
Load the directory under checkpoint/ with the same OmniASR fairseq2 recipe.
The tokenizer, training configuration, evaluation summary, and detailed logs
are available under metadata/.
Reproducibility
The exact mixed-replay configuration and evaluation outputs are included in
metadata/. manifest.json records every uploaded file, its size, and its
SHA-256 digest. These records can be used to verify a restored checkpoint
before resuming training or running inference.
Repository Contents
checkpoint/: resumable native fairseq2 checkpoint.metadata/: tokenizer, configuration, metrics, and evaluation logs.manifest.json: file sizes and SHA-256 integrity hashes.
Limitations
Performance can vary with recording conditions, speaker demographics, code switching, regional vocabulary, and domains not represented in the evaluation sets. WER values should not be assumed to transfer unchanged to other Darija corpora. The model may produce omissions, substitutions, or plausible but incorrect text. Review transcriptions before using them in consequential workflows. Dataset and deployment licensing must be reviewed separately for the intended use.
Version
This repository contains Lmaana v0, the retained general-purpose baseline for subsequent Lmaana experiments.