Lmaana v0

Lmaana v0 is a 1B-parameter-class research checkpoint for automatic speech recognition in Moroccan Darija. It combines a Dataset13-adapted OmniASR model with mixed replay over Dataset13 and MoulSot. The objective is to learn the MoulSot domain without losing the speech representations already acquired on Dataset13.

Model Summary

Property Value
Task Automatic speech recognition
Language Moroccan Darija (ary)
Architecture OmniASR CTC 1b_v2 (wav2vec2_asr)
Framework fairseq2
Tokenization Character tokenizer (omniASR_tokenizer_written_v2)
Decoding used for evaluation Greedy CTC, no external language model
Checkpoint type Native resumable training checkpoint

Why Mixed Replay?

Adapting the Dataset13 model only on MoulSot improved MoulSot WER from 46.0435% to 43.9947%, but Dataset13 WER deteriorated from 44.1165% to 50.0524%. This is the catastrophic-forgetting problem: specialization on the new corpus damages performance on the original domain.

Lmaana v0 instead replays Dataset13 examples while learning from MoulSot. The result is a general-purpose checkpoint that improves MoulSot while preserving Dataset13 almost exactly.

Training Approach

  1. Starting point: OmniASR CTC 1b_v2 after 6,000 Dataset13 adaptation steps.
  2. Data construction: Dataset13 and MoulSot were exposed as one mixture-Parquet dataset. The local preparation uses hard links, avoiding a second physical copy of the audio payload.
  3. Replay sampling: corpus and language sampling coefficients were both set to 0.5 (beta_corpus and beta_language).
  4. Optimization: all encoder parameters remained trainable, with a 5e-7 learning rate and gradient accumulation over 8 batches.
  5. Selection: validation and checkpointing ran every 250 steps. The retained checkpoint is step 1,000, selected using validation WER.

Main Training Parameters

Parameter Value
Training steps 1,000
Learning rate 5e-7
Frozen encoder steps 0
Gradient accumulation 8 batches
Maximum audio length 320,000 elements
Maximum batch size 1,280,000 elements
Example shuffle window 1,000
Validation interval 250 steps
Checkpoint interval 250 steps
Audio normalization Enabled

Evaluation Protocol

The selected checkpoint was evaluated separately on the held-out Dataset13 and MoulSot test partitions with the same fairseq2 evaluation pipeline. Reported WER and UER use greedy CTC decoding without an external language model. The test partitions were not used for checkpoint selection.

Final Test Results

Test set CTC loss UER WER Evaluated examples
Dataset13 183.9310 17.7329% 44.1211% 6,576
MoulSot 76.4528 14.8204% 45.4075% 1,960

Comparison With Previous Checkpoints

Checkpoint Dataset13 WER MoulSot WER Interpretation
Dataset13 step 6,000 44.1165% 46.0435% Strong starting baseline
MoulSot-only step 1,000 50.0524% 43.9947% Better specialist, substantial forgetting
Lmaana v0 44.1211% 45.4075% Best general-purpose trade-off

Against the starting checkpoint, Lmaana v0 changes Dataset13 WER by only +0.0046 point while improving MoulSot WER by 0.6360 point. The MoulSot-only checkpoint remains preferable only when MoulSot specialization matters more than cross-domain retention.

Intended Use

  • Research and evaluation of Moroccan Darija speech recognition.
  • Continued fairseq2/OmniASR training from the native checkpoint.
  • Development of Darija transcription systems with domain-specific testing.

This checkpoint has not been validated for high-stakes or fully automated decision-making.

Checkpoint Format

This repository stores a complete native fairseq2/OmniASR training checkpoint, including optimizer and trainer state. It is not a Transformers from_pretrained() export.

from huggingface_hub import snapshot_download

local_path = snapshot_download(repo_id="sailu4/lmaana-v0")
print(local_path)

Load the directory under checkpoint/ with the same OmniASR fairseq2 recipe. The tokenizer, training configuration, evaluation summary, and detailed logs are available under metadata/.

Reproducibility

The exact mixed-replay configuration and evaluation outputs are included in metadata/. manifest.json records every uploaded file, its size, and its SHA-256 digest. These records can be used to verify a restored checkpoint before resuming training or running inference.

Repository Contents

  • checkpoint/: resumable native fairseq2 checkpoint.
  • metadata/: tokenizer, configuration, metrics, and evaluation logs.
  • manifest.json: file sizes and SHA-256 integrity hashes.

Limitations

Performance can vary with recording conditions, speaker demographics, code switching, regional vocabulary, and domains not represented in the evaluation sets. WER values should not be assumed to transfer unchanged to other Darija corpora. The model may produce omissions, substitutions, or plausible but incorrect text. Review transcriptions before using them in consequential workflows. Dataset and deployment licensing must be reviewed separately for the intended use.

Version

This repository contains Lmaana v0, the retained general-purpose baseline for subsequent Lmaana experiments.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support