techiaith
/

whisper-large-v3-ft-commonvoice-cy-en

Automatic Speech Recognition

Generated from Trainer

Model card Files Files and versions Metrics Training metrics

DewiBrynJones commited on Nov 6, 2024

Commit

cb11aa5

·

verified ·

1 Parent(s): 6466ac2

Update README.md

Files changed (1) hide show

README.md +16 -49

README.md CHANGED Viewed

@@ -9,61 +9,28 @@ metrics:
 model-index:
 - name: whisper-large-v3-ft-cv-cy-en
   results: []
 ---
-<!-- This model card has been generated automatically according to the information the Trainer had access to. You
-should probably proofread and complete it, then remove this comment. -->
 # whisper-large-v3-ft-cv-cy-en
-This model is a fine-tuned version of [openai/whisper-large-v3](https://huggingface.co/openai/whisper-large-v3) on the DewiBrynJones/commonvoice_18_0_cy_en train main dataset.
-It achieves the following results on the evaluation set:
-- Loss: 0.2744
-- Wer: 0.1474
-## Model description
-More information needed
-## Intended uses & limitations
-More information needed
-## Training and evaluation data
-More information needed
-## Training procedure
-### Training hyperparameters
-The following hyperparameters were used during training:
-- learning_rate: 1e-05
-- train_batch_size: 16
-- eval_batch_size: 16
-- seed: 42
-- gradient_accumulation_steps: 2
-- total_train_batch_size: 32
-- optimizer: Use adamw_torch with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
-- lr_scheduler_type: linear
-- lr_scheduler_warmup_steps: 500
-- training_steps: 5000
-- mixed_precision_training: Native AMP
-### Training results
-| Training Loss | Epoch  | Step | Validation Loss | Wer    |
-|:-------------:|:------:|:----:|:---------------:|:------:|
-| 0.4825        | 0.7075 | 1000 | 0.2708          | 0.1810 |
-| 0.2262        | 1.4149 | 2000 | 0.2486          | 0.1594 |
-| 0.0867        | 2.1224 | 3000 | 0.2506          | 0.1511 |
-| 0.0973        | 2.8299 | 4000 | 0.2444          | 0.1490 |
-| 0.0303        | 3.5373 | 5000 | 0.2744          | 0.1474 |
-### Framework versions
-- Transformers 4.46.1
-- Pytorch 2.5.1+cu124
-- Datasets 3.1.0
-- Tokenizers 0.20.1

 model-index:
 - name: whisper-large-v3-ft-cv-cy-en
   results: []
+datasets:
+- techiaith/commonvoice_18_0_cy_en
+language:
+- cy
+- en
+pipeline_tag: automatic-speech-recognition
 ---
 # whisper-large-v3-ft-cv-cy-en
+This model is a fine-tuned version of [openai/whisper-large-v3](https://huggingface.co/openai/whisper-large-v3) on the
+[techiaith/commonvoice_18_0_cy_en](https://huggingface.co/datasets/techiaith/commonvoice_18_0_cy_en) dataset. Both the
+English and Welsh data have been used to fine-tune the whisper model for transcribing both languages as well as improved
+language detection.
+It achieves a success rate of *98.86% for language detection* on recordings from a [Common Voice bilingual test set](https://huggingface.co/datasets/techiaith/commonvoice_18_0_cy_en/viewer/default/test)
+While, it achieves the following WER results for transcribing using the same test set:
+- Welsh: 26.20
+- English: 15.37
+- Average: 20.70
+N.B. the desired transcript language is not given to the fine-tuned model during testing.