Automatic speech recognition needs to reflect how people actually speak — including regional dialects and local recording conditions that are often underrepresented in pretraining data. NVIDIA provides a practical fine‑tuning recipe for Nemotron 3.5 ASR using the NeMo framework to adapt the model to Saudi Arabic dialects (Najdi and Hijazi).
Why this matters
A multilingual pretrained model can score well on broad benchmarks yet fail in real deployments where dialectal variation and recording conditions differ. Fine‑tuning on a target dialect can markedly improve accuracy for that dialect, but naive fine‑tuning risks degrading performance on other languages or domains unless steps are taken to preserve prior capabilities.
Model, data and environment
- Model: NVIDIA Nemotron 3.5 ASR (Cache-Aware FastConformer-RNNT streaming, multilingual prompt conditioning).
- Framework: NVIDIA NeMo, PyTorch, Python, OmegaConf.
- Datasets used: SADA 2022 (target dialect data: Najdi, Hijazi) and FLEURS (replay data in English and Arabic).
- Hardware for the baseline 12,000-step experiment: two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs.
When to use this workflow
This pipeline is suited for cases where you have enough labeled speech to specialize a pretrained ASR but not enough to train from scratch: dialekt adaptation, domain specific transcription, and deployments that must retain existing language skills.
Key techniques
- Minimal curation: remove clearly unusable labels and misaligned clips but keep scarce, difficult examples.
- Replay mixing: interleave a small, weighted stream of previously learned data to reduce catastrophic forgetting.
- Partial encoder unfreezing: restrict how much of the encoder is updated to save compute and memory when needed.
- Length bucketing: group utterances by duration to reduce padding and OOM risk for streaming encoders.
- Decoding optimizations: increase attention lookahead or use beam search at inference for offline accuracy gains without retraining (at the cost of latency and compute).
These are not universal defaults — replay protects only what it represents, and partial unfreezing requires retuning when the data mix changes.
Fine‑tuning walkthrough
- Curate the target corpus without throwing away the problem
- Select the dialects to deploy: SAUDI_DIALECTS = {"najdi", "hijazi"}.
- Drop items with non-learnable annotations (SADA marks inaudible speech with غيرواضح / غير واضح), extremely short or long clips, and transcripts with implausible character rate. Example thresholds: MIN_DURATION = 0.5s, MAX_DURATION = 30.0s, MIN_CHAR_RATE = 1.5, MAX_CHAR_RATE = 35.0.
- Normalize Arabic text (remove diacritics, unify alef variants, replace certain characters, strip punctuation). After these checks, 103,559 of 125,490 utterances were retained — 133.7 hours (82.5% of the starting set). The intent is to remove structurally bad examples, not dialectal or noisy samples that the base model handles poorly.
- The NeMo Curator pipeline (MonoConversionStage, UTMOSFilterStage, SIGMOSFilterStage) was run first in score-only mode to inspect distributions; cutoffs were set corpus‑specifically (UTMOS ≥ 1.25, SIGMOS noise ≥ 1.5, SIGMOS overall ≥ 1.5), passing about 85% of duration-valid samples.
- Start with one dataset and watch model behavior
- Begin with a single representative target dataset (SADA) and a conventional full fine‑tune to create an interpretable baseline and verify pipeline correctness.
- The pretrained checkpoint produced 49.5% WER on the validation split and 59% WER on the full training corpus. Validation WER improved modestly to 46.7% over continued runs and plateaued after ~45 epochs, suggesting diminishing returns from continuing the same full fine‑tune configuration.
- What worked: narrower target, replay stream, and bucketed batches
- Narrow the training target to only Najdi and Hijazi, with the minimal curation above (103,559 utt retained).
- Replay mix: 90% SADA Saudi speech, 7% FLEURS English, 3% FLEURS Arabic. Declare these weights in an input_cfg rather than concatenating manifests so the replay stream is present throughout training.
- Duration bucketing: enable use_bucketing=True, num_buckets=30, batch_duration=400.0 to group similar-length utterances and reduce padding.
- Combined effect (12,000 steps, ~4.5 hours on two GPUs): substantial WER/CER improvements (see results below).
- Adaptation depth: how much of the encoder to update
- The model has 24 encoder layers. Options tested: unfreeze top 6, top 8, or all 24 layers; decoder, joint network and prompt embeddings are always trained.
- In the top-8 recipe 230.4M params were trainable and 407.6M frozen.
- Results on Najdi+Hijazi test split:
- pretrained baseline WER: 55.05%, CER: 31.63%
- top 6 unfreeze: WER 33.42%, CER 14.10%
- top 8 unfreeze: WER 32.32%, CER 13.53%
- all 24 unfreeze: WER 29.96%, CER 12.18%
- More trainable capacity improved results for this data volume (~134 hours). Freezing remains a viable cost/memory tradeoff when data is scarcer.
- Improve inference without retraining: larger context and beam search
- The checkpoint exposes multiple attention lookahead sizes. Switching from [56,3] to [56,13] reduced WER by 1.31 absolute points without retraining, at the cost of ~800 ms additional latency.
- Beam search (MALSD) also yielded gains; e.g., MALSD beam‑4 with [56,3] reduced WER from 29.96% to 28.81% at ~0.59× greedy runtime; beam‑8 reached 28.62% at ~0.64×. The best tested combination (MALSD beam‑8, [56,13]) achieved WER 27.25%.
- Note: strip_lang_tags=True must be set so locale tags are not emitted as insertions.
- Latency examples: [56,0] ≈ 80 ms, [56,3] ≈ 320 ms, [56,13] ≈ 1.12 s.
Quantitative results (before → after fine‑tuning)
- SADA Najdi + Hijazi WER: 55.05% → 29.96%
- SADA Najdi + Hijazi CER: 31.63% → 12.18%
- Full SADA WER: 58.84% → 35.61%
- Full SADA CER: 35.40% → 15.97%
- FLEURS English WER (replay retention): 11.04% → 10.42%
- FLEURS Arabic WER: 12.67% → 11.41% (All evaluations used the NeMo evaluation script.)
Applying the workflow to other languages
The loop is the same: curate, mix replay, pick adaptation depth, bucket by duration, evaluate on independent target and regression sets. For other languages change the normalizer/tokenizer, verify Unicode handling, and choose appropriate metrics (character/token/morpheme) where WER is misleading. Replay data should reflect capabilities to retain and be real rather than synthetic when possible.
Extend with speaker diarization
Fine‑tuning improves transcription; diarization adds speaker labels and timestamps. NVIDIA Nemotron 3 Diarization (recently released) supports up to 8 speakers and can be combined with the ASR timestamps to produce speaker‑attributed transcripts useful for meetings, interviews and contact centers.
Recommendations and next steps
To make dialect adaptation practical: curate lightly, mix replay by weight, add the exact examples you need the system to handle, and update as much of the encoder as your data supports. Then optimize decoding and evaluate each capability on independent test sets before deployment.
Resources
- Finetuning notebook (NVIDIA Riva tutorials): asr-finetune-nemotron-3.5-asr-streaming-prompt.ipynb
- Finetuning skill (NVIDIA skills repo): nemotron-asr-finetune
(Reported numbers and configurations are from the described SADA-based experiments and should be adapted to your data and deployment constraints.)



