astra-atc-models / par /ASR /README.md
RanenSim's picture
feat: add par asr
98f8619
|
Raw History Blame Contribute Delete
3.13 kB
metadata
language:
  - en
license: other
library_name: nemo
base_model: nvidia/parakeet-tdt-0.6b-v2
pipeline_tag: automatic-speech-recognition
tags:
  - nemo
  - parakeet
  - tdt
  - automatic-speech-recognition
  - air-traffic-control
  - par
  - singapore
  - military
metrics:
  - wer

Parakeet-TDT 0.6B v2 - PAR Talkdown

NVIDIA Parakeet-TDT 0.6B v2 fine-tuned for Precision Approach Radar (PAR) talkdown speech: fast, clipped controller calls and pilot replies, as used by the PAR training simulator.

Usage

model.nemo is self-contained (weights, config and tokenizer). Audio is 16 kHz mono.

from nemo.collections.asr.models import EncDecRNNTBPEModel

model = EncDecRNNTBPEModel.restore_from("par/ASR/model.nemo")
model.eval().cuda()
print(model.transcribe(["audio.wav"])[0].text)

Tested with nemo_toolkit[asr]==2.7.3.

Output Format

Lowercase spoken words with no punctuation, for a downstream phrase matcher. Numbers, callsign numbers and abbreviations are spoken digit by digit and letter by letter. ICAO "tree" and "fife" come out as three and five, "niner" and "nine" are kept as spoken, and hesitations come out as uh or um.

Results

WER on held-out test sets, best epoch (14 of 15). Every epoch is in training/results.txt.

Test set Clips WER
PAR, held-out phrases (half sped up with radio noise) 3,390 0.17%
PAR, real recordings 20 0.00%
General English (Loquacious test) 461 8.43%
Singapore English (MNSC) 267 0.95%
Real ATC radio (DeepDML) 142 4.31%

The PAR held-out phrases are synthetic and split by sentence, but share templates with training phrases, and the 20 real recordings are one speaker whose other recordings were trained on. New speakers are the real test. The model was not trained on ASTRA phraseology and outputs PAR callsigns for unfamiliar ones, so use the ASTRA ASR for ASTRA. Gate it with VAD or push-to-talk, since noise-only audio can still produce words.

Training

Fine-tuned from nvidia/parakeet-tdt-0.6b-v2 on 79,242 clips (93.7 h). 10% of PAR phrases were held out by sentence across all voices for the PAR test set; the other test sets come from separate splits of each source.

Source Hours
PAR synthetic speech (six TTS voices) 51.4
ASTRA PAR-like synthetic speech and real recordings 9.3
General English replay (Loquacious: audiobooks, VoxPopuli, Common Voice, YODAS) 22.5
Singapore English (MNSC) 7.6
Real ATC radio (DeepDML) 2.6
Real PAR recordings 0.3

Augmentation: speed copies of PAR speech from 0.8x to 2.5x (most at 1.3x to 2.0x, capped at 7 words/s); radio band-pass, drive, noise, codec and push-to-talk clicks; and on about 4% of clips one light effect (mic hiss, clipping, fade, one distant talker or quiet babble).

AdamW at 1e-4 with one epoch of warmup and cosine decay to 1e-6, weight decay 1e-3, batch 64, bf16, 15 epochs, best checkpoint by PAR WER. About 10.5 h on one RTX 4090. Loss and WER per step and epoch are in training/train.log.