Instructions to use aether-raid/astra-atc-models with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use aether-raid/astra-atc-models with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("aether-raid/astra-atc-models") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
Download par/ASR/README.md from aether-raid/astra-atc-models: direct link, hf CLI and curl.
- Browser
- Download file 3.13 kB
-
https://huggingface.co/aether-raid/astra-atc-models/resolve/main/par/ASR/README.md
- Command line
-
hf download hf://aether-raid/astra-atc-models/par/ASR/README.md
-
curl -L -o README.md https://huggingface.co/aether-raid/astra-atc-models/resolve/main/par/ASR/README.md
language:
- en
license: other
library_name: nemo
base_model: nvidia/parakeet-tdt-0.6b-v2
pipeline_tag: automatic-speech-recognition
tags:
- nemo
- parakeet
- tdt
- automatic-speech-recognition
- air-traffic-control
- par
- singapore
- military
metrics:
- wer
Parakeet-TDT 0.6B v2 - PAR Talkdown
NVIDIA Parakeet-TDT 0.6B v2 fine-tuned for Precision Approach Radar (PAR) talkdown speech: fast, clipped controller calls and pilot replies, as used by the PAR training simulator.
Usage
model.nemo is self-contained (weights, config and tokenizer). Audio is 16 kHz mono.
from nemo.collections.asr.models import EncDecRNNTBPEModel
model = EncDecRNNTBPEModel.restore_from("par/ASR/model.nemo")
model.eval().cuda()
print(model.transcribe(["audio.wav"])[0].text)
Tested with nemo_toolkit[asr]==2.7.3.
Output Format
Lowercase spoken words with no punctuation, for a downstream phrase matcher. Numbers, callsign numbers and
abbreviations are spoken digit by digit and letter by letter. ICAO "tree" and "fife" come out as three and five,
"niner" and "nine" are kept as spoken, and hesitations come out as uh or um.
Results
WER on held-out test sets, best epoch (14 of 15). Every epoch is in training/results.txt.
| Test set | Clips | WER |
|---|---|---|
| PAR, held-out phrases (half sped up with radio noise) | 3,390 | 0.17% |
| PAR, real recordings | 20 | 0.00% |
| General English (Loquacious test) | 461 | 8.43% |
| Singapore English (MNSC) | 267 | 0.95% |
| Real ATC radio (DeepDML) | 142 | 4.31% |
The PAR held-out phrases are synthetic and split by sentence, but share templates with training phrases, and the 20 real recordings are one speaker whose other recordings were trained on. New speakers are the real test. The model was not trained on ASTRA phraseology and outputs PAR callsigns for unfamiliar ones, so use the ASTRA ASR for ASTRA. Gate it with VAD or push-to-talk, since noise-only audio can still produce words.
Training
Fine-tuned from nvidia/parakeet-tdt-0.6b-v2 on 79,242 clips (93.7 h). 10% of PAR phrases were held out by
sentence across all voices for the PAR test set; the other test sets come from separate splits of each source.
| Source | Hours |
|---|---|
| PAR synthetic speech (six TTS voices) | 51.4 |
| ASTRA PAR-like synthetic speech and real recordings | 9.3 |
| General English replay (Loquacious: audiobooks, VoxPopuli, Common Voice, YODAS) | 22.5 |
| Singapore English (MNSC) | 7.6 |
| Real ATC radio (DeepDML) | 2.6 |
| Real PAR recordings | 0.3 |
Augmentation: speed copies of PAR speech from 0.8x to 2.5x (most at 1.3x to 2.0x, capped at 7 words/s); radio band-pass, drive, noise, codec and push-to-talk clicks; and on about 4% of clips one light effect (mic hiss, clipping, fade, one distant talker or quiet babble).
AdamW at 1e-4 with one epoch of warmup and cosine decay to 1e-6, weight decay 1e-3, batch 64, bf16, 15 epochs, best checkpoint by PAR WER. About 10.5 h on one RTX 4090. Loss and WER per step and epoch are in training/train.log.