Instructions to use alefiury/G-DFlowTTS-NeuCodec-Emilia-YODAS with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use alefiury/G-DFlowTTS-NeuCodec-Emilia-YODAS with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="alefiury/G-DFlowTTS-NeuCodec-Emilia-YODAS", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("alefiury/G-DFlowTTS-NeuCodec-Emilia-YODAS", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Mask, Sample, Revise: A Revisable CTMC Inference Stack for Guided Discrete Flow Matching Text-to-Speech
Accepted at the IEEE Spoken Language Technology Workshop (SLT 2026).
G-DFlowTTS (Guided Discrete Flow Matching TTS) is a zero-shot, alignment-free text-to-speech model based on Discrete Flow Matching (DFM). A DiT with adaptive layer norm predicts NeuCodec audio codes conditioned on GPT-2 text tokens, and a Continuous-Time Markov Chain (CTMC) sampler infills them in parallel after an acoustic prompt. This checkpoint was trained on NeuCodec Emilia-YODAS (English).
In the paper we propose Mask, Sample, Revise, an inference-time CTMC stack for DFM-TTS that requires no post-hoc fine-tuning:
- SC-ReMask (Schedule-Constrained CTMC Remasking): a new remasking strategy for Discrete Flow Matching. It adapts the remasking formulation of ReMDM, originally designed for masked discrete diffusion, to the CTMC formulation of DFM. The two formulations are not directly compatible, so remasking is re-derived as an explicit token-to-mask CTMC transition whose rate is constrained by the probability-path schedule and added to the tau-leaping hazard. Generated tokens can return to the mask state and be revised, which yields good results with fewer inference steps.
- Discrete guidance: predictor-free guidance (PFG) mixes conditional and unconditional CTMC rates to strengthen text conditioning.
- Prompt-matched conditional coupling: training paths that match the prompted infilling task.
On LibriSpeech test-clean at 32 sampling steps, the full stack reduces WER from 75.44% (unguided baseline) to 8.39%, outperforming unguided and guidance-only samplers that use substantially more steps.
Model
| Architecture | DiT (adaLN), 12 blocks, hidden size 768, 12 heads |
| Parameters | 233M |
| Audio tokens | neuphonic/neucodec (50 Hz, 65,536 codes) |
| Text tokenizer | openai-community/gpt2 |
| Source distribution | Masked |
| Scheduler | Polynomial (n = 1) |
| Loss | Cross-entropy |
| Text dropout | 0.1 |
| Input sampling rate | 16,000 Hz (reference audio) |
| Output sampling rate | 24,000 Hz |
Training Dataset
| Dataset | Hours | Utterances | Language | License |
|---|---|---|---|---|
| NeuCodec Emilia-YODAS (English) | > 78k | > 30M | English | CC BY 4.0 |
The dataset contains the English subset of Emilia-YODAS pre-encoded with NeuCodec.
Usage
Load the model directly from the Hugging Face Hub with Transformers. No clone or manual snapshot download is required.
import soundfile as sf
from transformers import AutoModel
model = AutoModel.from_pretrained(
"alefiury/G-DFlowTTS-NeuCodec-Emilia-YODAS",
trust_remote_code=True,
).to("cuda")
ref_audio, sr = sf.read("reference.wav")
wav = model.synthesize(
text="I just received wonderful news about the promotion I have been waiting for!",
ref_audio=ref_audio,
ref_sampling_rate=sr,
ref_text="Kids are talking by the door.",
steps=32,
# Predictor-free guidance
use_pfg=True,
gamma=1.5,
# SC-ReMask: schedule-constrained CTMC remasking
use_sc_remask=True,
sc_remask_eta_rescale=0.5,
sc_remask_eta_cap=0.5,
sc_remask_tswitch=0.0,
seed=0,
)
sf.write("output.wav", wav.numpy(), model.config.sampling_rate)
The settings above are the best configuration reported in the paper (Mask, Sample, Revise): 32 sampling steps, PFG with γ = 1.5, and always-on SC-ReMask (t_switch = 0, η_rescale = η_cap = 0.5).
ref_text must be the exact transcription of the reference audio. Reference
audio is converted to mono and resampled to 16,000 Hz (resampling requires
torchaudio). The output length is estimated from the reference speaking rate;
use speed to adjust it or duration to set it explicitly in 50 Hz frames.
trust_remote_code=True is required because the G-DFlowTTS architecture and
sampler are shipped with this model repository. NeuCodec is loaded
automatically from its original Hugging Face repository.
Sampling options
Extra keyword arguments of synthesize are forwarded to model.sample_codes:
| Argument | Default | Description |
|---|---|---|
steps |
128 | Number of CTMC sampling steps |
speed |
1.0 | Speaking rate used by the length heuristic |
duration |
None |
Number of 50 Hz frames to generate (overrides speed) |
x1_temp, temp_schedule |
1.0, "dfm36" |
Sampling temperature and its schedule ("dfm36" or "constant") |
use_pfg, gamma |
False, 1.5 |
Prediction-free guidance |
use_tsr, tsr_k, tsr_sigma |
False, 1.0, 0.1 |
Temporal score rescaling |
use_sc_remask |
False |
SC-ReMask: schedule-constrained CTMC remasking (token-to-mask transitions that make generated tokens revisable) |
sc_remask_eta_rescale, sc_remask_eta_cap, sc_remask_tswitch |
0.5, 0.5, 0.0 | SC-ReMask schedule: σ = η_rescale · min(η_cap, σ_max), disabled for t < t_switch |
sc_remask_use_conf, sc_remask_conf_threshold, sc_remask_beta, sc_remask_strength |
False, 0.35, 2.0, 1.0 |
Optional confidence weighting: remask low-confidence tokens more |
The model was trained with text dropout, so prediction-free guidance (use_pfg=True) is supported.
The raw forward pass returns audio-code logits:
model(x_t, text_ids, time, drop_text=False, audio_att_mask=None).
License
Released under the cc-by-nc-4.0 license.
The training data (neuphonic/emilia-yodas-english-neucodec) is CC BY 4.0, and NeuCodec remains subject to its own license.
- Downloads last month
- 4