Mask, Sample, Revise: A Revisable CTMC Inference Stack for Guided Discrete Flow Matching Text-to-Speech

arXiv GitHub SLT 2026

Accepted at the IEEE Spoken Language Technology Workshop (SLT 2026).

G-DFlowTTS (Guided Discrete Flow Matching TTS) is a zero-shot, alignment-free text-to-speech model based on Discrete Flow Matching (DFM). A DiT with adaptive layer norm predicts NeuCodec audio codes conditioned on GPT-2 text tokens, and a Continuous-Time Markov Chain (CTMC) sampler infills them in parallel after an acoustic prompt. This checkpoint was trained on NeuCodec Emilia-YODAS (English).

In the paper we propose Mask, Sample, Revise, an inference-time CTMC stack for DFM-TTS that requires no post-hoc fine-tuning:

  • SC-ReMask (Schedule-Constrained CTMC Remasking): a new remasking strategy for Discrete Flow Matching. It adapts the remasking formulation of ReMDM, originally designed for masked discrete diffusion, to the CTMC formulation of DFM. The two formulations are not directly compatible, so remasking is re-derived as an explicit token-to-mask CTMC transition whose rate is constrained by the probability-path schedule and added to the tau-leaping hazard. Generated tokens can return to the mask state and be revised, which yields good results with fewer inference steps.
  • Discrete guidance: predictor-free guidance (PFG) mixes conditional and unconditional CTMC rates to strengthen text conditioning.
  • Prompt-matched conditional coupling: training paths that match the prompted infilling task.

On LibriSpeech test-clean at 32 sampling steps, the full stack reduces WER from 75.44% (unguided baseline) to 8.39%, outperforming unguided and guidance-only samplers that use substantially more steps.

Model

Architecture DiT (adaLN), 12 blocks, hidden size 768, 12 heads
Parameters 233M
Audio tokens neuphonic/neucodec (50 Hz, 65,536 codes)
Text tokenizer openai-community/gpt2
Source distribution Masked
Scheduler Polynomial (n = 1)
Loss Cross-entropy
Text dropout 0.1
Input sampling rate 16,000 Hz (reference audio)
Output sampling rate 24,000 Hz

Training Dataset

Dataset Hours Utterances Language License
NeuCodec Emilia-YODAS (English) > 78k > 30M English CC BY 4.0

The dataset contains the English subset of Emilia-YODAS pre-encoded with NeuCodec.

Usage

Load the model directly from the Hugging Face Hub with Transformers. No clone or manual snapshot download is required.

import soundfile as sf
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "alefiury/G-DFlowTTS-NeuCodec-Emilia-YODAS",
    trust_remote_code=True,
).to("cuda")

ref_audio, sr = sf.read("reference.wav")

wav = model.synthesize(
    text="I just received wonderful news about the promotion I have been waiting for!",
    ref_audio=ref_audio,
    ref_sampling_rate=sr,
    ref_text="Kids are talking by the door.",
    steps=32,
    # Predictor-free guidance
    use_pfg=True,
    gamma=1.5,
    # SC-ReMask: schedule-constrained CTMC remasking
    use_sc_remask=True,
    sc_remask_eta_rescale=0.5,
    sc_remask_eta_cap=0.5,
    sc_remask_tswitch=0.0,
    seed=0,
)

sf.write("output.wav", wav.numpy(), model.config.sampling_rate)

The settings above are the best configuration reported in the paper (Mask, Sample, Revise): 32 sampling steps, PFG with γ = 1.5, and always-on SC-ReMask (t_switch = 0, η_rescale = η_cap = 0.5).

ref_text must be the exact transcription of the reference audio. Reference audio is converted to mono and resampled to 16,000 Hz (resampling requires torchaudio). The output length is estimated from the reference speaking rate; use speed to adjust it or duration to set it explicitly in 50 Hz frames.

trust_remote_code=True is required because the G-DFlowTTS architecture and sampler are shipped with this model repository. NeuCodec is loaded automatically from its original Hugging Face repository.

Sampling options

Extra keyword arguments of synthesize are forwarded to model.sample_codes:

Argument Default Description
steps 128 Number of CTMC sampling steps
speed 1.0 Speaking rate used by the length heuristic
duration None Number of 50 Hz frames to generate (overrides speed)
x1_temp, temp_schedule 1.0, "dfm36" Sampling temperature and its schedule ("dfm36" or "constant")
use_pfg, gamma False, 1.5 Prediction-free guidance
use_tsr, tsr_k, tsr_sigma False, 1.0, 0.1 Temporal score rescaling
use_sc_remask False SC-ReMask: schedule-constrained CTMC remasking (token-to-mask transitions that make generated tokens revisable)
sc_remask_eta_rescale, sc_remask_eta_cap, sc_remask_tswitch 0.5, 0.5, 0.0 SC-ReMask schedule: σ = η_rescale · min(η_cap, σ_max), disabled for t < t_switch
sc_remask_use_conf, sc_remask_conf_threshold, sc_remask_beta, sc_remask_strength False, 0.35, 2.0, 1.0 Optional confidence weighting: remask low-confidence tokens more

The model was trained with text dropout, so prediction-free guidance (use_pfg=True) is supported.

The raw forward pass returns audio-code logits: model(x_t, text_ids, time, drop_text=False, audio_att_mask=None).

License

Released under the cc-by-nc-4.0 license. The training data (neuphonic/emilia-yodas-english-neucodec) is CC BY 4.0, and NeuCodec remains subject to its own license.

Downloads last month
4
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train alefiury/G-DFlowTTS-NeuCodec-Emilia-YODAS

Collection including alefiury/G-DFlowTTS-NeuCodec-Emilia-YODAS

Paper for alefiury/G-DFlowTTS-NeuCodec-Emilia-YODAS