MDLM (LM1B)

MDLM (masked diffusion language model) with a factorized reverse process and a dense DiT backbone. This is the factorized baseline for E-MoE, trained with the same data, tokenizer and budget.

This checkpoint accompanies the paper E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models. Code: github.com/Arseny5/e-moe.

Model

Architecture DiT, 12 blocks, hidden size 768, 12 heads
Trainable parameters 139.3M
Active parameters per token 139.3M
Sequence length 128 tokens
Tokenizer bert-base-uncased
Diffusion absorbing-state (masked) diffusion, log-linear noise schedule, no time conditioning
Weights EMA (decay 0.9999), float32, model.safetensors

Training data

LM1B (One Billion Word Benchmark), tokenized with bert-base-uncased and packed (wrapped) into blocks of 128 tokens.

Training

  • Steps: 1M, global batch 512 sequences (128 per GPU × 4 GPUs), ≈ 65.5B tokens (~72 epochs)
  • Optimizer: AdamW, lr 3e-4, betas (0.9, 0.999), weight decay 0, 2.5k warmup steps, gradient clipping 1.0
  • Precision: bf16 mixed precision; attention via PyTorch SDPA (FlashAttention kernels), rotary embeddings from flash-attn
  • Hardware: 4 × NVIDIA H200, ~80 h for 1M steps
  • Training objective: the standard MDLM continuous-time ELBO (SUBS parameterization).

Results

Generative perplexity (↓, judged by GPT-2 Large) / unigram sample entropy, unconditional generation of 128 tokens on LM1B. Mean ± std over 5 disjoint groups of 1,000 samples. Best gen-PPL per row in bold.

NFE MDLM (this model) SEDD VADD MDLM-MoE E-MoE
1 1440 ± 13 / 4.368 1604 ± 14 / 4.383 1279 ± 20 / 4.349 1293 ± 13 / 4.327 644 ± 6 / 4.352
2 993 ± 19 / 4.370 1065 ± 16 / 4.379 769 ± 11 / 4.334 816 ± 14 / 4.328 379 ± 5 / 4.349
4 469 ± 4 / 4.355 461 ± 2 / 4.347 372 ± 2 / 4.329 383 ± 2 / 4.323 235 ± 3 / 4.342
8 260 ± 1 / 4.351 242 ± 5 / 4.333 223.6 ± 5.5 / 4.326 214.7 ± 1.2 / 4.324 174.8 ± 2.1 / 4.340
16 179.4 ± 3.1 / 4.350 167.4 ± 3.5 / 4.329 165.7 ± 1.5 / 4.325 149.0 ± 3.1 / 4.322 148.4 ± 2.0 / 4.339
32 149.9 ± 1.9 / 4.349 138.0 ± 1.2 / 4.328 141.1 ± 2.5 / 4.324 125.1 ± 1.3 / 4.326 135.3 ± 1.8 / 4.338
64 133.7 ± 2.0 / 4.349 124.2 ± 1.5 / 4.326 126.6 ± 0.6 / 4.321 112.4 ± 1.1 / 4.325 130.1 ± 0.8 / 4.339
128 127.6 ± 1.3 / 4.349 118.9 ± 1.2 / 4.325 117.8 ± 0.6 / 4.317 107.0 ± 1.4 / 4.323 124.9 ± 1.6 / 4.337

Autoregressive baseline (same backbone and budget, 128 steps): 68.3 ± 0.9 / 4.320.

Usage

The checkpoint contains the backbone weights (model.safetensors) and the architecture/training configuration (config.json):

import json
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file

repo = "ArsenyIvanov/mdlm-lm1b"
state_dict = load_file(hf_hub_download(repo, "model.safetensors"))
config = json.load(open(hf_hub_download(repo, "config.json")))

The model code and sampling scripts that load these weights will be released in github.com/Arseny5/e-moe.

Training checkpoint

training_checkpoint/step_1000000.ckpt is the full PyTorch Lightning checkpoint at step 1M: raw and EMA weights, AdamW optimizer state, learning-rate scheduler and loop state, for resuming training. It is a pickle file: load it only from this trusted repository, e.g. torch.load(path, map_location="cpu", weights_only=False). For inference, use model.safetensors (EMA weights).

License

Apache License 2.0.

Citation

@article{ivanov2026emoe,
  title   = {E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models},
  author  = {Ivanov, Arseny and Kolesov, Alexander and Korotin, Alexander and
             Oseledets, Ivan and Goncharov, Mikhail},
  journal = {arXiv preprint arXiv:2609.37533},
  year    = {2026}
}
Downloads last month
7
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including ArsenyIvanov/mdlm-lm1b

Paper for ArsenyIvanov/mdlm-lm1b