Instructions to use AEON-7/AEON-DFlash2-Qwen3.8-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AEON-7/AEON-DFlash2-Qwen3.8-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AEON-7/AEON-DFlash2-Qwen3.8-27B")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("AEON-7/AEON-DFlash2-Qwen3.8-27B") model = AutoModel.from_pretrained("AEON-7/AEON-DFlash2-Qwen3.8-27B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AEON-7/AEON-DFlash2-Qwen3.8-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AEON-7/AEON-DFlash2-Qwen3.8-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/AEON-DFlash2-Qwen3.8-27B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/AEON-7/AEON-DFlash2-Qwen3.8-27B
- SGLang
How to use AEON-7/AEON-DFlash2-Qwen3.8-27B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AEON-7/AEON-DFlash2-Qwen3.8-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/AEON-DFlash2-Qwen3.8-27B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AEON-7/AEON-DFlash2-Qwen3.8-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/AEON-DFlash2-Qwen3.8-27B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use AEON-7/AEON-DFlash2-Qwen3.8-27B with Docker Model Runner:
docker model run hf.co/AEON-7/AEON-DFlash2-Qwen3.8-27B
Request access to AEON DFlash 2
Early-access drafter. Requests are reviewed every few hours; supporters include the community access word for automatic approval.
AEON DFlash 2 v1.0 is an early release. Requests are reviewed every few hours. Supporters on Patreon: enter the community access word from the pinned post and you are approved automatically on the next check. Everyone else: type none and your request is reviewed by hand. The model is Apache-2.0 licensed; keep the LICENSE and NOTICE files and credit AEON-7 and z-lab / Inco AI when you share it or anything built from it.
Log in or Sign Up to review the conditions and access this model content.
- AEON DFlash 2 for Qwen3.8-27B AEON Ultimate
- Access
- At a glance
- Which weights do I use?
- Quickstart by architecture
- Speed: measured on one DGX Spark (lk3, the previous stage)
- Acceptance, this release (lk4 = S5 step 753)
- Notes and caveats
- More tests and data coming
- Quickstart: one DGX Spark
- Use it with any inference engine
- Where it runs
- What's in this repo
- What it is
- Data
- Support the work
- Attribution & sharing
- License
- Access
AEON DFlash 2 for Qwen3.8-27B AEON Ultimate
v1.0 (lk4) · released 2026-10-05 · Early Release · Gated Access
Measured with the previous stage (lk3): +21% faster than the stock DFlash 2 drafter at T=0.6, +35% at T=1.0, 3.24× faster than plain decoding for one user.† Lossless by construction.
Measured on one NVIDIA DGX Spark with the published AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4-MIXED build, against stock z-lab/Qwen3.8-27B-DFlash2 (BF16).
† The "vs plain decoding" multiples are unpaired: AEON ran 537 prompts at T=0.6, the no-speculation baseline 72 prompts (4 per category, about 10.1 tok/s in every category). "Lossless" holds by construction: rejection sampling has the target verify every drafted token. Our small accuracy check is not proof of it; that check was dominated by 512-token output caps on the no-speculation baseline (see Notes and caveats).
Read this first. The served speed numbers on this page were measured with lk3, the previous training stage. This release (lk4) accepts +0.7% more drafted tokens than lk3 offline ([+0.47, +1.00], 95% CI), so expect equal or slightly better speed. lk4's own served numbers are coming: a full category × concurrency benchmark is next on the DGX Spark. More extensive tests and data to come.
Weights update, 2026-10-05. Within the first hours of release the v1.0 weights were updated from training step 690 to step 753 (lk4, chosen by the release rule's tie-break). Offline acceptance is identical (+0.70% vs lk3, +4.87% vs stock z-lab). If you downloaded earlier, re-download; the new sha256 values are under What's in this repo.
Built on DFlash 2 by Inco AI & z-lab.
AEON DFlash 2 drafts up to nine tokens in one parallel pass. The 27B model checks them all at once, keeps what it agrees with, and moves on. It still decides every token. You just get them sooner.
It is a speculative-decoding drafter, not a standalone chat model. It runs next to the target it was trained for,
AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4-MIXED,
and it was fine-tuned from z-lab/Qwen3.8-27B-DFlash2
(a mirror of incoai/Qwen3.8-27B-DFlash2, Apache-2.0) on how that target is actually served.
Two precisions, two folders. The repo root holds the drafter as an NVFP4 W4A16 pack (4-bit MLP weights, 16-bit activations, 1.93 GB) for NVIDIA Blackwell with vLLM.
The BF16 weights (3.85 GB) live only in the bf16/ folder, for every other architecture and engine. See Which weights do I use? and Quickstart by architecture.
Access
Request access on the model page. Requests are reviewed every few hours. Supporters: include the community access word from the pinned post for automatic approval.
Initial gated access for Patreon supporters: early-access post · become a member.
At a glance
| +21.1% | tok/s over stock z-lab DFlash 2 at T=0.6, 1 user [+19.6, +22.5] (lk3, measured) |
| +34.6% | tok/s over stock z-lab at T=1.0, the Qwen / z-lab sampling protocol [+31.5, +37.7] (lk3, measured) |
| 3.24× | faster than plain decoding at 1 user, T=0.6: 32.7 vs 10.1 tok/s (lk3, measured; unpaired†) |
| 217 tok/s | total output at 16 concurrent users on one DGX Spark, vs 181 stock z-lab and 112 plain (lk3, measured) |
| +4.87% | offline acceptance over stock z-lab BF16 for this release [+4.38, +5.51] |
| 18 / 18 | benchmark categories faster than stock z-lab at T=0.6, every confidence interval above zero |
| 1.9 GB | NVFP4 W4A16 drafter in the repo root, for NVIDIA Blackwell; half the 3.85 GB BF16 copy, which lives only in bf16/ (for every other architecture) |
| Lossless by construction | rejection sampling has the target verify every token, so the output distribution is the target's own. A property of the method, not a measurement |
Why it's great
- Trained on the model you actually serve. It learned the published NVFP4-MIXED target's real served output distribution at its serving sampling settings (T=0.6, top-p 0.95, top-k 20), thinking and tool calls included, with an objective that directly maximizes how many drafted tokens the target keeps.
- Native NVFP4 (repo root). Trained with NVFP4 quantization-aware training on the exact grid it serves on. The published W4A16 pack in the repo root matches the training grid bit for bit (export check: every FP4 code and scale equal).
bf16/holds the same values expanded to BF16. - Built for 9 drafts per step (block 10), where z-lab's published default is 7.
- Tuned for DGX Spark. A per-concurrency draft-length lattice plus probabilistic drafting.
- Drop-in. The
bf16/copy has the same 81 tensors (names, shapes, dtypes) asz-lab/Qwen3.8-27B-DFlash2, so it loads wherever that drafter runs.
Recommended configuration (the measured production config)
| Setting | Value | Why |
|---|---|---|
| Drafter folder | Repo root (NVFP4 W4A16) on NVIDIA Blackwell with vLLM; bf16/ everywhere else |
Which weights do I use? |
| Draft tokens (K) | 9, never above 9 |
The training block is 10 (1 + 9) |
| Per-concurrency lattice | [[1,2,9],[3,8,7],[9,12,6],[13,16,4]] |
Lowers K as more requests share a step (docs/HOSTING.md) |
| Draft attention backend | TRITON_ATTN (inside --speculative-config) |
The top-level flag does not always reach the drafter |
| Draft sampling | probabilistic |
Lossless by construction; the measured setting (vLLM defaults to greedy) |
| vLLM model runner | VLLM_USE_V2_MODEL_RUNNER=1 |
DFlash 2 exists only in the V2 runner; on V1 it silently drafts as DFlash 1 |
| Speculative methods | DFlash only | Don't combine with MTP |
| Prefix caching | Stock upstream vLLM with hybrid-GDN targets, every release so far (through 0.31): off (--no-enable-prefix-caching) |
vllm#58894, fix pending in vllm#55601. The DGX Spark command below matches the measured config on the AEON image (exceptions in Quickstart step 2) |
Which weights do I use?
This repo holds one drafter in two precisions. Download only the one your hardware needs.
| Folder | Precision | Size | Use it on | Pair it with |
|---|---|---|---|---|
Repo root (model.safetensors) |
NVFP4 W4A16: 4-bit FP4 weights in the 15 MLP linears, every other tensor BF16, activations BF16 | 1.93 GB | NVIDIA Blackwell with vLLM. Tested: DGX Spark (GB10) on the AEON vLLM 0.29 image. Expected but untested: RTX 50-series, RTX PRO 6000, B200 / B300 | …-NVFP4-MIXED target |
bf16/ (bf16/model.safetensors) |
BF16, every tensor | 3.85 GB | Every other architecture and engine: NVIDIA Hopper / Ampere / Ada, AMD ROCm, Apple Silicon, and every engine other than vLLM | …-BF16 target on NVIDIA; …-MXFP4-MXFP6-ROCm target on AMD |
- The BF16 weights live only in
bf16/. The rootmodel.safetensorsis the NVFP4 W4A16 pack, not BF16. Both folders hold the same trained drafter:bf16/is the root pack expanded to BF16, with 0 differing elements. - What W4A16 means here. The 15 MLP linears' weights are stored in 4 bits and expanded on the fly by vLLM's Marlin kernel (FlashInfer CuTe-DSL on B200 / B300 in vLLM ≥ 0.30). The math runs with BF16 activations, not on FP4 tensor cores. The pack halves the drafter's bytes, which helps on bandwidth-bound Blackwell boxes such as the DGX Spark. At load vLLM may log "Your GPU does not have native support for FP4 computation…"; that line is generic and expected, on Blackwell too.
- Older NVIDIA GPUs. On Ampere, Ada and Hopper (A100, RTX 30 / 40, L40S, H100 / H200), vLLM can also load the root pack through Marlin W4A16 (checked against vLLM's source; untested). Use
bf16/there: it is the path documented below. Turing (T4) is not supported, because the drafter needs BF16. - AMD ROCm: vLLM cannot load the root pack (its NVFP4 Marlin kernel is CUDA-only). Use
bf16/.
Which files are where:
AEON-7/AEON-DFlash2-Qwen3.8-27B
├── model.safetensors NVFP4 W4A16 drafter, 1.93 GB ← NVIDIA Blackwell + vLLM
├── config.json with quantization_config (modelopt_mixed, W4A16_NVFP4)
├── hf_quant_config.json NVFP4 pack description
├── FIDELITY_GATE.json pack checks
├── LOAD_SHAPE_AUDIT.json
├── bf16/
│ ├── model.safetensors BF16 drafter, 3.85 GB ← everything else
│ ├── config.json same config, without quantization_config
│ ├── DEQUANT_REPORT.json NVFP4 → BF16 exactness report
│ └── README.md, LICENSE, NOTICE
├── configs/ scripts/ recipes/ docs/ images/
└── README.md, AGENTS.md, CHANGELOG.md, CREDITS.md, EVOLUTION.md, LICENSE, NOTICE
Download only the folder you need. The files are gated: request access first, then hf auth login.
# NVIDIA Blackwell: repo root, NVFP4 W4A16 (~1.9 GB). Drafter dir = ~/models/aeon-dflash2
hf download AEON-7/AEON-DFlash2-Qwen3.8-27B --local-dir ~/models/aeon-dflash2 --exclude "bf16/*"
# Everything else: bf16/ only (~3.85 GB). Drafter dir = ~/models/aeon-dflash2/bf16
hf download AEON-7/AEON-DFlash2-Qwen3.8-27B --local-dir ~/models/aeon-dflash2 --include "bf16/*"
With --include "bf16/*" the weights land in ~/models/aeon-dflash2/bf16/. Point your engine at that subfolder, not at ~/models/aeon-dflash2, which then holds no weights.
Quickstart by architecture
| Your hardware | Target model | Drafter folder | Engine | Status |
|---|---|---|---|---|
| A. NVIDIA Blackwell (DGX Spark / GB10, RTX 50-series, RTX PRO 6000, B200) | NVFP4-MIXED | repo root (NVFP4 W4A16) | AEON vLLM image | DGX Spark: tested. Others: untested |
| B. NVIDIA Hopper / Ampere / Ada (H100, H200, A100, L40S, RTX 30 / 40) | BF16 | bf16/ |
stock vLLM ≥ 0.28.0 | Untested |
| C. AMD ROCm (Radeon AI PRO R9700 and the other GPUs the ROCm kit covers) | MXFP4-MXFP6-ROCm | bf16/ |
vLLM 0.31.0 ROCm, from the target's kit | Untested on AMD with this release |
| D. Other engines (SGLang, llama.cpp, TensorRT-LLM, TensorFold, MLX) | depends on the engine | bf16/ |
see the engine table | Untested with AEON targets |
On every vLLM path: VLLM_USE_V2_MODEL_RUNNER=1, K=9 and never above, the per-concurrency lattice, TRITON_ATTN inside --speculative-config, probabilistic draft sampling, and no MTP at the same time.
A. NVIDIA Blackwell: NVFP4 W4A16 (root)
DGX Spark (GB10): tested.
hf download AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4-MIXED --local-dir ~/models/aeon-27b
hf download AEON-7/AEON-DFlash2-Qwen3.8-27B --local-dir ~/models/aeon-dflash2 --exclude "bf16/*"
docker pull ghcr.io/aeon-7/aeon-vllm-ultimate:2026-09-18-v0.29.0-omni
Serve with the full command in Quickstart: one DGX Spark, step 2. It mounts the repo root (~/models/aeon-dflash2, the NVFP4 W4A16 pack) as /draft. Weights total about 26.6 GB (24.7 GB target + 1.93 GB drafter).
curl -sf http://localhost:8000/health && echo OK
docker logs aeon-dflash2 2>&1 | grep -E "DFlash2|MarlinNvFp4LinearKernel" | head # drafter loaded on the W4A16 (Marlin) path
curl -s http://localhost:8000/metrics | grep -E '^vllm:spec_decode_num_(drafts|accepted_tokens)_total'
# τ = 1 + accepted ÷ drafts. Measured (lk3): about 4.1 at 1 user, T=0.6
RTX 5090 / RTX PRO 6000 (sm_120): untested, being validated (recipes/rtx-5090.md). The Spark image is arm64 and won't run there. The NVFP4-MIXED card's RTX image is ghcr.io/aeon-7/aeon-vllm-ultimate-rtx:latest; it has not yet been checked for DFlash 2 support, which needs vLLM 0.28.0 or newer. Use the same root download, /draft mount, environment and --speculative-config as the Spark command, with the memory settings from the target card's RTX seats. On a 32 GB RTX 5090 the weights alone take about 26.6 GB, so little is left for the KV cache; the target card's 5090 seat uses MTP instead.
B200 / B300 (sm_100 / sm_103): untested. No AEON image is published for them. Stock vLLM ≥ 0.30 runs the root pack's MLPs on FlashInfer CuTe-DSL W4A16 when FlashInfer provides that kernel (otherwise Marlin). Serving the NVFP4-MIXED target on stock vLLM may also need the target's ModelOpt patch (modelopt-54367.py, see its card).
B. NVIDIA Hopper / Ampere / Ada: bf16/ (untested)
For H100, H200, A100, L40S, RTX 6000 Ada, RTX 30 / 40 and other sm_80+ GPUs. The NVFP4-MIXED target is a Blackwell build, so pair bf16/ with the BF16 target. Use stock vLLM 0.28.0 or newer: DFlash 2 first shipped in 0.28.0, and the latest release is 0.31.0.
pip install -U huggingface_hub "vllm>=0.28.0"
hf auth login
hf download AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 --local-dir ~/models/aeon-27b-bf16
hf download AEON-7/AEON-DFlash2-Qwen3.8-27B --local-dir ~/models/aeon-dflash2 --include "bf16/*"
DRAFT="$HOME/models/aeon-dflash2/bf16" # the BF16 drafter: the bf16/ subfolder, not the root
VLLM_USE_V2_MODEL_RUNNER=1 vllm serve ~/models/aeon-27b-bf16 \
--served-model-name aeon --port 8000 \
--dtype bfloat16 --max-model-len 65536 --gpu-memory-utilization 0.90 \
--no-enable-prefix-caching \
--reasoning-parser qwen3 --tool-call-parser qwen3_coder --enable-auto-tool-choice \
--trust-remote-code \
--speculative-config '{"method":"dflash","model":"'"$DRAFT"'","num_speculative_tokens":9,"num_speculative_tokens_per_batch_size":[[1,2,9],[3,8,7],[9,12,6],[13,16,4]],"attention_backend":"TRITON_ATTN","draft_sample_method":"probabilistic"}' \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
# Two GPUs (for example 2x 48 GB): add --tensor-parallel-size 2
- Untested. This exact pairing (stock vLLM, BF16 target,
bf16/drafter) has not been run yet; stock-vLLM validation is tracked in recipes/vllm-upstream.md. Acceptance with the BF16 target is not measured: the drafter was trained on the NVFP4-MIXED target's outputs. - Prefix caching off: every vLLM release so far (through 0.31) needs
--no-enable-prefix-cachingwith this hybrid Gated-DeltaNet target (vllm#58894; fix pending in vllm#55601). - Don't add MTP. The BF16 target card's own command uses
"method":"mtp"; use one speculative method at a time. - Memory (weights only, estimates): 55.6 GB target + 3.85 GB drafter ≈ 59.4 GB. H200 (141 GB): comfortable. H100 / A100 80 GB: fits on one GPU with roughly 10 GB left for the KV cache at 0.90, so keep the context modest (65,536 above is a conservative suggestion, not a measurement). 48 GB cards: tensor parallel 2, about 29.7 GB of weights per GPU. The target's KV cache takes about 65.5 KB per token in BF16, plus a fixed Gated-DeltaNet state per running request.
- Add
--api-key <secret>if other people can reach the port. With data parallel above 1, vLLM ignores the lattice and uses a fixed K=9.
curl -sf http://localhost:8000/health && echo OK
curl -s http://localhost:8000/metrics | grep -E '^vllm:spec_decode_num_(drafts|accepted_tokens)_total'
# τ = 1 + accepted ÷ drafts. Well below 3 on normal prompts: check that "model" points at .../bf16
# and that VLLM_USE_V2_MODEL_RUNNER=1 is set (the startup log should name DFlash2DraftModel)
C. AMD ROCm: bf16/ (untested on AMD)
Use the kit that ships with the ROCm target, AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm (gated). Its card has the full setup, the pinned vLLM 0.31.0 ROCm image and a recipe per GPU. Its Recipe A (2x Radeon AI PRO R9700, TP=2) already runs DFlash 2 with the right settings: K=9 plus the lattice, TRITON_ATTN inside and outside the speculative config, the V2 runner and prefix caching off. The kit bundles an earlier BF16 drafter (lk3) in its dflash2/ folder. To run this release (lk4), download bf16/ from this repo and point the kit at it with AEON_DFLASH:
hf auth login
hf download AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm --local-dir ~/models/aeon-mxfp4-rocm
hf download AEON-7/AEON-DFlash2-Qwen3.8-27B --local-dir ~/models/aeon-dflash2 --include "bf16/*"
export AEON_DIR="$HOME/models/aeon-mxfp4-rocm"
# Set AEON_CACHE, AEON_IMAGE and RENDER_GID as in the ROCm card's "Common setup" step, then:
AEON_DFLASH="$HOME/models/aeon-dflash2/bf16" bash "$AEON_DIR/scripts/serve_2x_r9700_dflash2.sh"
- Other AMD GPUs:
python3 "$AEON_DIR/scripts/detect_setup.py"picks the recipe. Recipes K (4x R9700) and Q (2x W7900) also use DFlash 2: in their plaindocker runcommands, replace the-v "$AEON_DIR/dflash2":/draft:romount with-v "$HOME/models/aeon-dflash2/bf16":/draft:ro. - Never the repo root on ROCm: vLLM cannot load the NVFP4 pack there.
- Memory (weights only): 23.7 GB target + 3.85 GB drafter ≈ 27.5 GB.
- Untested: lk4 has not been run on AMD hardware. The kit's measured τ of 3.73 at T=0.6 was with its bundled lk3 drafter.
curl -sf http://localhost:8000/health && echo OK
# Startup log: "Resolved architecture: DFlash2DraftModel". Under load: "SpecDecoding metrics: Mean acceptance length"
# of about 3 or more. About 1 with several requests in flight means the draft is not on TRITON_ATTN.
D. Other engines: bf16/
hf download AEON-7/AEON-DFlash2-Qwen3.8-27B --local-dir ~/models/aeon-dflash2 --include "bf16/*"
# Drafter dir for every engine below: ~/models/aeon-dflash2/bf16
- SGLang, llama.cpp, TensorRT-LLM, TensorFold and the MLX tools all take
bf16/, with a draft width of 9 (block 10). Versions, flags and commands for each: Use it with any inference engine. None has been tested with an AEON target yet. - Target: the BF16 target for SGLang and TensorRT-LLM, and a GGUF conversion of it for llama.cpp. TensorFold serves only its own Qwen3.8-27B target formats, not the AEON checkpoints as shipped. On Apple Silicon no AEON MLX build is published, so you would convert the BF16 target with mlx-lm yourself (untested).
- Check: SGLang,
llama-serverandtrtllm-serveanswercurl -sf http://localhost:<port>/health. Then read the acceptance length from the engine's speculative-decoding log or metrics: about 3 or more on normal prompts at T=0.6. Much lower usually means the drafter path or the draft width is wrong.
Speed: measured on one DGX Spark (lk3, the previous stage)
AEON DFlash 2 (lk3, tuned config: K=9 + lattice + probabilistic drafting) vs stock z-lab DFlash 2 (BF16, unmodified, K=7, its published config)
vs no speculation. Same DGX Spark, same published target build, same image and serving flags; only the drafter and --speculative-config change.
One user, all categories pooled
| Sampling | No speculation | Stock z-lab tok/s (τ) | AEON tok/s (τ) | AEON vs no speculation† | Stock z-lab vs no speculation† | AEON vs stock z-lab [95% CI] |
|---|---|---|---|---|---|---|
| T=0.6 (served default; 537 prompts, 18 categories) | 10.1 | 27.0 (3.56) | 32.7 (4.12) | 3.24× | 2.67× | +21.1% [+19.6, +22.5] |
| T=1.0 (Qwen / z-lab protocol; 180 prompts, 8 categories) | 10.1 | 24.9 (3.26) | 33.4 (4.18) | 3.30× | 2.45× | +34.6% [+31.5, +37.7] |
| T=0 (greedy; same 180 prompts) | 10.2 | 31.5 (4.11) | 35.9 (4.47) | 3.53× | 3.09× | +14.0% [+11.1, +16.9] |
τ = acceptance length: tokens produced per forward pass of the 27B target. Decode tok/s per request, thinking included.
† "vs no speculation" multiples are unpaired. AEON ran 537 prompts at T=0.6 (180 at T=1.0 and T=0); the no-speculation baseline ran only 72 prompts at T=0.6 (4 per category) and 16 at T=1.0 and T=0 (2 per category). Its decode speed is nearly constant (about 10.1 tok/s in every category), so each multiple divides by that category's no-speculation mean rather than pairing prompt by prompt.
By category, T=0.6 (sorted by gain; every category is faster, every interval above zero)
| Category (source) | Prompts | No speculation tok/s | Stock z-lab tok/s (τ) | AEON tok/s (τ) | AEON vs no speculation† | AEON vs stock z-lab [95% CI] |
|---|---|---|---|---|---|---|
| MBPP (z-lab task set) | 24 | 10.1 | 28.1 (3.68) | 36.4 (4.55) | 3.60× | +29.6% [+25.0%, +34.2%] |
| GSM8K (z-lab task set) | 24 | 10.1 | 37.9 (4.95) | 48.0 (5.99) | 4.74× | +26.7% [+22.3%, +32.0%] |
| MATH-500 (z-lab task set) | 24 | 10.1 | 33.7 (4.42) | 42.4 (5.31) | 4.18× | +25.7% [+20.9%, +31.2%] |
| Humanities (MMLU-Pro) | 32 | 10.1 | 27.1 (3.55) | 33.9 (4.24) | 3.35× | +25.3% [+19.6%, +30.6%] |
| Tool calling (BFCL v3) | 32 | 10.2 | 41.0 (5.37) | 51.1 (6.39) | 5.02× | +24.7% [+17.1%, +32.7%] |
| HumanEval (z-lab task set) | 24 | 10.1 | 27.4 (3.59) | 34.0 (4.25) | 3.36× | +24.1% [+16.1%, +33.1%] |
| MT-Bench, turn 1 (z-lab task set) | 24 | 10.2 | 25.5 (3.34) | 31.6 (3.94) | 3.12× | +23.9% [+18.4%, +30.3%] |
| Math (GSM8K, MATH-500) | 32 | 10.1 | 35.9 (4.70) | 44.3 (5.54) | 4.38× | +23.7% [+19.3%, +28.2%] |
| Coding (HumanEval, MBPP, HumanEvalPack) | 33 | 10.1 | 29.5 (3.86) | 36.3 (4.53) | 3.59× | +22.9% [+17.0%, +29.1%] |
| Multilingual (WMT24++, MGSM) | 32 | 10.1 | 27.4 (3.60) | 33.5 (4.19) | 3.31× | +22.2% [+16.8%, +27.8%] |
| RAG and QA (HotpotQA, NQ-Open) | 32 | 10.1 | 27.9 (3.67) | 33.8 (4.23) | 3.35× | +21.3% [+14.9%, +28.7%] |
| Multi-turn (MT-Bench, turn 2) | 32 | 10.1 | 24.5 (3.22) | 29.6 (3.71) | 2.93× | +20.7% [+14.2%, +27.5%] |
| Reasoning (BIG-Bench Hard) | 32 | 10.1 | 31.2 (4.09) | 37.6 (4.69) | 3.72× | +20.4% [+15.4%, +25.6%] |
| STEM (MMLU-Pro) | 32 | 10.1 | 29.4 (3.85) | 35.3 (4.42) | 3.50× | +20.1% [+13.7%, +26.2%] |
| Summarization (CNN/DailyMail) | 32 | 10.1 | 32.0 (4.21) | 38.3 (4.79) | 3.78× | +19.6% [+14.8%, +24.7%] |
| Roleplay (RoleBench) | 32 | 10.1 | 20.0 (2.63) | 23.7 (2.97) | 2.35× | +18.5% [+14.9%, +22.5%] |
| Long context (GovReport, ~8K and ~15K-token inputs) | 32 | 10.0 | 20.8 (2.91) | 24.1 (3.23) | 2.42× | +15.7% [+11.8%, +19.8%] |
| Creative writing (Arena-Hard) | 32 | 10.1 | 21.3 (2.80) | 24.3 (3.03) | 2.40× | +13.8% [+7.2%, +20.4%] |
By category, T=1.0 (z-lab task set plus tools, writing and multilingual)
| Category (source) | Prompts | No speculation tok/s | Stock z-lab tok/s (τ) | AEON tok/s (τ) | AEON vs no speculation† | AEON vs stock z-lab [95% CI] |
|---|---|---|---|---|---|---|
| HumanEval (z-lab task set) | 24 | 10.1 | 23.5 (3.08) | 33.6 (4.20) | 3.32× | +43.1% [+36.0%, +50.2%] |
| MBPP (z-lab task set) | 24 | 10.1 | 23.6 (3.09) | 33.0 (4.13) | 3.25× | +39.7% [+31.9%, +47.4%] |
| MATH-500 (z-lab task set) | 24 | 10.1 | 28.0 (3.67) | 38.9 (4.87) | 3.84× | +39.0% [+30.6%, +47.2%] |
| Tool calling (BFCL v3) | 20 | 10.2 | 35.8 (4.69) | 49.5 (6.16) | 4.84× | +38.4% [+24.4%, +53.0%] |
| GSM8K (z-lab task set) | 24 | 10.2 | 34.2 (4.48) | 47.1 (5.87) | 4.64× | +37.9% [+29.2%, +45.3%] |
| MT-Bench, turn 1 (z-lab task set) | 24 | 10.1 | 23.1 (3.02) | 30.4 (3.80) | 3.00× | +31.7% [+24.8%, +39.9%] |
| Creative writing (Arena-Hard) | 20 | 10.1 | 18.9 (2.47) | 24.2 (3.01) | 2.39× | +28.1% [+19.3%, +36.2%] |
| Multilingual (WMT24++, MGSM) | 20 | 10.1 | 25.7 (3.37) | 31.8 (3.98) | 3.14× | +23.8% [+15.0%, +32.8%] |
By category, T=0 (greedy)
| Category (source) | Prompts | No speculation tok/s | Stock z-lab tok/s (τ) | AEON tok/s (τ) | AEON vs no speculation† | AEON vs stock z-lab [95% CI] |
|---|---|---|---|---|---|---|
| Tool calling (BFCL v3) | 20 | 10.3 | 42.9 (5.58) | 54.5 (6.80) | 5.30× | +26.9% [+21.0%, +32.1%] |
| MATH-500 (z-lab task set) | 24 | 10.2 | 35.1 (4.59) | 42.1 (5.26) | 4.15× | +19.9% [+16.2%, +23.9%] |
| Multilingual (WMT24++, MGSM) | 20 | 10.2 | 31.2 (4.08) | 36.8 (4.58) | 3.62× | +17.9% [+12.7%, +23.6%] |
| GSM8K (z-lab task set) | 24 | 10.2 | 42.3 (5.51) | 49.7 (6.17) | 4.88× | +17.3% [+11.6%, +23.6%] |
| HumanEval (z-lab task set) | 24 | 10.2 | 30.4 (3.97) | 35.2 (4.39) | 3.46× | +16.0% [+10.2%, +22.7%] |
| MBPP (z-lab task set) | 24 | 10.2 | 31.2 (4.08) | 35.3 (4.39) | 3.47× | +12.9% [+7.4%, +19.6%] |
| MT-Bench, turn 1 (z-lab task set) | 24 | 10.2 | 29.2 (3.82) | 31.9 (3.98) | 3.13× | +9.1% [+1.7%, +17.4%] |
| Creative writing (Arena-Hard) | 20 | 10.1 | 23.8 (3.11) | 25.6 (3.20) | 2.52× | +7.5% [-3.6%, +18.7%] |
At T=0 the gain is smallest on creative writing, where the interval crosses zero.
Concurrency (T=0.6; 48 math, coding and writing prompts, outputs capped at 512 tokens; two replicate runs per arm)
| Concurrent users | No speculation total tok/s | Stock z-lab total tok/s (τ) | AEON total tok/s (τ) | AEON vs no speculation | AEON vs stock z-lab |
|---|---|---|---|---|---|
| 1 | 10.1 | 27.0 (3.56) | 32.7 (4.12) | 3.24׆ | +21.1% |
| 4 | 36.2 | 79.8 (3.39) | 98.4 (3.95) | 2.72× | +23.3% |
| 16 | 111.6 | 180.8 (3.45) | 217.1 (3.38) | 1.95× | +20.1% |
| Concurrent users | TTFT p50, no spec / stock z-lab / AEON | Time per output token p50, no spec / stock z-lab / AEON |
|---|---|---|
| 4 | 0.35 s / 0.49 s / 0.47 s | 102.6 ms / 44.5 ms / 35.7 ms |
| 16 | 1.73 s / 0.85 s / 0.71 s | 119.7 ms / 68.9 ms / 56.6 ms |
The 1-user row is the pooled T=0.6 result above (all 537 prompts). The 4- and 16-user rows are closed-loop runs over the 48-prompt set; "vs stock z-lab" there is a ratio of the two arms' totals.
The chart plots steady-state tok/s (ramp-up and drain excluded): 102 / 247 at 4 / 16 users for AEON, 84 / 208 for stock z-lab. The table gives whole-run totals.
Where the gain comes from (T=0.6, the 196 prompts all three arms ran)
| Step | tok/s | τ | Gain [95% CI] |
|---|---|---|---|
| Stock z-lab (BF16, K=7, published config) | 26.8 | 3.54 | baseline |
| + our serving config (z-lab weights packed to NVFP4 W4A16, K=9 lattice, probabilistic drafting) | 31.6 | 3.97 | +17.7% [+15.5, +20.0] |
| + our training (AEON DFlash 2, lk3) | 32.7 | 4.11 | +3.5% [+1.6, +5.5] |
The serving config does most of the work, and you can apply it to the stock drafter too (docs/HOSTING.md). Training adds a further, statistically clear +3.5% on top. lk4 adds about +0.7% more acceptance over lk3 (offline, below).
Full method, datasets and caveats: docs/BENCHMARKS.md.
Acceptance, this release (lk4 = S5 step 753)
Offline acceptance on the published NVFP4-MIXED build: 213 held-out windows of the target's own served traffic, weighted by workload type, with item-cluster bootstrap 95% confidence intervals. Change in acceptance length τ. lk4 is the step-753 save of training stage S5, chosen by the release rule's tie-break (step 690, which shipped for the first hours of v1.0, scored the same overall).
| lk4 (this release) vs | Change in τ [95% CI] |
|---|---|
| Stock z-lab BF16 | +4.87% [+4.38, +5.51] |
| z-lab weights packed to NVFP4 W4A16 | +5.10% [+4.62, +5.74] |
| lk3, the previous stage (whose served numbers are above) | +0.70% [+0.47, +1.00] |
For reference, lk3 measured +4.13% [+3.69, +4.66] over stock z-lab BF16 on the same windows.
| Part of the response | lk4 vs stock z-lab | lk4 vs lk3 |
|---|---|---|
| Tool calls | +8.54% [+7.47, +9.90] | +1.41% [+0.74, +2.15] |
| Thinking | +4.93% [+4.25, +5.72] | +0.68% [+0.38, +0.99] |
| Final answer | +4.02% [+3.21, +4.79] | +0.47% [−0.04, +1.05] |
| Workload (held-out windows) | lk4 vs stock z-lab | lk4 vs lk3 |
|---|---|---|
| Assistant chat (100) | +5.59% [+4.95, +6.34] | +0.82% [+0.55, +1.11] |
| Coding agent (12) | +5.38% [+4.38, +6.60] | +1.11% [+0.72, +1.78] |
| Voice assistant (85) | +2.50% [+1.15, +4.09] | +0.33% [−0.72, +1.56] |
| Persona chat (16) | +3.89% [+2.80, +7.04] | −0.04% [−0.41, +0.76] |
What to expect served: lk3's served speed above, plus roughly lk4's +0.7% acceptance gain. That is an expectation, not a measurement. lk4's own served benchmark is next.
Notes and caveats
- Stock z-lab needs
TRITON_ATTNon this image. On the AEON vLLM 0.29 image, the stock z-lab drafter would not boot until its--speculative-configincluded"attention_backend":"TRITON_ATTN". Every z-lab arm above ran that way, with everything else at its published config (K=7, greedy draft sampling). - Accuracy parity: no significant difference, but the check is small and is not proof. Paired against no speculation on 34 scored prompts at T=0.6 (drawn from GSM8K, MATH-500, BBH, BFCL, MMLU-Pro, MGSM, HotpotQA and NQ-Open): AEON −2.9 points [−11.8, +5.9], stock z-lab +2.9 [−8.8, +14.7], McNemar p = 1.00 for both. That is too small to rule out small effects, and it was dominated by 512-token output caps on the no-speculation baseline: with thinking on, many of its answers were cut off, which is why its accuracy is only 35.3%. It says little about quality either way. "Lossless" rests on construction (rejection sampling has the target verify every token), not on this check. A larger accuracy check with full-length outputs is part of the next benchmark.
- Not byte-identical, even at T=0. No arm reproduced the no-speculation text exactly at T=0 (0 of 16 for both AEON and stock z-lab). A different number of tokens per step changes floating-point rounding and can flip near-ties. The distribution is unchanged.
- BF16 SSM state cache: not recommended.
--mamba-ssm-cache-dtype bfloat16gave +15.5% total throughput at 16 users (248 vs 215 tok/s, one replicate). Quality: inconclusive overall (−1.8 pts pooled over 388 scored items, 95% CI [−4.4, +0.5], crosses zero); one chained-IFEval subset dropped 12.5 pts (95.8% → 83.3%, n=48, McNemar p=0.07); not recommended. Keep the FP32 default. - Home-field advantage, stated plainly. z-lab's drafter was trained for stock Qwen3.8-27B in BF16. Here it drafts for a modified NVFP4-MIXED model it never saw. Ours was trained for exactly this target.
- Different from z-lab's published numbers. Their card uses SGLang on an H200 with the stock target. Different hardware, engine and target.
More tests and data coming
- Full lk4 benchmark on the DGX Spark: every category at 1, 4, 8 and 16 concurrent users, with acceptance, tok/s, time to first token, time per output token, and prefill at 1K, 8K and 32K-token inputs.
- Continued training: the next stage is underway, on fresh on-policy data from the published build. A new version is published only when it shows an overall acceptance gain, and every version is tracked in EVOLUTION.md.
- More platforms: RTX 5090 / RTX PRO 6000 and stock upstream vLLM are being validated (recipes/).
Quickstart: one DGX Spark
This is quickstart A in full: NVIDIA Blackwell, with the NVFP4 W4A16 drafter from the repo root. Other hardware: Quickstart by architecture.
Using an AI agent to deploy? Point it at AGENTS.md.
1. Get access and download. Request access on the model page first, then:
python3 -m venv ~/hf && . ~/hf/bin/activate && pip install -U huggingface_hub
hf auth login
hf download AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4-MIXED --local-dir ~/models/aeon-27b
hf download AEON-7/AEON-DFlash2-Qwen3.8-27B --local-dir ~/models/aeon-dflash2 --exclude "bf16/*"
docker pull ghcr.io/aeon-7/aeon-vllm-ultimate:2026-09-18-v0.29.0-omni
Use --local-dir: Hugging Face cache snapshot folders are symlinks and break inside Docker. The image is public (arm64).
--exclude "bf16/*" downloads the repo root (the NVFP4 W4A16 drafter, 1.93 GB) and skips the 3.85 GB BF16 copy in bf16/, which Blackwell doesn't need.
2. Serve. /draft = this repo's root (the NVFP4 W4A16 pack, not bf16/). This matches the measured production config, except: default memory 0.80 (we measured at 0.57 with other services on the box), aliases omitted (extra --served-model-name entries), restart policy optional:
GPU_UTIL=0.80 # dedicated Spark. Use 0.57 if other GPU services share the box (our measured setting).
docker run -d --name aeon-dflash2 --gpus all --network host \
-e VLLM_USE_V2_MODEL_RUNNER=1 \
-e VLLM_ENABLE_CUDA_COMPATIBILITY=0 \
-e VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 \
-e VLLM_USE_FLASHINFER_SAMPLER=1 \
-e VLLM_ALLREDUCE_USE_FLASHINFER=0 \
-v ~/models/aeon-27b:/model:ro \
-v ~/models/aeon-dflash2:/draft:ro \
-v ~/models/aeon-27b/modelopt-54367.py:/usr/local/lib/python3.12/site-packages/vllm/model_executor/layers/quantization/modelopt.py:ro \
--entrypoint vllm ghcr.io/aeon-7/aeon-vllm-ultimate:2026-09-18-v0.29.0-omni \
serve /model \
--served-model-name aeon --host 0.0.0.0 --port 8000 \
--gpu-memory-utilization $GPU_UTIL \
--max-model-len 262144 \
--max-num-seqs 16 \
--max-num-batched-tokens 16384 \
--kv-cache-dtype fp8 \
--enable-chunked-prefill \
--tool-call-parser qwen3_coder --enable-auto-tool-choice \
--reasoning-parser qwen3 \
--limit-mm-per-prompt '{"image":4,"video":2}' \
--attention-backend TRITON_ATTN \
--compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE"}' \
--trust-remote-code \
--speculative-config '{"method":"dflash","model":"/draft","num_speculative_tokens":9,"num_speculative_tokens_per_batch_size":[[1,2,9],[3,8,7],[9,12,6],[13,16,4]],"attention_backend":"TRITON_ATTN","draft_sample_method":"probabilistic"}' \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
modelopt-54367.pyships in the target's model folder. The measured container mounts it; the target's card notes the same patch is already baked into this image tag, so the mount is a harmless pin.- Memory:
0.80on a dedicated Spark (the target card's setting). We measured at0.57because speech services share our Spark. Don't go above 0.80: running out of unified memory can freeze the machine. - Never K > 9. Don't combine with MTP. No
--quantizationflag. - Add
--restart unless-stoppedto start again after a reboot, and--api-key <secret>if other people can reach port 8000.
3. Health check. The first boot takes several minutes (docker logs -f aeon-dflash2).
curl -sf http://localhost:8000/health && echo OK
curl -s http://localhost:8000/v1/models # lists "aeon" when ready
4. Test request.
curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"aeon","messages":[{"role":"user","content":"Explain speculative decoding in two sentences."}],"max_tokens":512,"chat_template_kwargs":{"enable_thinking":false}}'
5. Measure τ. On an otherwise idle server:
scripts/test_request.sh # streams a thinking-on request; prints TTFT, decode tok/s and τ
curl -s http://localhost:8000/metrics | grep -E '^vllm:spec_decode_num_(drafts|accepted_tokens)_total'
# τ = 1 + accepted ÷ drafts (take the difference before and after a request)
At 1 user and T=0.6 we measured τ ≈ 4.1 pooled over all categories (lk3), from about 3 on roleplay and creative writing to over 6 on tool calls.
Much lower than 3 on normal prompts means something is off: check the /draft mount and "draft_sample_method":"probabilistic".
Docker Compose, Python client and troubleshooting: docs/QUICKSTART.md. Recipes per platform: recipes/.
Use it with any inference engine
The drop-in rule: any engine that runs z-lab/Qwen3.8-27B-DFlash2 runs this drafter. Point its drafter path at this repo's bf16/ folder instead and set the draft width to 9. Each engine names that setting differently (see the table).
Which folder to use:
- Repo root (NVFP4 W4A16): vLLM on NVIDIA Blackwell. Verified on DGX Spark with the AEON vLLM 0.29 image; stock upstream vLLM is being validated.
bf16/(exact BF16 copy of the same weights, the only BF16 weights in the repo): everything else, including vLLM on NVIDIA Hopper / Ampere / Ada and on AMD ROCm, SGLang (the NVFP4 pack is untested there), llama.cpp (convert frombf16/), TensorRT-LLM, TensorFold (which re-packs to 4-bit at load) and MLX.
| Engine | Minimum version | Status | Draft-width setting | Folder |
|---|---|---|---|---|
| vLLM (recommended) | 0.28.0 (V2 runner) | Supported; the path this release is built and tested on (DGX Spark) | "num_speculative_tokens": 9 |
root (NVIDIA Blackwell); bf16/ (other NVIDIA GPUs, ROCm) |
| SGLang | 0.5.19 | Supported by format; not yet tested with an AEON target | --speculative-num-draft-tokens 10 (block = 1 + 9) |
bf16/ |
| llama.cpp | build b10658 | Supported by format; needs a GGUF target and a GGUF conversion of this drafter | --spec-draft-n-max 9 |
bf16/ → GGUF |
| TensorRT-LLM | 1.3.0rc28 (pre-release), NVIDIA only | Supported by format; AEON targets untested | max_draft_len: 9 |
bf16/ |
| TensorFold | v0.6.5 | Loads this drafter; serves its own Qwen3.8-27B target formats | None: read from the config (block 10) | bf16/ |
| MLX (Apple Silicon) | n/a | Not supported (yet) in upstream mlx-lm; z-lab's dflash package and the oMLX fork run z-lab's drafter |
n/a | bf16/ |
| Ollama | n/a | Not supported (yet): needs an unmerged pull request | n/a | n/a |
| LM Studio | n/a | Unverified | n/a | n/a |
vLLM (NVIDIA Blackwell: the repo root; NVIDIA Hopper / Ampere / Ada and AMD ROCm: bf16/)
VLLM_USE_V2_MODEL_RUNNER=1 vllm serve <AEON target dir> \
--speculative-config '{"method":"dflash","model":"<drafter dir>","num_speculative_tokens":9,"num_speculative_tokens_per_batch_size":[[1,2,9],[3,8,7],[9,12,6],[13,16,4]],"attention_backend":"TRITON_ATTN","draft_sample_method":"probabilistic"}'
# Stock upstream vLLM (every release so far, through 0.31) with this hybrid-GDN target: add --no-enable-prefix-caching
# Hopper / Ampere / Ada: point "model" at the bf16/ folder and use the BF16 target (Quickstart B)
# ROCm: also pass --attention-backend TRITON_ATTN and point "model" at the bf16/ folder (Quickstart C)
SGLang
The block (verify window) is 10. If you leave the flag unset, SGLang reads it from the config. z-lab's snippet hard-codes 8, which would run this drafter narrower than it was trained.
python -m sglang.launch_server --model-path <AEON target dir> \
--speculative-algorithm DFLASH --speculative-draft-model-path <drafter>/bf16 \
--speculative-num-draft-tokens 10
llama.cpp
Convert the drafter from bf16/. Don't reuse z-lab's DFlash2 GGUFs: they hold z-lab's weights with block 8.
python convert_hf_to_gguf.py <drafter>/bf16 --target-model-dir <AEON target HF dir> \
--outtype bf16 --outfile aeon-dflash2.gguf
llama-server -m <AEON target>.gguf -md aeon-dflash2.gguf \
--spec-type draft-dflash --spec-draft-n-max 9 -fa on --jinja
TensorRT-LLM (trtllm-serve <target dir> --config spec.yaml)
speculative_config:
decoding_type: DFlash
max_draft_len: 9
speculative_model: <drafter>/bf16
TensorFold
TensorFold's DFlash2 loader reads this drafter and re-packs its linears to 4-bit at load. This can lower acceptance, but not output fidelity.
tensorfold serve <supported target> --drafter <drafter>/bf16
Only vLLM has been run with an AEON target. The other engines run z-lab's Qwen3.8-27B DFlash2, so by format they run this drafter too, but we haven't tested them with AEON targets yet.
Where it runs
| Setup | Status |
|---|---|
| NVIDIA DGX Spark (GB10) + AEON vLLM 0.29 image, repo root (NVFP4 W4A16) | Verified (lk3 measured; lk4 benchmark next) |
| RTX 5090 / RTX PRO 6000, repo root (NVFP4 W4A16) | Coming soon (being validated) |
| Stock upstream vLLM ≥ 0.30 | Coming soon (being validated) |
NVIDIA Hopper / Ampere / Ada (stock vLLM, bf16/ + BF16 target) |
Untested (quickstart B) |
AMD ROCm (vLLM, bf16/) |
Untested. The earlier lk3 BF16 build ships bundled in …-MXFP4-MXFP6-ROCm (dflash2/) |
What's in this repo
| Path | What |
|---|---|
Repo root: model.safetensors, config.json, hf_quant_config.json |
The NVFP4 W4A16 pack, for vLLM on NVIDIA Blackwell: 15 MLP linears in FP4 (E2M1, FP8 scale per 16 weights), everything else BF16. 1,926,977,716 bytes, sha256 02d956cfe9fbb0259400fdccc194de0f45db2a9bb12d7aa808ec91ebd36e575d (lk4 = S5 step 753) |
FIDELITY_GATE.json, LOAD_SHAPE_AUDIT.json |
Pack checks: dequantized-as-served vs BF16 reference, and the load-shape audit |
bf16/: model.safetensors, config.json, DEQUANT_REPORT.json |
The only BF16 weights in this repo, for every other architecture and engine: exact BF16 dequantization of the same weights (0 differing elements). 3,848,817,920 bytes, sha256 266075471fd5a72682f27da2b69ed4354846cad6e474e901ea2c8af04b4748ed |
configs/, scripts/, recipes/, docs/ |
Serving configs, test scripts, platform recipes, guides |
images/ |
Benchmark charts |
What it is
- A DFlash 2 draft model. It reads hidden states from five target layers and proposes up to 9 tokens per step; the target checks them in one forward pass.
- Fine-tuned from
z-lab/Qwen3.8-27B-DFlash2by Inco AI / z-lab (Apache-2.0). Its config differs only inblock_size10,causal: falseanddraft_vocab_size. - NVFP4 weights (W4A16) in the repo root; BF16 in
bf16/. In the root pack the 15 MLP linears hold FP4 weights. Attention and every other layer stay BF16. Activations are BF16, and vLLM's Marlin kernel runs the math (FlashInfer CuTe-DSL on B200 / B300 in vLLM ≥ 0.30): weight-only FP4, not FP4 tensor-core math. - Lossless by construction. Speculative decoding with rejection sampling preserves the target's output distribution: the target verifies every token.
How it works, in plain words: docs/HOW_IT_WORKS.md · Version history: EVOLUTION.md · CHANGELOG.md · FAQ
Data
Training prompts were drawn from public datasets (see CREDITS.md) plus synthetic agentic conversations. Every training target was generated by the target model itself.
Support the work
AEON-7 models, drafters and tools are built and trained independently, on my own hardware. If they're useful to you, please consider supporting development:
Become a member on Patreon → patreon.com/cw/AeonForge7/membership
Milestones unlock bigger work: reaching 500 paid supporters will fund fine-tuning larger models and bigger project releases. Supporters also get early access to new releases, such as the AEON DFlash2 drafter.
Attribution & sharing
You may share or build on this model under Apache-2.0. Keep the LICENSE and NOTICE files, and credit AEON-7 and z-lab / Inco AI.
- AEON-7: fine-tuning for the AEON Ultimate target, NVFP4 (W4A16) packing, serving config.
- Inco AI / z-lab: DFlash 2 and the original Qwen3.8-27B-DFlash2 (blog, code).
- Qwen team, Alibaba: Qwen3.8-27B, the base of the target model.
- Full acknowledgements, including vLLM, SpecForge, NVIDIA Model Optimizer, Marlin, FlashInfer, every training dataset and every benchmark dataset, with licenses: CREDITS.md.
@misc{aeon7_2026_dflash2_qwen38_27b,
title = {AEON DFlash 2 for Qwen3.8-27B AEON Ultimate},
author = {AEON-7},
year = {2026},
note = {Version 1.0 (lk4, S5 step 753). Fine-tuned from z-lab/Qwen3.8-27B-DFlash2},
url = {https://huggingface.co/AEON-7/AEON-DFlash2-Qwen3.8-27B}
}
@misc{inco2026dflash2,
title = {DFlash 2: Keep Drafting Parallel},
author = {Inco AI},
year = {2026},
month = {August},
url = {https://inco.ai/blog/dflash2/}
}
@inproceedings{chen2026dflash,
title = {DFlash: Block Diffusion for Flash Speculative Decoding},
author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
booktitle = {ICML},
year = {2026}
}
License
Apache-2.0. A modified (fine-tuned) version of z-lab / incoai Qwen3.8-27B-DFlash2, also Apache-2.0. See the LICENSE and NOTICE files.
- Downloads last month
- 32



