MiniMAX-M2.7-MXFP4-CANN

MXFP4 (W4A4) quantized MiniMax-M2.7 for Ascend CANN / vLLM-Ascend.

MiniMax-M2.7 的 MXFP4 量化版本,面向昇腾 NPU(CANN)推理部署。对 Attention QKV 与 MoE Expert 权重进行 MXFP4 量化,采用 PTQ(SmoothQuant OSPlus)+ QAT 两阶段方案,量化全程在 昇腾 910C 上完成。

在推荐推理配置下(VLLM_MXFP4_ACT_QDQ_SCALE_FACTOR=5.8),7 项评测任务相对 BF16 基线的**平均性能保有率为 99.3%**。

Model Details

Item Value
Architecture MiniMax-M2 MoE (MiniMaxM2ForCausalLM)
Total / activated params 229B / ~12B (top-8 of 256 experts)
Layers 62
Hidden size 3072
Attention GQA, 48 heads / 8 KV heads, head_dim=128, QK-Norm
Context length 204,800
Quantized modules Attention Q / K / V projections, MoE expert weights
Weight format MXFP4 (FP4-E2M1 + E8M0 block scale, group_size=32)
Activation Online MXFP4 QDQ (group_size=32)
Quantization pipeline PTQ (SmoothQuant OSPlus) + QAT
Target runtime Ascend CANN + vLLM-Ascend

This checkpoint is intended for Ascend NPU serving via the official CANN vLLM recipe. It is not a drop-in replacement for generic CUDA vLLM MXFP4 loaders.

Quantization

Quantization is a two-stage pipeline, executed end-to-end on Ascend 910C:

  1. PTQ — SmoothQuant OSPlus
    Search per-layer smoothing scales on calibration data, fuse the scales into BF16 weights, then convert fused weights to MXFP4 (RTN, block size 32). This reduces activation outliers before 4-bit quantization.

  2. QAT
    Quantization-aware training on the PTQ checkpoint to recover accuracy lost in the 4-bit conversion, especially on knowledge-intensive and long-context tasks.

What is quantized

  • Attention: q_proj / k_proj / v_proj
  • MoE: expert FFN weights (W4A4 MXFP4)

Typical high-precision leftovers (kept closer to original precision): embeddings, RMSNorm, router / gate, and lm_head.

At serve time, weights are loaded from the MXFP4 (Quark-style) checkpoint; activations are quantized online with MXFP4 QDQ. The vLLM-Ascend recipe can additionally apply KV-cache MXFP4 QDQ (token-wise, group_size=32, with the first VLLM_KV_MXFP4_ANCHOR tokens kept in BF16) plus block-diagonal Hadamard rotation on Q/K.

Evaluation

Evaluated with the cann-recipes-infer MiniMax MXFP4 vLLM recipe and:

VLLM_MXFP4_ACT_QDQ_SCALE_FACTOR=5.8

VLLM_MXFP4_ACT_QDQ_SCALE_FACTOR is the block-scale divisor used in activation MXFP4 QDQ:

scale = 2^round(log2(max_abs / scale_factor))

5.8 is the recommended setting for this checkpoint (MXFP4 E2M1 codebook max is 6.0; 5.8 is a slight tightening that we found preserves accuracy better than the default 6.0 on this model).

Model HumanEval+ GSM8K MATH500 GPQA Diamond DROP LiveCodeBench LongBench v2 Average
M2.7 baseline (BF16) 91.6 95.55 91.26 88.88 89.71 60.28 54.87 81.736
M2.7 W4A4 (this model) 90.85 95.26 90.6 86.26 88.78 63.07 53.362 81.169
Retention (quant / baseline) 0.992 0.997 0.993 0.971 0.990 1.043 0.973 0.993

Average retention across the seven tasks is 99.3% of the BF16 baseline. LiveCodeBench is slightly above the baseline (sampling variance is possible on generation benchmarks).

How to Use

Follow the official CANN recipe (MiniMax-M2.5 and MiniMax-M2.7 share the same architecture; point MODEL_PATH at this repo):

Recipe: https://gitcode.com/cann/cann-recipes-infer/tree/master/integration/vllm/minimax_m2.5_mxfp4

Hardware

Item Requirement
Device Atlas A3 / Ascend 910C (Ascend 910_93)
Cards 16 NPUs (TP=16, Expert Parallel on)
Disk Enough space for this MXFP4 checkpoint

Recommended image

docker pull quay.io/ascend/vllm-ascend:v0.18.0rc1-a3

Serve

cd cann-recipes-infer/integration/vllm/minimax_m2.5_mxfp4
source set_env.sh
bash patch_vllm/apply.sh

MODEL_PATH=/path/to/guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN \
SERVED_MODEL_NAME=MiniMAX-M2.7-MXFP4-CANN \
TP_SIZE=16 \
VLLM_MXFP4_ACT_QDQ_SCALE_FACTOR=5.8 \
VLLM_KV_MXFP4_ANCHOR=32 \
bash run_vllm_w4a4.sh

VLLM_MXFP4_ACT_QDQ_SCALE_FACTOR=5.8 is required to reproduce the accuracy numbers above. Do not omit it if you care about the reported retention.

Sanity check

curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMAX-M2.7-MXFP4-CANN",
    "messages": [{"role": "user", "content": "Hello"}],
    "max_tokens": 256
  }'

Decoding parameters

Inherited from MiniMax-M2.7:

  • temperature=1.0
  • top_p=0.95
  • top_k=40

Default system prompt:

You are a helpful assistant. Your name is MiniMax-M2.7 and is built by MiniMax.

More environment variables (TP_SIZE, KV-cache MXFP4 switches, FlashComm, tool/reasoning parsers, etc.) are documented in the recipe README.

Intended Use & Limitations

  • Intended use: offline / online inference of MiniMax-M2.7 on Ascend NPUs with vLLM-Ascend, including coding, math, QA, and long-context workloads covered by the table above.
  • Not intended: training from this MXFP4 checkpoint; serving on non-Ascend backends without the CANN vLLM patches.
  • Limitations: 4-bit W4A4 still shows a larger gap on GPQA Diamond (97.1% retention) and LongBench v2 (97.3% retention) than on GSM8K / MATH500. Generation benchmarks can fluctuate with sampling.

License

This repository is a quantized derivative of MiniMaxAI/MiniMax-M2.7. Use of the weights is subject to the original MiniMax license:

https://github.com/MiniMax-AI/MiniMax-M2.7/blob/main/LICENSE

Acknowledgements

Downloads last month
57
Safetensors
Model size
115B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN

Quantized
(112)
this model

Evaluation results