Instructions to use guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN
- SGLang
How to use guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN with Docker Model Runner:
docker model run hf.co/guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN
MiniMAX-M2.7-MXFP4-CANN
MXFP4 (W4A4) quantized MiniMax-M2.7 for Ascend CANN / vLLM-Ascend.
MiniMax-M2.7 的 MXFP4 量化版本,面向昇腾 NPU(CANN)推理部署。对 Attention QKV 与 MoE Expert 权重进行 MXFP4 量化,采用 PTQ(SmoothQuant OSPlus)+ QAT 两阶段方案,量化全程在 昇腾 910C 上完成。
在推荐推理配置下(VLLM_MXFP4_ACT_QDQ_SCALE_FACTOR=5.8),7 项评测任务相对 BF16 基线的**平均性能保有率为 99.3%**。
- Base model: MiniMaxAI/MiniMax-M2.7
- Quantization: MXFP4 W4A4(Attention QKV + MoE Experts)
- Algorithm: PTQ (SmoothQuant OSPlus) → QAT
- Quantization hardware: Ascend 910C
- Inference: cann-recipes-infer / vLLM MiniMax MXFP4
Model Details
| Item | Value |
|---|---|
| Architecture | MiniMax-M2 MoE (MiniMaxM2ForCausalLM) |
| Total / activated params | 229B / ~12B (top-8 of 256 experts) |
| Layers | 62 |
| Hidden size | 3072 |
| Attention | GQA, 48 heads / 8 KV heads, head_dim=128, QK-Norm |
| Context length | 204,800 |
| Quantized modules | Attention Q / K / V projections, MoE expert weights |
| Weight format | MXFP4 (FP4-E2M1 + E8M0 block scale, group_size=32) |
| Activation | Online MXFP4 QDQ (group_size=32) |
| Quantization pipeline | PTQ (SmoothQuant OSPlus) + QAT |
| Target runtime | Ascend CANN + vLLM-Ascend |
This checkpoint is intended for Ascend NPU serving via the official CANN vLLM recipe. It is not a drop-in replacement for generic CUDA vLLM MXFP4 loaders.
Quantization
Quantization is a two-stage pipeline, executed end-to-end on Ascend 910C:
PTQ — SmoothQuant OSPlus
Search per-layer smoothing scales on calibration data, fuse the scales into BF16 weights, then convert fused weights to MXFP4 (RTN, block size 32). This reduces activation outliers before 4-bit quantization.QAT
Quantization-aware training on the PTQ checkpoint to recover accuracy lost in the 4-bit conversion, especially on knowledge-intensive and long-context tasks.
What is quantized
- Attention:
q_proj/k_proj/v_proj - MoE: expert FFN weights (W4A4 MXFP4)
Typical high-precision leftovers (kept closer to original precision): embeddings, RMSNorm, router / gate, and lm_head.
At serve time, weights are loaded from the MXFP4 (Quark-style) checkpoint; activations are quantized online with MXFP4 QDQ. The vLLM-Ascend recipe can additionally apply KV-cache MXFP4 QDQ (token-wise, group_size=32, with the first VLLM_KV_MXFP4_ANCHOR tokens kept in BF16) plus block-diagonal Hadamard rotation on Q/K.
Evaluation
Evaluated with the cann-recipes-infer MiniMax MXFP4 vLLM recipe and:
VLLM_MXFP4_ACT_QDQ_SCALE_FACTOR=5.8
VLLM_MXFP4_ACT_QDQ_SCALE_FACTOR is the block-scale divisor used in activation MXFP4 QDQ:
scale = 2^round(log2(max_abs / scale_factor))
5.8 is the recommended setting for this checkpoint (MXFP4 E2M1 codebook max is 6.0; 5.8 is a slight tightening that we found preserves accuracy better than the default 6.0 on this model).
| Model | HumanEval+ | GSM8K | MATH500 | GPQA Diamond | DROP | LiveCodeBench | LongBench v2 | Average |
|---|---|---|---|---|---|---|---|---|
| M2.7 baseline (BF16) | 91.6 | 95.55 | 91.26 | 88.88 | 89.71 | 60.28 | 54.87 | 81.736 |
| M2.7 W4A4 (this model) | 90.85 | 95.26 | 90.6 | 86.26 | 88.78 | 63.07 | 53.362 | 81.169 |
| Retention (quant / baseline) | 0.992 | 0.997 | 0.993 | 0.971 | 0.990 | 1.043 | 0.973 | 0.993 |
Average retention across the seven tasks is 99.3% of the BF16 baseline. LiveCodeBench is slightly above the baseline (sampling variance is possible on generation benchmarks).
How to Use
Follow the official CANN recipe (MiniMax-M2.5 and MiniMax-M2.7 share the same architecture; point MODEL_PATH at this repo):
Recipe: https://gitcode.com/cann/cann-recipes-infer/tree/master/integration/vllm/minimax_m2.5_mxfp4
Hardware
| Item | Requirement |
|---|---|
| Device | Atlas A3 / Ascend 910C (Ascend 910_93) |
| Cards | 16 NPUs (TP=16, Expert Parallel on) |
| Disk | Enough space for this MXFP4 checkpoint |
Recommended image
docker pull quay.io/ascend/vllm-ascend:v0.18.0rc1-a3
Serve
cd cann-recipes-infer/integration/vllm/minimax_m2.5_mxfp4
source set_env.sh
bash patch_vllm/apply.sh
MODEL_PATH=/path/to/guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN \
SERVED_MODEL_NAME=MiniMAX-M2.7-MXFP4-CANN \
TP_SIZE=16 \
VLLM_MXFP4_ACT_QDQ_SCALE_FACTOR=5.8 \
VLLM_KV_MXFP4_ANCHOR=32 \
bash run_vllm_w4a4.sh
VLLM_MXFP4_ACT_QDQ_SCALE_FACTOR=5.8 is required to reproduce the accuracy numbers above. Do not omit it if you care about the reported retention.
Sanity check
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "MiniMAX-M2.7-MXFP4-CANN",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 256
}'
Decoding parameters
Inherited from MiniMax-M2.7:
temperature=1.0top_p=0.95top_k=40
Default system prompt:
You are a helpful assistant. Your name is MiniMax-M2.7 and is built by MiniMax.
More environment variables (TP_SIZE, KV-cache MXFP4 switches, FlashComm, tool/reasoning parsers, etc.) are documented in the recipe README.
Intended Use & Limitations
- Intended use: offline / online inference of MiniMax-M2.7 on Ascend NPUs with vLLM-Ascend, including coding, math, QA, and long-context workloads covered by the table above.
- Not intended: training from this MXFP4 checkpoint; serving on non-Ascend backends without the CANN vLLM patches.
- Limitations: 4-bit W4A4 still shows a larger gap on GPQA Diamond (97.1% retention) and LongBench v2 (97.3% retention) than on GSM8K / MATH500. Generation benchmarks can fluctuate with sampling.
License
This repository is a quantized derivative of MiniMaxAI/MiniMax-M2.7. Use of the weights is subject to the original MiniMax license:
https://github.com/MiniMax-AI/MiniMax-M2.7/blob/main/LICENSE
Acknowledgements
- Base model: MiniMaxAI/MiniMax-M2.7
- Inference recipe: cann-recipes-infer (vLLM / vLLM-Ascend MXFP4 patches)
- Downloads last month
- 57
Model tree for guanwenyu1995/MiniMAX-M2.7-MXFP4-CANN
Base model
MiniMaxAI/MiniMax-M2.7Evaluation results
- pass@1 on HumanEval+self-reported90.850
- accuracy on GSM8Kself-reported95.260
- accuracy on MATH-500self-reported90.600
- accuracy on GPQA Diamondself-reported86.260
- f1/accuracy on DROPself-reported88.780
- pass@1 on LiveCodeBenchself-reported63.070
- accuracy on LongBench v2self-reported53.362