Skip to content

[vLLM-ATOM] Support EAGLE3 spec decoding for MiniMax-M3 - #1518

Merged
valarLip merged 48 commits into
ROCm:mainfrom
kliuae:kliuae/plugin_m3_eagle
Aug 1, 2026
Merged

[vLLM-ATOM] Support EAGLE3 spec decoding for MiniMax-M3#1518
valarLip merged 48 commits into
ROCm:mainfrom
kliuae:kliuae/plugin_m3_eagle

Conversation

@kliuae

@kliuae kliuae commented Jul 8, 2026

Copy link
Copy Markdown
Contributor

Motivation

This PR adds EAGLE 3/3.1 speculative decoding support for MiniMax-M3 in vllm-atom.
It depends on and currently contains changes from #1201, and will be rebased once the dependency gets merged.

Technical Details

Test Plan

Target model: MiniMaxAI/MiniMax-M3-MXFP8
Draft model: Inferact/MiniMax-M3-EAGLE3

Server command

VLLM_ROCM_USE_AITER=1 VLLM_USE_V1=1 \
vllm serve MiniMaxAI/MiniMax-M3-MXFP8 \
    --host localhost \
    --port 8000 \
    --no-trust-remote-code \
    --dtype auto \
    --kv-cache-dtype auto \
    --max-model-len 32768 \
    --max-num-batched-tokens 32768 \
    --no-enable-prefix-caching \
    --compilation-config '{"cudagraph_mode": "FULL_AND_PIECEWISE"}' \
    --tensor-parallel-size 4 \
    --block-size 128 \
    --additional-config '{"online_quant_config": {"global_quant_config": "ptpc_fp8", "exclude_layer": ["lm_head", "model.embed_tokens", "vision_tower", "multi_modal_projector", "patch_merge_mlp", "*block_
sparse_moe"]}}' \
    --language-model-only \
    --speculative-config.method eagle3 \
    --speculative-config.model Inferact/MiniMax-M3-EAGLE3 \
    --speculative-config.num_speculative_tokens 3

lm_eval

lm_eval --model local-completions --model_args model=MiniMaxAI/MiniMax-M3-MXFP8,base_url=http://0.0.0.0:8000/v1/completions,tokenizer_backend=None,tokenized_requests=False,num_concurrent=128,max_retries=10,max_gen_toks=2048,timeout=60000 --batch_size auto --tasks gsm8k --num_fewshot 20

Test Result

lm_eval, EAGLE=3

Tasks Version Filter n-shot Metric Value Stderr
gsm8k 3 flexible-extract 20 exact_match _ 0.9431 _ 0.0064
strict-match 20 exact_match _ 0.9431 _ 0.0064

gsm8k: Per-position acceptance rate: 0.876, 0.752, 0.626, Avg Draft acceptance rate: 75.1%

Target model: MiniMaxAI/MiniMax-M3-MXFP8
Draft model: Inferact/MiniMax-M3-EAGLE3

Mode ISL/OSL Conc EAGLE3 Req/s delta Total tok/s delta Output tok/s delta median TTFT (ms) delta TPOT mean/med/p99 (ms) delta ITL mean/med/p99 (ms) delta
ATOM native 1k/1k 64 no 4.235 - 8672.6 - 4336.3 - 906 - 13.84 - 13.73 -
vLLM plugin 1k/1k 64 no 3.785 -10.6% 7752.2 -10.6% 3876.1 -10.6% 1487 +64% 14.36 +3.8% 14.02 +2.1%
ATOM native 1k/1k 64 spec=3 4.950 - 10144.5 - 5075.6 - 149 - 9.37 - 24.51 -
vLLM plugin 1k/1k 64 spec=3 4.515 -8.8% 9247.7 -8.8% 4623.8 -8.9% 177 +19% 10.24 +9.3% 25.38 +3.6%

Submission Checklist

kliuae and others added 30 commits June 13, 2026 00:18
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: miloice <kuanfu.liu@embeddedllm.com>
Signed-off-by: miloice <kuanfu.liu@embeddedllm.com>
Signed-off-by: miloice <kuanfu.liu@embeddedllm.com>
Fix PTPC FP8 MoE loading to preserve offline checkpoint bits and wire MiniMax-M3 sparse MHA metadata/backend support for vLLM serving.

Co-authored-by: Cursor <cursoragent@cursor.com>
Reuse vLLM-provided output buffers in sparse MHA prefill/decode and align the adapter with the page-shuffled KV cache layout used by MiniMax-M3 serving.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Keep mixed decode/prefill/extend batches phase-local and separate index-cache top-k metadata from main KV-cache sparse block emission to prevent cross-request fp8 accuracy drift.
Route MiniMax-M3 vLLM-plugin attention through the vendored community implementation by default so serving has a stable correctness baseline while the ATOM plugin attention integration is debugged.
Remove the community MiniMax-M3 vLLM backend and registry switches so serving uses ATOM-owned dense and sparse attention paths.
Keep MiniMax-M3 sparse attention on the dedicated adapter path and rename the backend metadata pieces to match their model-specific ownership.
Restore the model-local sparse attention implementation to match main so the vLLM PR only carries the adapter-side changes needed for MiniMax-M3.
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Co-authored-by: whx-sjtu <xiaowang990929@gmail.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
return mla_layers, mha_layers


def _build_heterogeneous_kv_cache_groups(kv_cache_spec):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do we need this? Why can't just let vllm side handle the management of kv cache groups?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is to match native ATOM's use case of MLA target with MHA draft (e.g. Kimi-K2.6 with lightseekorg/kimi-k2.6-eagle3 drafter). This requires separate kv cache groups for MLA and MHA respectively, but this configuration is currently not supported in vllm which crashes when it tries to unify page sizes of its cache pools. Therefore, we build separate kv cache groups for plugin and patch their peripheries like memory usage calculation to accommodate this use case.

@kliuae
kliuae marked this pull request as ready for review July 16, 2026 08:14
@zejunchen-zejun
zejunchen-zejun requested a review from yhl-amd July 17, 2026 02:02
yhl-amd
yhl-amd previously approved these changes Jul 17, 2026
@zufayu
zufayu requested a review from ZhangLirong-amd July 17, 2026 03:00
whx-sjtu
whx-sjtu previously approved these changes Jul 17, 2026
kliuae added 4 commits July 20, 2026 21:01
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
…3_eagle

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
@kliuae
kliuae dismissed stale reviews from whx-sjtu and yhl-amd via 4dadc9a July 27, 2026 02:05
@zejunchen-zejun

Copy link
Copy Markdown
Collaborator

Hi, @kliuae
Please help resolve the format issue
Thank you

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
kliuae added 2 commits July 27, 2026 11:32
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
@zejunchen-zejun

Copy link
Copy Markdown
Collaborator

Hi, @kliuae
Could you help rebase this PR? We plan to merge it soon.
Thank you

Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Comment thread atom/models/llama.py Outdated
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
@ZhangLirong-amd

Copy link
Copy Markdown
Collaborator

please fix the ci code format

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@valarLip
valarLip merged commit e48347f into ROCm:main Aug 1, 2026
40 of 57 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants