[vLLM-ATOM] Support EAGLE3 spec decoding for MiniMax-M3 - #1518
Conversation
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: miloice <kuanfu.liu@embeddedllm.com>
Signed-off-by: miloice <kuanfu.liu@embeddedllm.com>
Fix PTPC FP8 MoE loading to preserve offline checkpoint bits and wire MiniMax-M3 sparse MHA metadata/backend support for vLLM serving. Co-authored-by: Cursor <cursoragent@cursor.com>
Reuse vLLM-provided output buffers in sparse MHA prefill/decode and align the adapter with the page-shuffled KV cache layout used by MiniMax-M3 serving. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Keep mixed decode/prefill/extend batches phase-local and separate index-cache top-k metadata from main KV-cache sparse block emission to prevent cross-request fp8 accuracy drift.
Route MiniMax-M3 vLLM-plugin attention through the vendored community implementation by default so serving has a stable correctness baseline while the ATOM plugin attention integration is debugged.
Remove the community MiniMax-M3 vLLM backend and registry switches so serving uses ATOM-owned dense and sparse attention paths.
Keep MiniMax-M3 sparse attention on the dedicated adapter path and rename the backend metadata pieces to match their model-specific ownership.
Restore the model-local sparse attention implementation to match main so the vLLM PR only carries the adapter-side changes needed for MiniMax-M3.
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com> Co-authored-by: whx-sjtu <xiaowang990929@gmail.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
| return mla_layers, mha_layers | ||
|
|
||
|
|
||
| def _build_heterogeneous_kv_cache_groups(kv_cache_spec): |
There was a problem hiding this comment.
Why do we need this? Why can't just let vllm side handle the management of kv cache groups?
There was a problem hiding this comment.
This is to match native ATOM's use case of MLA target with MHA draft (e.g. Kimi-K2.6 with lightseekorg/kimi-k2.6-eagle3 drafter). This requires separate kv cache groups for MLA and MHA respectively, but this configuration is currently not supported in vllm which crashes when it tries to unify page sizes of its cache pools. Therefore, we build separate kv cache groups for plugin and patch their peripheries like memory usage calculation to accommodate this use case.
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
…3_eagle Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
|
Hi, @kliuae |
|
Hi, @kliuae |
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
|
please fix the ci code format |
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Motivation
This PR adds EAGLE 3/3.1 speculative decoding support for MiniMax-M3 in vllm-atom.
It depends on and currently contains changes from #1201, and will be rebased once the dependency gets merged.
Technical Details
Test Plan
Target model: MiniMaxAI/MiniMax-M3-MXFP8
Draft model: Inferact/MiniMax-M3-EAGLE3
Server command
lm_eval
Test Result
lm_eval, EAGLE=3
gsm8k: Per-position acceptance rate: 0.876, 0.752, 0.626, Avg Draft acceptance rate: 75.1%
Target model: MiniMaxAI/MiniMax-M3-MXFP8
Draft model: Inferact/MiniMax-M3-EAGLE3
Submission Checklist