[Klaud Cold] Update minimaxm3-fp4-b300-vllm-agentic-mtp vLLM image to nightly-8a728663c1c3eeace834a95f5654fa653cc1998c and move to cluster:b300-dsxe - #2883
Conversation
25e3f5b to
51d7399
Compare
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
2 similar comments
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34177505317 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34182377798 |
… nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 and move to cluster:b300-dsxe Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…breaks the FA4 CuTe fp8-KV descale path used by the EAGLE3 draft on Blackwell Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
769479d to
3f0bc9a
Compare
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34182377798 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34255435497 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34255435497 |
Summary
Update the vLLM image for
minimaxm3-fp4-b300-vllm-agentic-mtpfromvllm/vllm-openai:nightly-1dc464d42681d22f38caf1fdc1eb632dc4421c45tovllm/vllm-openai:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36, and move the recipe from the retiredcluster:b300-nvrunner tocluster:b300-dsxe.sha256:254eebf919e8b7b0d530d97fccc606c36f380ff6724d64bc190951bec1aee838; same tag as the B200/H100/H200 MiniMax-M3 bumps ([Klaud Cold] Update minimaxm3-fp4-b200-vllm-agentic-mtp vLLM image to nightly-8a728663c1c3eeace834a95f5654fa653cc1998c #2860, [Klaud Cold] Update minimaxm3-fp8-h100-vllm-agentic-mtp vLLM image to nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 #2874, [Klaud Cold] Update minimaxm3-fp8-h200-vllm-agentic-mtp vLLM image to nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 #2875). The script does not set VLLM_USE_BREAKABLE_CUDAGRAPH=0, so it is not exposed to the piecewise-graph init failure seen on the ROCm gfx942 siblings.cluster:b300-nvwas retired in feat(runners): add the B300 DSXE cluster and retire the B300 NV launcher / 添加 B300 DSXE 集群并下线 B300 NV 启动脚本 #2826 (launcher and runner labels removed), so a sweep on it can never be scheduled. Repointed tocluster:b300-dsxe, the same move [AgentX] update B300 GLM image and extend concurrency coverage / 更新 B300 GLM 镜像并扩展并发覆盖 #2829 makes for the GLM-5.2 FP4 B300 sibling. Because the runner changed this is not an append-only bump; the whole curve reruns on DSXE.Recipes touched:
minimaxm3-fp4-b300-vllm-agentic-mtpTest plan
🤖 Generated with Claude Code
Note
Low Risk
Benchmark-only config and changelog updates; the main operational impact is rerunning the full curve on DSXE after the runner change, not production serving paths.
Overview
Updates
minimaxm3-fp4-b300-vllm-agentic-mtpso it can run again and stay on a working vLLM build on Blackwell.The recipe moves from
cluster:b300-nvtocluster:b300-dsxebecause the NV fleet was retired (#2826); that is not an append-only image bump—the full agentic-coding sweep reruns on DSXE. Scenario grid, EAGLE3/MTP settings, andbenchmarks/single_node/agentic/minimaxm3_fp4_b300_mtp.share unchanged.The vLLM image is updated from the old 2026-08-30 nightly to
vllm/vllm-openai:nightly-8a728663c1c3eeace834a95f5654fa653cc1998c(2026-09-04). Newer nightlies from 2026-09-05+ pull a flash-attn sync that breaks EAGLE3 + fp8 KV on B300 during CUDA-graph profiling, so this PR re-pins to the newest nightly still on the prior FA pin (same commit as the MI355X MiniMax AgentX recipe).perf-changelog.yamlrecords the image and runner changes for theagentic-codingscenario.Reviewed by Cursor Bugbot for commit 3f0bc9a. Bugbot is set up for automated code reviews on this repo. Configure here.
Update: re-pinned to the 2026-09-04 nightly
The 2026-09-07 nightly failed at engine init on B300 (run 34168437161, TP4 vllm-simple c36): during CUDA-graph memory profiling the EAGLE3 draft's FLASH_ATTN backend takes the FA4 CuTe path on Blackwell and, with the fp8 KV cache, its descale tensors fail
to_cute_tensorwithRuntimeError: Expected strides[leading_dim] == 1, but got 0. Cause: vllm-project/vllm@4ee259551 ("Sync FA with upstream", #54819, 2026-09-05) moved vllm-flash-attn from 06bdd47c to 506341a1; every nightly from 2026-09-05 on carries it, and no fix has landed on vllm main as of 2026-09-08T03:00Z. Re-pinned tovllm/vllm-openai:nightly-8a728663c1c3eeace834a95f5654fa653cc1998c(2026-09-04, pushed 2026-09-04T06:18:14Z, digestsha256:f5df5cc3302b5f404848c4eca88d7bf7ed5226e151c056da22816d7734644d67): the newest nightly still on FA 06bdd47c, five days newer than the recipe's previous pin, and the same vllm commit the MI355X MiniMax-M3 vLLM AgentX recipe already runs. Recipe otherwise unchanged. The Hopper (#2874/#2875) and ROCm (#2872/#2873) MiniMax siblings do not take the FA4 CuTe path and stay on the 09-07 nightly.