chore: bump mlx-swift-lm to 9d5041a for the MTP rollback fixes (#184) - #201
Merged
Merged
Conversation
Picks up SharpAI/mlx-swift-lm#72: rejected MTP drafts are rolled back layer by layer, including sliding-window layers after the ring wraps (#184). The range also has SharpAI/mlx-swift-lm#68 (upstream #603 sync, not used by SwiftLM). README: replace the "avoid --mtp" warning with the measured result. Rollback is exact; the remaining differences from plain decoding are batch-shape bf16 numerics and the text-only vs vision model path. Fixes #184 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
solderzzc
force-pushed
the
chore/bump-mlx-swift-lm-mtp-rollback
branch
from
September 29, 2026 21:42
fc54d97 to
41e5ce4
Compare
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bumps
mlx-swift-lmfrom 5596071 to 9d5041a to pick up SharpAI/mlx-swift-lm#72, and rewords the #184 note in the README.SharpAI/mlx-swift-lm#72 rolls back rejected MTP drafts layer by layer, including sliding-window layers after the ring has wrapped. Before it, Gemma 4 with
--mtpdrifted from the plain model's temperature-0 output once the context passed the 1,024-token window (#184). It also:kv_bits) plus the Gemma 4 assistant--draft-modeland--mtpThe other PR in the range is SharpAI/mlx-swift-lm#68. It merges upstream ml-explore/mlx-swift-lm#603 (model cache extraction in
MLXFoundationModels, which SwiftLM does not use) and ml-explore/mlx-swift-lm#620, which is already in our fork and adds no diff.Output check (M5 Pro,
mlx-community/gemma-4-26b-a4b-it-4bit)Each config generated 200 tokens at temperature 0 on the five
m6_benchprompts, with--ctx-size 16384. The prompts are 560, 2,590, 2,627, 11,111 and 11,139 tokens (m6_bench targets ~530 / 2.3K / 9.5K). Each output was then replayed through the plain model in mlx-lm. For every token we measured how far its logit is below the top logit at that position (0 = the plain model's greedy choice).--mtp--mtp, bf16 assistant--mtp, QAT 4-bit assistantWith the bf16 assistant, every pick is now within one logit of the top (max 0.875, against 0.625 without
--mtp). The QAT assistant has one pick above that in each 11.1K prompt (1.75 and 1.375).These are numerics, not a rollback bug:
--mtploads Gemma 4 as a text-only LLM, while the run without it loads the vision model, so the two runs already use different kernels.The output is still not token-identical to plain decoding. With either assistant, all five prompts leave SwiftLM's own no-
--mtptext at some near-tie (tokens 21–125 for bf16). The README now says this and gives the measured gaps.Speed
There is no consistent change in decode speed on M5 Pro. Decode tok/s at 560 / 2.6K / 2.6K / 11.1K / 11.1K tokens:
--mtp--mtp, bf16--mtp, QATEach cell is a single run, and single runs vary a lot. Only no
--mtpat 560 tokens was repeated: four requests gave 69.2–70.4 tok/s on the current pin and 69.7–70.4 on this PR, so the 44.5 was a one-off. The M6 numbers in the README were not re-measured.Testing
swift build -c releasesucceeds with 9d5041a.kv_bitswith--mtp,--draft-model, and the SwiftLM test suite (left to CI).AI disclosure: the investigation, measurements and this PR were done with Claude Code (AI).
Fixes #184
🤖 Generated with Claude Code