fix: use the prompt cache for Gemma 4 text requests on VLM loads - #202
Merged
Merged
Conversation
Gemma 4 26B A4B auto-loads as a VLM, and every VLM load skipped the prompt cache, so each text turn re-prefilled its full conversation. - Allow the cache for MLXVLM.Gemma4 text requests (its text path is stateless); Qwen-VL & co. stay skipped. Video inputs now also count as multimodal. - Slice the input along the sequence axis (keeps [1, T] rank and the mask). - Detach saved/restored RotatingKVCache buffers: decode writes through them. - For caches with sliding-window layers, snapshot synchronously at the turn boundary instead of one decode token late, and miss when a wrapped ring would need a trim deeper than one slot. - Replace the tautological skip-predicate tests with the real function. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
solderzzc
added a commit
to CodeAndCanvas728/SwiftLM
that referenced
this pull request
Oct 4, 2026
…SharpAI#202 Gemma 4 VLM text-only requests are now cached, so the flag only matters for other VLM families. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #200
Problem
Gemma 4 26B A4B auto-loads as a VLM, and every VLM load skipped the prompt cache, so each text turn re-prefilled its whole conversation.
Changes
MLXVLM.Gemma4(its text-only path is stateless). Qwen-VL and the other VLM/Omni models stay skipped, since they need the LMOutput state (ropeDeltas) of the cached prefix. Video inputs now also count as multimodal.[1, T]rank and the mask (an axis-0 slice cut the batch axis on a hit).RotatingKVCachesnapshots on save and restore: decode steps write into the ring buffers in place, which corrupted saved snapshots.<|turn>/<|im_start|>) instead of one decode token late. A wrapped ring that would need a trim deeper than one slot is treated as a miss.onPrefillDoneno longer saves for requests that skip the cache (it previously also saved multimodal prompts).shouldSkipPromptCachefunction; add ring-buffer, slicing and boundary tests.Behaviour change
Non-Gemma LLMs with a wrapped sliding-window cache (
--ctx-size) may now miss where the old late save produced a misaligned snapshot.Testing
PromptCacheTests+PromptCacheRotatingTests: 23 tests pass. With the detach disabled, two of the ring tests fail.--audio: image request x2, audio request x2, then text x2. The only HIT is the repeated text request; nothing hits after an image or audio request.--ctx-size 2048(RotatingKVCache), ~1.8k-token prompt (ring not wrapped): HIT on repeat and on turn 2, and the turn-2 output equals a cold run.--prefill-size(256 vs 512/1024), so this is the sliding window's sensitivity to prefill chunking rather than a restore error. This was inferred from outputs, not by comparing logits.AI disclosure
Written with AI assistance (Claude Code, Claude Sonnet 5.5) and reviewed by AI agents; the author has not yet reviewed it line by line.
🤖 Generated with Claude Code