Skip to content

fix: use the prompt cache for Gemma 4 text requests on VLM loads - #202

Merged
solderzzc merged 1 commit into
mainfrom
fix/gemma4-vlm-prompt-cache
Oct 4, 2026
Merged

solderzzc merged 1 commit into
mainfrom
fix/gemma4-vlm-prompt-cache

Conversation

@solderzzc

@solderzzc solderzzc commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Fixes #200

Problem

Gemma 4 26B A4B auto-loads as a VLM, and every VLM load skipped the prompt cache, so each text turn re-prefilled its whole conversation.

Changes

  • Use the prompt cache for text requests on MLXVLM.Gemma4 (its text-only path is stateless). Qwen-VL and the other VLM/Omni models stay skipped, since they need the LMOutput state (ropeDeltas) of the cached prefix. Video inputs now also count as multimodal.
  • Slice the input along the sequence axis, keeping the [1, T] rank and the mask (an axis-0 slice cut the batch axis on a hit).
  • Detach RotatingKVCache snapshots on save and restore: decode steps write into the ring buffers in place, which corrupted saved snapshots.
  • For caches with sliding-window layers, snapshot synchronously at the turn boundary (last <|turn> / <|im_start|>) instead of one decode token late. A wrapped ring that would need a trim deeper than one slot is treated as a miss.
  • onPrefillDone no longer saves for requests that skip the cache (it previously also saved multimodal prompts).
  • Replace the tautological skip-predicate tests with the real shouldSkipPromptCache function; add ring-buffer, slicing and boundary tests.

Behaviour change

Non-Gemma LLMs with a wrapped sliding-window cache (--ctx-size) may now miss where the old late save produced a misaligned snapshot.

Testing

  • PromptCacheTests + PromptCacheRotatingTests: 23 tests pass. With the detach disabled, two of the ring tests fail.
  • gemma-4-e2b-it-4bit and gemma-4-26b-a4b-it-4bit, ~2.5k / ~4.3k token system prompt, two turns, temperature 0: HIT on repeat and on turn 2, and the turn-2 output equals a cold run on a fresh server.
  • gemma-4-e2b with --audio: image request x2, audio request x2, then text x2. The only HIT is the repeated text request; nothing hits after an image or audio request.
  • Qwen2.5-0.5B-Instruct-4bit with --ctx-size 2048 (RotatingKVCache), ~1.8k-token prompt (ring not wrapped): HIT on repeat and on turn 2, and the turn-2 output equals a cold run.
  • Same model, ~2.5k-3k-token prompt (ring wrapped): HITs occur, but the cached turn-2 output differs from a cold run. A cold run alone also changes with --prefill-size (256 vs 512/1024), so this is the sliding window's sensitivity to prefill chunking rather than a restore error. This was inferred from outputs, not by comparing logits.
  • CI: build_and_unit_test and the full integration matrix pass.

AI disclosure

Written with AI assistance (Claude Code, Claude Sonnet 5.5) and reviewed by AI agents; the author has not yet reviewed it line by line.

  • I have read this PR description

🤖 Generated with Claude Code

Gemma 4 26B A4B auto-loads as a VLM, and every VLM load skipped the prompt
cache, so each text turn re-prefilled its full conversation.

- Allow the cache for MLXVLM.Gemma4 text requests (its text path is stateless);
  Qwen-VL & co. stay skipped. Video inputs now also count as multimodal.
- Slice the input along the sequence axis (keeps [1, T] rank and the mask).
- Detach saved/restored RotatingKVCache buffers: decode writes through them.
- For caches with sliding-window layers, snapshot synchronously at the turn
  boundary instead of one decode token late, and miss when a wrapped ring
  would need a trim deeper than one slot.
- Replace the tautological skip-predicate tests with the real function.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@solderzzc
solderzzc merged commit f2409ce into main Oct 4, 2026
14 checks passed
solderzzc added a commit to CodeAndCanvas728/SwiftLM that referenced this pull request Oct 4, 2026
…SharpAI#202

Gemma 4 VLM text-only requests are now cached, so the flag only matters for
other VLM families.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Gemma 4 26B A4B is automatically detected as VLM even without --vision/--audio, preventing prompt KV-cache in text-only usage

1 participant