Skip to content

perf: bound the streaming stop scan, real usage counts, --no-token-echo - #203

Merged
solderzzc merged 1 commit into
SharpAI:mainfrom
CodeAndCanvas728:pr/stream-hot-path
Oct 4, 2026
Merged

solderzzc merged 1 commit into
SharpAI:mainfrom
CodeAndCanvas728:pr/stream-hot-path

Conversation

@CodeAndCanvas728

@CodeAndCanvas728 CodeAndCanvas728 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Three small fixes to the per-token request path in Server.swift.

Changes

1. Bounded stop-sequence scan. Both streaming handlers (chat and text completions) ran checkStopSequences on the whole accumulated response for every chunk, which is O(n²) in response length. The chat path always carries five built-in stop strings, so this cost applies to every chat request.

  • checkStopSequences now takes an optional lookback, and the streaming paths search only the text appended since their last scan plus the longest stop string, with one extra character for grapheme merges.
  • This is safe because everything earlier was already scanned and didn't match.
  • The chat path counts characters since its last scan rather than using the chunk length. JSON-mode buffering skips the scan for its first chunks, so the next scan has to cover them.
  • Non-streaming paths still scan once in full. The non-streaming chat path no longer calls the function twice.

2. usage.completion_tokens reports real tokens. All four handlers counted .chunk events. Tokens the decoder buffers, such as tool-call bodies and partial text, emit no chunk, so they were never counted. Each path now takes GenerateCompletionInfo.generationTokenCount when .info arrives. This also corrects timings.predicted_n, tok/s and the gen_tokens log line. An early stop-sequence finish still reports the chunk count, because .info never arrives on that path; there's a comment saying so.

3. --no-token-echo. This flag turns off the per-token print + fflush(stdout) in the chat handlers. The default is unchanged. It's documented in the README flags table.

Testing

  • swift test --filter SwiftLMTests: 210 tests, 0 failures. New StopSequenceTests cases:
    • Equivalence: the bounded scan gives the same result as the old full rescan, over 8 texts × 7 chunkings. The texts include CJK, emoji, ZWJ sequences, combining marks and overlapping stops.
    • A stop straddling two chunks is still found.
    • Chunks skipped by JSON-mode buffering are still covered.
    • A grapheme merge across a chunk boundary is handled.
  • Live smoke against mlx-community/Ministral-8B-Instruct-2410-4bit:
    • The streaming chat stop cuts cleanly with no leak.
    • The streaming text-completions stop hits correctly.
    • In a tool-call response where only one chunk reached the client, completion_tokens is now 20; the old code would have reported 1.
    • --no-token-echo leaves the srv log lines intact.

Built and tested with Xcode 27.0 (27A266a), using the submodule as pinned.

There is no CHANGELOG or version file in this repo, so nothing was bumped.

🤖 Generated with Claude Code

Three fixes on the per-token request path:

- Stop sequences: the streaming chat and text handlers re-scanned the
  whole accumulated response on every chunk (O(n^2) in response length;
  chat always carries five built-in stops). checkStopSequences takes an
  optional lookback and the streaming paths search only the text since
  their last scan plus the longest stop string. The chat path tracks
  characters since the last scan rather than the chunk length, because
  JSON-mode buffering skips the scan for its first chunks.
- usage.completion_tokens counted .chunk events, so tokens the decoder
  buffers (tool-call bodies, partial text) were never counted. All four
  paths now take GenerateCompletionInfo.generationTokenCount when .info
  arrives. An early stop-sequence finish still reports the chunk count,
  since .info never arrives there.
- --no-token-echo turns off the per-token print/fflush to stdout. Default
  is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014J4GsqznbKNX8tcKzNRxEr
@solderzzc

Copy link
Copy Markdown
Member

Reviewed (with AI assistance, Claude Code) and checked against current main (#202 is merged; no conflicts). I found no blocking issues: the bounded lookback matches a full rescan for grapheme merges, multi-byte text, split stops and JSON-mode buffering.

Two notes, neither blocking:

  • completion_tokens now reports the real generated-token count, so tool-call-only responses go from 0 to a real number. Worth a line in the release notes.
  • With --no-token-echo the srv generate: id 0 | prefix is still printed on its own line.

Merging.

@solderzzc
solderzzc merged commit af3236c into SharpAI:main Oct 4, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants