Skip to content

fix(eval): count logical Live requests across response segments - #7427

Open
yang0228 wants to merge 13 commits into
google:mainfrom
yang0228:codex/7351-live-inference-count-draft
Open

yang0228 wants to merge 13 commits into
google:mainfrom
yang0228:codex/7351-live-inference-count-draft

Conversation

@yang0228

@yang0228 yang0228 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Link to Issue or Description of Change

Closes #7351. Depends on #7323; please land the usage fix first. The Live metric semantics were confirmed by the evaluation owner: one model generation/request contributes one inference, regardless of text or audio chunking. Acceptance of the optional saved-schema field remains part of this review.

Problem: A single Gemini Live generation counts as multiple inference calls when its answer is chunked, when thought switches to answer, or when text is flushed before a tool call. Filtering partial events does not identify these boundaries and can discard usage.

Solution: Reconstruct requests from the original event sequence before conversion drops completion markers. Track each author/Live connection independently; close on completion, interruption, tool response, or a new client message. Input transcription retains its connection identity and does not close a generation. Keep ordinary unary counting and SSE final-response boundaries, including mixed invocations. Stamp merged final text with model_version so persisted sessions that omit partial chunks retain their count. Content, tools, and usage remain available; no partial-event filter is used.

Review follow-up (2026-10-08):

  • Added a typed-client-message fixture with no transcription or live_session_id between an unfinished generation and a new model response: 2 calls, with the new user message and final answer retained.
  • Added output-audio coverage for identical bytes in 1 / 2 / 3 chunks, with and without standalone usage: 1 call, unchanged audio bytes, and 15 tokens when reported.
  • Replaced the usage-merge tag with a plain TODO in fix(eval): preserve live usage without extra inference calls #7323 at 3517a033b, then merged that prerequisite update here. The cleanup does not change production behavior.
  • The seven new regression cases are in 19b61d22b. Removing the typed-message boundary in an isolated source copy produces 1 != 2; restoring event-based counting produces 2/3 != 1 in four audio cases.

Incremental review and landing order

The GitHub diff targets main and includes #7323. The original incremental comparison isolates the original two logical-counting commits. Review the follow-up tests separately at 19b61d22b; subsequent merge commits synchronize upstream or the prerequisite without rewriting published history.

The production transcription fix remains a separate commit, bc5f75588, in postprocess_live_flow, as confirmed in the maintainer review. After #7323 is imported onto main, this PR needs a rebase to remove its inherited diff.

Persisted eval JSON compatibility

InvocationEvent does not retain partial, turn_complete, interrupted, or live_session_id; pure completion markers are discarded. The projected events cannot distinguish a segmented response from multiple real requests, so the count must be reconstructed before projection.

Optional Invocation.inference_call_count preserves the result through eval JSON roundtrips. Missing/None retains event-based fallback, and explicit 0 is respected. New readers can read older eval files; older strict readers may reject newly written JSON containing this field. Please confirm acceptance of this saved-schema change during review.

Legacy tool responses without connection identity retain author-level closure. Requests without observable boundaries remain ambiguous; this change does not infer billing.

Testing Plan

Fresh validation of head 19b61d22b64fe26f085836d631a33d8499881e7b, based on main 315c3dd2cd2710b3ccc2fbb88c22125938d89adf plus #7323:

  • 252 related tests passed, with 27 existing warnings, on Python 3.12.13 with both the existing google-genai 2.24.0 environment and the isolated 2.29.0 environment.
  • Configured pre-commit hooks passed, including the commit hooks.
  • Built a wheel and installed it with frozen test dependencies into a separate clean Python 3.12 environment. Copied the regression file outside the checkout: 7 passed. ADK imported from that environment's site-packages.
  • All four currently supported Python versions ran the full unit suite in isolated tox environments. All seven new cases passed in each run.
python -m pytest -q \
  tests/unittests/evaluation/test_evaluation_generator.py \
  tests/unittests/evaluation/test_efficiency_evaluators.py \
  tests/unittests/models/test_gemini_llm_connection.py \
  tests/unittests/live/test_live_llm_flow.py \
  tests/unittests/flows/llm_flows/test_base_llm_flow_realtime.py \
  tests/unittests/flows/llm_flows/test_base_llm_flow_partial_handling.py
# 252 passed, 27 warnings

The checkout has no uv.lock; tox needed temporary runner and dependency overrides. No repository configuration was changed:

tox -p 2 -e py311,py312,py313,py314 \
  -x 'testenv.runner=uv-venv-runner' \
  -x 'testenv.deps=.[test]' \
  -x 'testenv.commands=pytest tests/unittests -n 2 -q'
Python Passed Failed Skipped
3.11.15 17,987 3 90
3.12.13 17,971 10 91
3.13.14 17,972 9 91
3.14.6 17,973 8 91

Each run also reported 25 xfailed, 2 xpassed, and 28 passed subtests. These are not all-green suites. The isolated environments use google-genai 2.29.0 / google-cloud-aiplatform 2.4.0. Kubernetes resolves to 36.0.3 on Python 3.11 and 37.0.0 on the other versions.

All final failure IDs were freshly reproduced using unchanged main 315c3dd2c in the same respective environments:

  • Python 3.11: webpage file 26 passed / 3 failed.
  • Python 3.12: GKE, import-loading, and webpage files 53 passed / 10 failed.
  • Python 3.13: complete unchanged-main suite 17,912 passed / 9 failed, matching all nine final failure IDs. The isolated auto-tracing file passes 60/60 on both main and the PR.
  • Python 3.14: GKE and webpage files 39 passed / 8 failed.
Names and causes of the final full-suite failures

All four versions fail these existing macOS system-proxy cases in tests/unittests/tools/test_load_web_page.py:

  • test_load_web_page_allows_public_nat64_ip
  • test_load_web_page_fetches_public_urls_by_pinning_the_resolved_ip
  • test_load_web_page_tries_another_resolved_address_after_connect_error

Python 3.12–3.14 additionally fail these TestGkeCodeExecutor cases in tests/unittests/code_executors/test_gke_code_executor.py. Kubernetes 37 validates mock-valued V1OwnerReference fields as strings:

  • test_execute_code_success
  • test_execute_code_job_failed
  • test_execute_code_job_failed_without_terminated_container
  • test_execute_code_timeout
  • test_execute_code_forks_to_job

Python 3.12 additionally fails test_entry_point_loads_only_allowlisted_packages[agent] and [runner] in tests/unittests/test_import_loading.py, because Homebrew loads sitecustomize.

Python 3.13 additionally fails tests/unittests/plugins/test_auto_tracing_plugin.py::test_signature_introspection_happens_once_per_wrap. Pytest reports an event-loop destructor's unraisable RecursionError as a runtime error. The same full-suite failure was reproduced on unchanged main; isolated runs of that test file pass.

Earlier validation recorded the 35-case text/tool/sub-agent/usage diagnostic matrix on source and an independently installed wheel. Those are earlier-head results, not a fresh 35-case matrix for this follow-up. The new regressions exercise the real Live adapter, converter, and evaluators with deterministic transport; they do not establish remote Gemini ordering, billing semantics, or release integration.

Checklist

  • Read CONTRIBUTING.md and link the issue.
  • Follow the maintainer's requested review changes and add effective regression tests.
  • Run configured pre-commit hooks and all supported-version unit suites.
  • Build and verify an independently installed wheel.
  • Evaluation owner confirms Live counting semantics.
  • All full unit suites pass locally (unchanged-main failures documented above).
  • Manually test against a remote Gemini service.
  • fix(eval): preserve live usage without extra inference calls #7323 has landed.
  • Evaluation owners accept the optional saved-schema field.

Merge standalone Live usage into compatible model events during evaluation
conversion, retaining unmatched reports without mutating session events.

Fixes google#7321
Keep input and output transcription events tied to their Live connection so normalized user audio does not appear to start another model request.
Reconstruct request boundaries before event conversion drops completion and connection metadata. Persist the optional count so saved eval sets retain it, while old files and unary invocations keep event-based fallback.

Fixes google#7351
@yang0228
yang0228 marked this pull request as ready for review October 7, 2026 06:35
Replace the implementation-specific tag with a TODO describing the existing quadratic scan and indexing boundary. No counting or usage behavior changes.
Pin the client-message boundary during an unfinished generation at two calls. Verify one, two and three audio chunks stay at one inference with or without standalone usage, preserving audio and token totals.

Refs google#7351
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Live inference_call_count_v1 changes with text chunk boundaries

2 participants