Skip to content

Live evaluation drops standalone usage metadata, so token_usage_v1 reports n/a聽#7321

Description

@yang0228

馃敶 Required Information

Describe the Bug:

Live usage metadata reaches ADK events but is discarded when those events are converted into evaluation invocations. As a result, token_usage_v1 reports an unavailable score (None) despite the model having reported token counts.

GeminiLlmConnection.receive() emits usage as a separate LlmResponse without content. EvaluationGenerator.convert_events_to_eval_invocations() retains content-bearing events and standalone grounding metadata, but not standalone usage metadata.

Steps to Reproduce:

  1. Use main at 044a1ec3f434cf2e3f7acfdbd52305b60d16f6e5 with the project's evaluation dependencies installed.
  2. Save the code below as repro_live_usage_only.py.
  3. Run PYTHONPATH=/path/to/adk-python/src python repro_live_usage_only.py.
  4. Observe that the input has one usage event, the converted invocation has none, and the assertion expecting 15 tokens fails.

The reproduction replaces only the network transport with local Gemini Live protocol messages. It runs the real ADK connection, event conversion, and token evaluator; no credentials or model service are required.

Expected Behavior:

The reported 10 prompt tokens and 5 response tokens survive conversion. token_usage_v1 reports 15, while the final answer remains Hello.

Observed Behavior:

The standalone usage event is dropped and the score is None.

{"input_usage_events": 1, "retained_usage_events": 0, "expected_tokens": 15, "actual_tokens": null}
AssertionError: Expected 15 tokens, got None

Environment Details:

  • ADK: source version 2.10.0, main commit 044a1ec3f434cf2e3f7acfdbd52305b60d16f6e5 (2026-09-26), loaded through PYTHONPATH.
  • OS: macOS 26.6.2, arm64.
  • Python: 3.12.13.
  • Dependencies: google-genai==2.24.0, pydantic==2.13.5.

Model Information:

  • LiteLLM: No.
  • Model: no live model call; gemini-live-test is a synthetic identifier supplied to the real GeminiLlmConnection with a local transport stub. A live Gemini E2E run has not been performed.

馃煛 Optional Information

Regression:

Confirmed on the current main revision above. A prior release has not been bisected.

Logs: Included above.

Screenshots / Video: N/A; deterministic Python reproduction.

Additional Context:

Minimal Reproduction Code:

import asyncio
import json

from google.genai import types
from google.adk.events import Event
from google.adk.models.gemini_llm_connection import GeminiLlmConnection
from google.adk.evaluation.evaluation_generator import EvaluationGenerator
from google.adk.evaluation._efficiency_evaluators import _TokenUsageV1Evaluator


class LocalTransport:
    session_id = "local-test"

    async def receive(self):
        yield types.LiveServerMessage(
            server_content=types.LiveServerContent(
                model_turn=types.Content(
                    role="model", parts=[types.Part(text="Hello")]
                )
            )
        )
        yield types.LiveServerMessage(
            usage_metadata=types.UsageMetadata(
                prompt_token_count=10,
                response_token_count=5,
                total_token_count=15,
            )
        )
        yield types.LiveServerMessage(
            server_content=types.LiveServerContent(turn_complete=True)
        )


async def main():
    connection = GeminiLlmConnection(
        LocalTransport(), model_version="gemini-live-test"
    )
    events = [Event(
        author="user", invocation_id="inv1",
        content=types.Content(role="user", parts=[types.Part(text="Hi")]),
    )]
    async for response in connection.receive():
        events.append(Event(
            author="agent", invocation_id="inv1",
            **response.model_dump(exclude_none=True),
        ))

    invocations = EvaluationGenerator.convert_events_to_eval_invocations(events)
    score = _TokenUsageV1Evaluator().evaluate_invocations(invocations).overall_score
    print(json.dumps({
        "input_usage_events": sum(e.usage_metadata is not None for e in events),
        "retained_usage_events": sum(
            e.usage_metadata is not None
            for i in invocations
            for e in i.intermediate_data.invocation_events
        ),
        "expected_tokens": 15,
        "actual_tokens": score,
    }))
    assert score == 15, f"Expected 15 tokens, got {score!r}"


asyncio.run(main())

How often has this issue occurred?:

Always in the local reproduction. The original connection-to-evaluator check was run twice with consistent results; the isolated reproduction above also fails as shown.

Contribution / assignment request

I would like to work on this issue. Is anyone already addressing it? If the scope is appropriate, could a maintainer assign it to @yang0228?

My proposed PR would preserve standalone usage metadata during evaluation conversion and add regression coverage for usage-only events, ordinary content-bearing usage, and absent usage, while preserving final-response selection. I will follow the required test and validation workflow. My Google CLA is already signed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

live[Component] This issue is related to live, voice and video chat

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions