Skip to content

Commit 9ae4705

Browse files
akshayliveclaude
andcommitted
fix(agents): Delegate token usage was silently zero on every live turn
test_delegate_live_token_usage_populated failed live (crashed=False, real text/tool output, token_usage=None): _parse_usage was never even reached with a dict, because it was looking in the wrong place entirely. Confirmed by reading the installed @uipath/delegate-sdk's bundled dist/index.mjs directly: sendMessage() always resolves to a plain string, never an object, and no event forwarded through agent.onEvent() ever carries a `usage` field. The SDK's per-turn token accounting lives only in its internal store, reachable through DelegateAgent.getLastTurnUsage() (and the session id through getSessionId()) -- called nowhere in this agent before now. - delegate_host.mjs: call both getters right after sendMessage() resolves and attach them to the send_ok message (usage, sessionId). - delegate_agent.py: _handle_send_ok reads them off send_ok's own top level instead of a nonexistent nested result dict. _parse_usage now reads the getter's real (confirmed, not guessed) shape -- promptTokens/completionTokens/promptTokensCached/cacheCreationTokens -- computing uncached_input_tokens = promptTokens - promptTokensCached (promptTokens is the OpenAI-style total, promptTokensCached the cache-read subset). - Updated tests to the real send_ok shape and confirmed bucket names. - .claude/notes/agents.md § Delegate agent records the finding so it isn't re-litigated as a guess next time. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
1 parent 52ea560 commit 9ae4705

4 files changed

Lines changed: 112 additions & 57 deletions

File tree

‎.claude/notes/agents.md‎

Lines changed: 20 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -809,6 +809,26 @@ own env-var-driven bootstrapping, not the `DelegateAgent` class itself, which we
809809
against a real backend). `tests/test_delegate_agent.py::TestStart::test_org_and_tenant_slug_forwarded_into_auth`
810810
pins it. Passing `DELEGATE_BACKEND_URL` directly instead of `DELEGATE_ENV` skips this whole path.
811811

812+
**Token usage comes from two getters called after `sendMessage()` resolves, not from any event or the
813+
resolved value itself — confirmed by reading the installed SDK's bundle, not by guessing.** A first
814+
pass guessed at a flat `usage` dict carried on the resolved `sendMessage()` value or on a forwarded
815+
event, tried several plausible snake_case/camelCase bucket-name spellings, and silently returned zero
816+
tokens every turn against a real backend (`tests/test_delegate_agent_live.py::test_delegate_live_token_usage_populated`
817+
failed live: `record.crashed is False`, real text/tool output, `token_usage=None`). Reading
818+
`node_modules/@uipath/delegate-sdk/dist/index.mjs` directly settled it: `sendMessage()` always resolves
819+
to a plain string (never an object), and no event this host forwards via `agent.onEvent()` ever carries
820+
a `usage` field — the SDK's own internal event vocabulary has no `"usage"` member (that name IS used
821+
internally, but only inside the SDK's own Zustand store reducer that updates its "Token usage" UI
822+
panel, never re-emitted through the public `onEvent` bus). The only way to reach it is
823+
`DelegateAgent.getLastTurnUsage(sessionId?)` (defaults to the just-used session), so
824+
`delegate_host.mjs`'s `handleSend` calls it — and `getSessionId()` — right after `sendMessage()`
825+
resolves, and attaches both to the `send_ok` message itself. That getter's shape, read straight off the
826+
SDK's own `setUsage` store action, is `{promptTokens, completionTokens, promptTokensCached,
827+
cacheCreationTokens, turnTokenUnits, contextBreakdown}` — `promptTokens` is the TOTAL input token count
828+
(OpenAI-style, cached + uncached), `promptTokensCached` the cache-READ subset of it, so
829+
`_parse_usage` computes `uncached_input_tokens = promptTokens - promptTokensCached`. This is now
830+
CONFIRMED, not a guess, so `_parse_usage`'s docstring no longer marks it `# UNVERIFIED`.
831+
812832
### Delegate agent pricing
813833

814834
Delegate-served models are keyed under the SDK's own hyphenated ids (e.g. `gpt-5-6-terra`) in

‎src/coder_eval/agents/delegate/delegate_host.mjs‎

Lines changed: 19 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -21,7 +21,13 @@
2121
// stdout (host -> coder_eval):
2222
// {"type": "init_ok"}
2323
// {"type": "init_error", "message": str}
24-
// {"type": "send_ok", "result": <sendMessage()'s resolved value, or null>}
24+
// {"type": "send_ok", "result": <sendMessage()'s resolved string, or null>,
25+
// "usage": <agent.getLastTurnUsage()'s result, or null>,
26+
// "sessionId": <agent.getSessionId()'s result, or null>}
27+
// -- usage/sessionId are read from these two getters AFTER sendMessage()
28+
// resolves, not from the resolved value itself (a plain string) or from
29+
// any forwarded event: CONFIRMED (reading the installed SDK's bundled
30+
// source) that no event this host forwards ever carries a `usage` field.
2531
// {"type": "send_error", "message": str}
2632
// {"type": "destroy_error", "message": str}
2733
// {"type": "protocol_error", "message": str} -- malformed/unknown stdin command; host keeps running
@@ -88,7 +94,18 @@ async function handleSend(msg) {
8894
return;
8995
}
9096
const result = await agent.sendMessage(msg.prompt, msg.sessionId || undefined);
91-
writeLine({ type: "send_ok", result: result ?? null });
97+
// CONFIRMED (reading @uipath/delegate-sdk@0.1.12's bundled dist/index.mjs):
98+
// sendMessage() resolves to a plain string (the final response text) --
99+
// never an object -- and no event this host forwards via agent.onEvent()
100+
// ever carries a `usage` field. The SDK's own per-turn token accounting is
101+
// internal state, reachable only through these two getters, called here
102+
// once the turn (and the backend's own internal "usage" store update) has
103+
// settled. Guarded with typeof, not called unconditionally: an older/newer
104+
// SDK build that drops either method must degrade to "no usage this turn",
105+
// never crash the host.
106+
const usage = typeof agent.getLastTurnUsage === "function" ? agent.getLastTurnUsage() : null;
107+
const sessionId = typeof agent.getSessionId === "function" ? agent.getSessionId() : null;
108+
writeLine({ type: "send_ok", result: result ?? null, usage: usage ?? null, sessionId: sessionId ?? null });
92109
}
93110

94111
async function handleDestroy() {

‎src/coder_eval/agents/delegate_agent.py‎

Lines changed: 45 additions & 39 deletions
Original file line numberDiff line numberDiff line change
@@ -11,9 +11,9 @@
1111
clear ``AgentConfigError`` at ``start()`` (Node.js, ``npm install
1212
@uipath/delegate-sdk``, UiPath auth). Deliberate scope reductions versus the
1313
UiPath-internal sibling agent's more hardened adapter (no multi-generation
14-
transcript splitting, no WAF/SSE/session-conflict/stall-resend recovery,
15-
best-effort token-bucket field names marked ``# UNVERIFIED``) and every
16-
``# UNVERIFIED`` spot's rationale live in one place, not scattered:
14+
transcript splitting, no WAF/SSE/session-conflict/stall-resend recovery) and
15+
every remaining ``# UNVERIFIED`` spot's rationale live in one place, not
16+
scattered:
1717
1818
Rationale: .claude/notes/agents.md § Delegate agent
1919
"""
@@ -206,39 +206,46 @@ def _resolve_bundled_skills_path(plugins: list[dict[str, Any]] | None) -> str |
206206

207207

208208
def _parse_usage(raw: Any) -> TokenUsage | None:
209-
"""Best-effort parse of an event/result's ``usage`` payload into ``TokenUsage``.
210-
211-
UNVERIFIED: the SDK confirms an ``usage`` field exists on at least some
212-
events, but not its internal bucket key names. Tries several plausible
213-
spellings (snake_case, as the internal sibling agent's protocol used; and
214-
camelCase, in case this layer differs) and falls back to 0 for anything it
215-
cannot read, mirroring the project's "warn on drift, never raise" contract.
209+
"""Parse the SDK's per-turn usage payload (from ``getLastTurnUsage()``) into ``TokenUsage``.
210+
211+
CONFIRMED (reading the installed ``@uipath/delegate-sdk@0.1.12``'s bundled
212+
``dist/index.mjs``): no event this host forwards ever carries a ``usage``
213+
field -- the SDK's per-turn token accounting lives only in its internal
214+
store, reachable through ``DelegateAgent.getLastTurnUsage()``, which
215+
``delegate_host.mjs`` calls after ``sendMessage()`` resolves and attaches
216+
to the ``send_ok`` message as ``usage``. That getter's shape, from the
217+
SDK's own ``setUsage`` store action: ``{promptTokens, completionTokens,
218+
promptTokensCached, cacheCreationTokens, turnTokenUnits,
219+
contextBreakdown}``. ``promptTokens`` is the TOTAL input token count
220+
(cached + uncached, OpenAI-style); ``promptTokensCached`` is the
221+
cache-READ subset of it, so ``uncached = promptTokens - promptTokensCached``.
222+
Falls back to 0 for anything absent (e.g. before the backend's first
223+
internal usage report), mirroring the project's "warn on drift, never
224+
raise" contract -- a future SDK release renaming one of these fields
225+
degrades to zero tokens for that bucket, not a crash.
216226
"""
217227
if not isinstance(raw, dict):
218228
return None
219229

220-
def _int(*keys: str) -> int:
221-
for key in keys:
222-
value = raw.get(key)
223-
if isinstance(value, bool):
224-
continue
225-
if isinstance(value, int) and value >= 0:
226-
return value
227-
return 0
228-
229-
input_tokens = _int("input_tokens", "inputTokens", "uncached_input_tokens")
230-
output_tokens = _int("output_tokens", "outputTokens")
231-
cache_creation = _int("cache_creation_input_tokens", "cacheCreationInputTokens", "cache_write", "cacheWrite")
232-
cache_read = _int("cache_read_input_tokens", "cacheReadInputTokens", "cache_read", "cacheRead")
233-
if input_tokens == 0 and output_tokens == 0 and cache_creation == 0 and cache_read == 0:
230+
def _int(key: str) -> int:
231+
value = raw.get(key)
232+
if isinstance(value, bool):
233+
return 0
234+
return value if isinstance(value, int) and value >= 0 else 0
235+
236+
prompt_total = _int("promptTokens")
237+
prompt_cached = _int("promptTokensCached")
238+
output_tokens = _int("completionTokens")
239+
cache_creation = _int("cacheCreationTokens")
240+
if prompt_total == 0 and output_tokens == 0 and prompt_cached == 0 and cache_creation == 0:
234241
if raw:
235242
logger.warning("delegate: usage payload matched none of the known bucket spellings: %r", sorted(raw))
236243
return None
237244
return TokenUsage(
238-
uncached_input_tokens=input_tokens,
245+
uncached_input_tokens=max(prompt_total - prompt_cached, 0),
239246
output_tokens=output_tokens,
240247
cache_creation_input_tokens=cache_creation,
241-
cache_read_input_tokens=cache_read,
248+
cache_read_input_tokens=prompt_cached,
242249
)
243250

244251

@@ -795,22 +802,21 @@ def _handle_tool_result(self, msg: dict[str, Any], state: _TurnState, emit: Call
795802
emit(ToolEndEvent(task_id=self.task_id, turn_id=state.turn_id, tool=telemetry, status=status))
796803

797804
def _handle_send_ok(self, msg: dict[str, Any], state: _TurnState) -> None:
805+
# CONFIRMED (reading the installed SDK's bundled source):
806+
# sendMessage()'s resolved value is always a plain string, never an
807+
# object -- `usage`/`sessionId` are NOT nested under it. delegate_host.mjs
808+
# instead reads them off `getLastTurnUsage()`/`getSessionId()` after
809+
# sendMessage() resolves and attaches them to this message's own
810+
# top level (see delegate_host.mjs's wire-protocol header comment).
798811
result = msg.get("result")
799812
if isinstance(result, str):
800813
state.final_response = result
801-
elif isinstance(result, dict):
802-
response = result.get("response") or result.get("content")
803-
if isinstance(response, str):
804-
state.final_response = response
805-
session_id = result.get("sessionId")
806-
if isinstance(session_id, str) and session_id:
807-
self._session_id = session_id
808-
usage = _parse_usage(result.get("usage"))
809-
if usage is not None:
810-
state.usage = usage
811-
model = result.get("model")
812-
if isinstance(model, str) and model:
813-
state.model_used = model
814+
session_id = msg.get("sessionId")
815+
if isinstance(session_id, str) and session_id:
816+
self._session_id = session_id
817+
usage = _parse_usage(msg.get("usage"))
818+
if usage is not None:
819+
state.usage = usage
814820

815821
def _close_open_tools(self, state: _TurnState, emit: Callable[[StreamEvent], None]) -> None:
816822
for tool_id, telemetry in list(state.open_tools.items()):

‎tests/test_delegate_agent.py‎

Lines changed: 28 additions & 16 deletions
Original file line numberDiff line numberDiff line change
@@ -251,7 +251,7 @@ async def test_happy_path_text_and_tool(self, patch_exec, tmp_path):
251251
_line({"type": "message", "content": "here is my answer"}),
252252
_line({"type": "tool_call", "toolName": "Bash", "toolId": "tool-1", "input": {"command": "ls"}}),
253253
_line({"type": "tool_result", "toolId": "tool-1", "output": "file.txt"}),
254-
_line({"type": "send_ok", "result": {"response": "here is my answer", "sessionId": "sess-1"}}),
254+
_line({"type": "send_ok", "result": "here is my answer", "sessionId": "sess-1"}),
255255
]
256256
agent, proc = await _started_agent(patch_exec, events, tmp_path)
257257
record = await agent.communicate("do something")
@@ -422,45 +422,57 @@ def on_event(self, event: Any) -> None:
422422
assert record.crashed is False
423423

424424
@pytest.mark.parametrize(
425-
("usage_payload", "expected_input", "expected_output"),
425+
("usage_payload", "expected_uncached_input", "expected_output", "expected_cache_read", "expected_cache_write"),
426426
[
427-
({"input_tokens": 10, "output_tokens": 5}, 10, 5),
428-
({"inputTokens": 10, "outputTokens": 5}, 10, 5),
429-
({"uncached_input_tokens": 10, "output_tokens": 5}, 10, 5),
427+
({"promptTokens": 10, "completionTokens": 5}, 10, 5, 0, 0),
428+
({"promptTokens": 10, "completionTokens": 5, "promptTokensCached": 4}, 6, 5, 4, 0),
429+
(
430+
{"promptTokens": 10, "completionTokens": 5, "promptTokensCached": 4, "cacheCreationTokens": 3},
431+
6,
432+
5,
433+
4,
434+
3,
435+
),
430436
],
431437
)
432438
async def test_usage_bucket_spellings_populate_token_usage(
433-
self, patch_exec, tmp_path, usage_payload, expected_input, expected_output
439+
self,
440+
patch_exec,
441+
tmp_path,
442+
usage_payload,
443+
expected_uncached_input,
444+
expected_output,
445+
expected_cache_read,
446+
expected_cache_write,
434447
):
435-
events = [_line({"type": "send_ok", "result": {"response": "done", "usage": usage_payload}})]
448+
events = [_line({"type": "send_ok", "result": "done", "usage": usage_payload})]
436449
agent, _ = await _started_agent(patch_exec, events, tmp_path)
437450
record = await agent.communicate("hi")
438451
assert record.token_usage is not None
439-
assert record.token_usage.uncached_input_tokens == expected_input
452+
assert record.token_usage.uncached_input_tokens == expected_uncached_input
440453
assert record.token_usage.output_tokens == expected_output
454+
assert record.token_usage.cache_read_input_tokens == expected_cache_read
455+
assert record.token_usage.cache_creation_input_tokens == expected_cache_write
441456

442457
async def test_usage_all_zero_is_none_and_warns(self, patch_exec, tmp_path, caplog):
443-
events = [_line({"type": "send_ok", "result": {"response": "done", "usage": {"weird_bucket": 3}}})]
458+
events = [_line({"type": "send_ok", "result": "done", "usage": {"weird_bucket": 3}})]
444459
agent, _ = await _started_agent(patch_exec, events, tmp_path)
445460
with caplog.at_level("WARNING"):
446461
record = await agent.communicate("hi")
447462
assert record.token_usage is None
448463
assert any("usage payload matched none" in r.message for r in caplog.records)
449464

450-
async def test_model_and_cost_wired_into_record(self, patch_exec, tmp_path):
465+
async def test_cost_wired_into_record_for_configured_model(self, patch_exec, tmp_path):
451466
events = [
452467
_line(
453468
{
454469
"type": "send_ok",
455-
"result": {
456-
"response": "done",
457-
"model": "virtuoso-1-5",
458-
"usage": {"input_tokens": 1_000_000, "output_tokens": 1_000_000},
459-
},
470+
"result": "done",
471+
"usage": {"promptTokens": 1_000_000, "completionTokens": 1_000_000},
460472
}
461473
)
462474
]
463-
agent, _ = await _started_agent(patch_exec, events, tmp_path)
475+
agent, _ = await _started_agent(patch_exec, events, tmp_path, model="virtuoso-1-5")
464476
record = await agent.communicate("hi")
465477
assert record.model_used == "virtuoso-1-5"
466478
assert record.token_usage is not None

0 commit comments

Comments
 (0)