Skip to content

openai-chat: keep the client's cache breakpoints on system and tool_result - #148

Merged
CMGS merged 1 commit into
mainfrom
fix/openai-wire-cache-breakpoints
Oct 6, 2026
Merged

CMGS merged 1 commit into
mainfrom
fix/openai-wire-cache-breakpoints

Conversation

@CMGS

@CMGS CMGS commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

Problem

A /v1/messages request served by an openai-chat model lost its prompt-cache breakpoints:

  • system blocks were flattened into a string system message, which dropped their cache_control.
  • each tool_result was flattened into a string role: tool message, which dropped its cache_control.

An agent loop through OpenRouter to a Claude model puts a breakpoint on the system prompt and another on the trailing tool_result, and it never read its cache. Turns that ended on a user text block, whose marker did pass through, paid the 1.25× cache-write price. Turns that ended on a tool_result read nothing. The billed total_tokens running above prompt + completion was that write premium, not an overcount: cost matches OpenRouter's reported cost.

Fix

  • The system message now carries the native system blocks as content parts.
  • A tool_result with cache_control becomes a one-part text array that keeps its marker. An unmarked one stays a string.
  • Tool-definition markers are still dropped. Gemini's OpenAI-compatible endpoint rejects a cache_control field on a tool (400), and any later breakpoint already covers the tools prefix.

Evidence

  • Gates (macOS): cargo fmt --check and cargo clippy --all-targets -- -D warnings are clean. cargo test passes 705, fails 0, ignores 4. The PG/Redis env-gated suites were not run because no storage code changed.

  • Fail-first: native_system_blocks_stay_inside_the_openai_wire_shapes and tool_result_cache_control_survives_the_openai_wire both fail on the unfixed source.

  • Live test: a release build sent requests through a recording proxy to OpenRouter, model anthropic/claude-sonnet-4.6. Each shape went out twice; the table shows the second call's cache read.

    breakpoint on before after
    system blocks 0 3227
    user text block 3238 3232
    trailing tool_result (+ system) 0 3879

    In a 4-turn growing agent loop, cache reads are 0 / 3885 / 3968 / 4051 tokens after the fix. On every ledger row, cost_micros is within 1 micro of the vendor-reported cost.

  • Vendor acceptance (direct calls): a system content-part array (with and without cache_control) and a tool message content array with cache_control both return 200 on OpenAI, DeepSeek, Qwen (compatible mode), Moonshot, xAI, MiniMax, OpenRouter and Gemini's OpenAI-compatible endpoint.

…esult

A /v1/messages request served by an openai-chat model flattened its system
blocks into a string system message and each tool_result into a string tool
message, so their cache_control markers never reached the vendor. Through
OpenRouter to a Claude model, an agent loop (system breakpoint plus a
breakpoint on the trailing tool_result) never read its cache: turns ending
on a user text block paid the 1.25x write, and tool_result turns read
nothing.

The system message now carries the native blocks as content parts, and a
marked tool_result becomes a one-part text array that keeps its marker.
OpenRouter honors both shapes; OpenAI, DeepSeek, Qwen, Moonshot, xAI,
Gemini's compatible endpoint and MiniMax accept them. Tool-level markers
stay dropped: Gemini's compatible endpoint rejects the field, and a later
breakpoint covers the tools prefix.
@CMGS
CMGS merged commit 13a3d90 into main Oct 6, 2026
2 checks passed
@CMGS
CMGS deleted the fix/openai-wire-cache-breakpoints branch October 6, 2026 02:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant