Problem
Live model runs are nondeterministic and need an API key, so they never run in
CI. Scripted runs are hand-written from scratch and drift from what a real model
actually emits. Neither gives a durable, key-free regression test for real model
behavior: today's live finding evaporates on the next run.
Proposal
Record a live run once, replay it forever, through the real tool loop.
EVAL_RECORD=1 captures each model turn's streamed chunks per scenario to
apps/sim/evals/agent-tool-use/fixtures/<scenario>.json.
- A replay completion feeds the recorded chunks back through
createOpenAICompatStreamingToolLoopStream — no provider key, no network.
- Replay runs in the normal
bun run test; live mode records the fixtures.
- Fixtures are committed and reviewed like snapshots: a diff shows a behavior
change, and the eval scores the real transcript instead of a hand-written one.
Acceptance criteria
Follow-up
Once fixtures exist, the context/memory and subagent suites can replay real
multi-turn transcripts instead of scripting every turn.
Problem
Live model runs are nondeterministic and need an API key, so they never run in
CI. Scripted runs are hand-written from scratch and drift from what a real model
actually emits. Neither gives a durable, key-free regression test for real model
behavior: today's live finding evaporates on the next run.
Proposal
Record a live run once, replay it forever, through the real tool loop.
EVAL_RECORD=1captures each model turn's streamed chunks per scenario toapps/sim/evals/agent-tool-use/fixtures/<scenario>.json.createOpenAICompatStreamingToolLoopStream— no provider key, no network.bun run test; live mode records the fixtures.change, and the eval scores the real transcript instead of a hand-written one.
Acceptance criteria
EVAL_RECORD=1 DEEPSEEK_API_KEY=... bun run test:evals:livewrites fixturesbun run testFollow-up
Once fixtures exist, the context/memory and subagent suites can replay real
multi-turn transcripts instead of scripting every turn.