Skip to content

feat(evals): record and replay live model transcripts #8424

Description

@sudoKrishna

Problem

Live model runs are nondeterministic and need an API key, so they never run in
CI. Scripted runs are hand-written from scratch and drift from what a real model
actually emits. Neither gives a durable, key-free regression test for real model
behavior: today's live finding evaporates on the next run.

Proposal

Record a live run once, replay it forever, through the real tool loop.

  • EVAL_RECORD=1 captures each model turn's streamed chunks per scenario to
    apps/sim/evals/agent-tool-use/fixtures/<scenario>.json.
  • A replay completion feeds the recorded chunks back through
    createOpenAICompatStreamingToolLoopStream — no provider key, no network.
  • Replay runs in the normal bun run test; live mode records the fixtures.
  • Fixtures are committed and reviewed like snapshots: a diff shows a behavior
    change, and the eval scores the real transcript instead of a hand-written one.

Acceptance criteria

  • EVAL_RECORD=1 DEEPSEEK_API_KEY=... bun run test:evals:live writes fixtures
  • Replay runs with no key and no network, and is collected by bun run test
  • A recorded scenario reproduces the live pass/fail it was captured from
  • Fixtures are committed and reviewed like snapshots
  • README documents recording, replaying, and refreshing a fixture

Follow-up

Once fixtures exist, the context/memory and subagent suites can replay real
multi-turn transcripts instead of scripting every turn.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    featureNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions