feat(evals): compare agent tool-use across models - #8472
sudoKrishna wants to merge 7 commits into
Conversation
Add a deterministic eval layer for the agent harness. Scenarios script the OpenAI-compatible streaming tool loop with model turns and stub tool results, then score tool selection, planning, retrieval, and recovery without a provider key. - apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report - `bun run test:evals` from apps/sim runs the suite and writes the report - picked up by the normal vitest run so a regression fails CI - README documents the contract and how to add a case
Replay the same scenarios against a real model. The model is the only thing that changes: runScenario now takes an optional completion transport and a live mode that relaxes exact assertions (ordered subsequence, minimum successes) and skips scripted-only recovery cases. - live.ts: OpenAI-compatible transport + DeepSeek factory - agent-tool-use.live.test.ts: K trials per scenario, gated on EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI - live report with pass rates, avg iterations, latency, failed checks - test:evals:live script and README knobs
…ve mode The first live DeepSeek run exposed brittle assertions, not harness bugs: the model chained the tools correctly but the checks were case-sensitive and required an internal order id. Match the retrieved value case-insensitively and let live runs accept the grounded status rather than the internal id.
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor, with only executeProviderRequest mocked at the provider boundary. This covers agent-block input wiring, variable resolution from Start outputs, and executor run/error handling, which the direct loop harness cannot see. - executor-harness.ts: workflow builder + runExecutorScenario - shares the scorer (scoreExpectations) and report with the loop suite - two scenarios: Start->Agent output, and <start.message> resolution - README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the Agent block has retry enabled, and the executor replays it. The run must complete with the second response. Verifies providerCalls === 2, and fails without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the Agent block has a fallback model, and the handler serves the answer from gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails without the fallback row (checked locally: got gpt-4o, run errored).
Run the same live scenarios across a list of models and write a scenario x model matrix. models.ts resolves provider:model specs (DeepSeek, OpenAI, Groq, OpenRouter) and reads each provider's key from <PROVIDER>_API_KEY. - agent-tool-use.compare.live.test.ts: EVAL_MODELS x scenarios x trials - report.ts: buildLiveComparisonReport + JSON/Markdown matrix - report.test.ts: key-free aggregation coverage - test:evals:compare script; README documents the spec format
|
@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel. A member of the Team first needs to authorize it. |
|
| toolsMockFns.mockExecuteTool.mockImplementation( | ||
| async (toolId: string, params: Record<string, unknown>): Promise<ToolResponse> => { | ||
| const startedAt = Date.now() | ||
| const response = resultQueues.get(toolId)?.shift() ?? { success: true, output: {} } |
There was a problem hiding this comment.
Live tool results can be fabricated
Live comparisons still draw tool results from the scripted transcript. If a model retries beyond the scripted calls or calls a tool with different arguments, it receives an unrelated queued result or a fabricated successful empty result. That can change its next answer and misstate its tool-use success in the comparison. Match live results to the actual invocation, and do not treat an unhandled call as successful.
| "test:watch": "vitest", | ||
| "test:coverage": "vitest run --coverage", | ||
| "test:evals": "EVAL_REPORT_PATH=test-results/evals/agent-tool-use.json vitest run evals/agent-tool-use", | ||
| "test:evals:live": "EVAL_LIVE=1 vitest run --mode live evals/agent-tool-use", |
There was a problem hiding this comment.
Single-model command runs comparison
test:evals:live selects the whole eval directory, so it also runs the new comparison test. With a DeepSeek key, the documented single-model command makes the comparison’s default DeepSeek calls too, duplicating paid API work and producing an unexpected second report. Target only the single-model test file here.
| }) | ||
| } | ||
| }, | ||
| TIMEOUT_MS |
There was a problem hiding this comment.
This timeout covers the entire test case, while its EVAL_TRIALS runs execute sequentially. The default three trials share one 180-second deadline even though each trial can make multiple model requests. A slow, otherwise valid comparison can therefore time out before its trials finish. The single-model live test uses the same pattern; give each case a budget that accounts for its trials and requests.
|
|
||
| function createScriptedCompletion(scenario: AgentToolUseScenario): OpenAICompatCreateCompletion { | ||
| let turnIndex = 0 | ||
| return async () => { |
There was a problem hiding this comment.
Scripted eval misses feedback regressions
The scripted completion advances through fixed turns without reading the tool-result messages sent back to the model. If the loop stops including a result or error in the next request, the scripted retry and answer can still occur and the eval can stay green. Assert what the completion receives on later turns so this deterministic suite covers the feedback behavior it claims to test.
| for (const [label, modelRuns] of byModel) { | ||
| const results = modelRuns.map((run) => run.result) | ||
| const passed = results.filter((result) => result.passed).length | ||
| const [provider, model] = label.split('/') |
There was a problem hiding this comment.
Splitting the label on every / loses part of a model ID that contains a slash, as supported OpenRouter model IDs can. The JSON summary then names the wrong model even though its label and matrix column retain the full ID, making the report harder to use reliably. Preserve the provider and model from the run instead of parsing them from the label.
| @@ -0,0 +1,249 @@ | |||
| import { mkdirSync, writeFileSync } from 'node:fs' | |||
| import { dirname } from 'node:path' | |||
| import type { AgentToolUseResult, LiveScenarioSummary } from './types' | |||
There was a problem hiding this comment.
Relative imports violate app convention
This new file imports ./types, but the Sim import directive requires absolute imports. The same pattern appears in models.ts and harness.ts. Change these to @/evals/agent-tool-use/... imports; this repository requirement must be satisfied before merging.
Context Used: Import patterns for the Sim application (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Summary
Run the same live scenarios across several models and produce a scenario × model
matrix, so tool-use reliability, latency, and token cost can be compared. This
is the measurement needed before changing model routing or
sim-auto.Stacked on #8409 (the eval harness) — base branch is
feat/agent-tool-use-evals.Closes #8471
What changed
models.ts— resolvesprovider:modelspecs against a small registry(DeepSeek, OpenAI, Groq, OpenRouter) and reads each provider's key from
<PROVIDER>_API_KEYagent-tool-use.compare.live.test.ts—EVAL_MODELS× scenarios × trials,opt-in and never in CI
report.ts—buildLiveComparisonReport+ JSON/Markdown matrixreport.test.ts— key-free aggregation and file-writing coveragetest:evals:comparescript; README documents the spec formatHow to run
cd apps/sim EVAL_MODELS=deepseek:deepseek-chat,deepseek:deepseek-reasoner \ DEEPSEEK_API_KEY=... bun run test:evals:compareThe report is
test-results/evals/agent-tool-use-compare.{json,md}: per-modelpass rate, iterations, latency, tokens, plus a per-scenario pass-rate matrix.