Skip to content

feat(evals): compare agent tool-use across models - #8472

Open
sudoKrishna wants to merge 7 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-model-compare
Open

sudoKrishna wants to merge 7 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-model-compare

Conversation

@sudoKrishna

Copy link
Copy Markdown

Summary

Run the same live scenarios across several models and produce a scenario × model
matrix, so tool-use reliability, latency, and token cost can be compared. This
is the measurement needed before changing model routing or sim-auto.

Stacked on #8409 (the eval harness) — base branch is feat/agent-tool-use-evals.

Closes #8471

What changed

  • models.ts — resolves provider:model specs against a small registry
    (DeepSeek, OpenAI, Groq, OpenRouter) and reads each provider's key from
    <PROVIDER>_API_KEY
  • agent-tool-use.compare.live.test.ts — EVAL_MODELS × scenarios × trials,
    opt-in and never in CI
  • report.ts — buildLiveComparisonReport + JSON/Markdown matrix
  • report.test.ts — key-free aggregation and file-writing coverage
  • test:evals:compare script; README documents the spec format

How to run

cd apps/sim
EVAL_MODELS=deepseek:deepseek-chat,deepseek:deepseek-reasoner \
  DEEPSEEK_API_KEY=... bun run test:evals:compare

The report is test-results/evals/agent-tool-use-compare.{json,md}: per-model
pass rate, iterations, latency, tokens, plus a per-scenario pass-rate matrix.

Add a deterministic eval layer for the agent harness. Scenarios script the
OpenAI-compatible streaming tool loop with model turns and stub tool results,
then score tool selection, planning, retrieval, and recovery without a
provider key.

- apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report
- `bun run test:evals` from apps/sim runs the suite and writes the report
- picked up by the normal vitest run so a regression fails CI
- README documents the contract and how to add a case
Replay the same scenarios against a real model. The model is the only thing
that changes: runScenario now takes an optional completion transport and a
live mode that relaxes exact assertions (ordered subsequence, minimum
successes) and skips scripted-only recovery cases.

- live.ts: OpenAI-compatible transport + DeepSeek factory
- agent-tool-use.live.test.ts: K trials per scenario, gated on
  EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI
- live report with pass rates, avg iterations, latency, failed checks
- test:evals:live script and README knobs
…ve mode

The first live DeepSeek run exposed brittle assertions, not harness bugs:
the model chained the tools correctly but the checks were case-sensitive and
required an internal order id. Match the retrieved value case-insensitively
and let live runs accept the grounded status rather than the internal id.
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor,
with only executeProviderRequest mocked at the provider boundary. This covers
agent-block input wiring, variable resolution from Start outputs, and executor
run/error handling, which the direct loop harness cannot see.

- executor-harness.ts: workflow builder + runExecutorScenario
- shares the scorer (scoreExpectations) and report with the loop suite
- two scenarios: Start->Agent output, and <start.message> resolution
- README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the
Agent block has retry enabled, and the executor replays it. The run must
complete with the second response. Verifies providerCalls === 2, and fails
without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the
Agent block has a fallback model, and the handler serves the answer from
gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails
without the fallback row (checked locally: got gpt-4o, run errored).
Run the same live scenarios across a list of models and write a scenario x
model matrix. models.ts resolves provider:model specs (DeepSeek, OpenAI, Groq,
OpenRouter) and reads each provider's key from <PROVIDER>_API_KEY.

- agent-tool-use.compare.live.test.ts: EVAL_MODELS x scenarios x trials
- report.ts: buildLiveComparisonReport + JSON/Markdown matrix
- report.test.ts: key-free aggregation coverage
- test:evals:compare script; README documents the spec format
@sudoKrishna
sudoKrishna requested a review from a team as a code owner September 30, 2026 17:18
@vercel

vercel Bot commented Sep 30, 2026

Copy link
Copy Markdown

@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 2/5

[Low risk] Adds evaluation test suite for agent tool handling.

The PR is not ready to merge because live comparisons can measure fabricated tool outcomes and the live commands can make unintended or prematurely timed-out API runs.

Findings

  1. P1 Live tool results can be fabricated ▶
  2. P1 Single-model command runs comparison ▶
  3. P1 Trials share one timeout ▶
  4. P2 Scripted eval misses feedback regressions ▶
  5. P2 Model IDs are truncated ▶
  6. P2 Relative imports violate app convention ▶

Summary

The PR adds deterministic and opt-in live agent tool-use evals, a provider registry, and a scenario-by-model comparison report.

  • The live harness can give models scripted or fabricated tool results, undermining comparison accuracy.
  • The single-model command also launches the comparison, and the live test timeout does not account for sequential trials.
  • The scripted suite needs a feedback assertion; report model IDs and import conventions need correction.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  S[Scenarios] --> T[Scripted eval]
  S --> L[Live model trials]
  T --> H[Streaming tool loop]
  L --> H
  H --> Q[Script-derived tool-result queues]
  H --> R[Scored runs]
  R --> M[Scenario × model report]
Loading

Reviews (1) · Last reviewed commit: "feat(evals): compare agent tool-use acro..."

toolsMockFns.mockExecuteTool.mockImplementation(
async (toolId: string, params: Record<string, unknown>): Promise<ToolResponse> => {
const startedAt = Date.now()
const response = resultQueues.get(toolId)?.shift() ?? { success: true, output: {} }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Live tool results can be fabricated

Live comparisons still draw tool results from the scripted transcript. If a model retries beyond the scripted calls or calls a tool with different arguments, it receives an unrelated queued result or a fabricated successful empty result. That can change its next answer and misstate its tool-use success in the comparison. Match live results to the actual invocation, and do not treat an unhandled call as successful.

Comment thread apps/sim/package.json
"test:watch": "vitest",
"test:coverage": "vitest run --coverage",
"test:evals": "EVAL_REPORT_PATH=test-results/evals/agent-tool-use.json vitest run evals/agent-tool-use",
"test:evals:live": "EVAL_LIVE=1 vitest run --mode live evals/agent-tool-use",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Single-model command runs comparison

test:evals:live selects the whole eval directory, so it also runs the new comparison test. With a DeepSeek key, the documented single-model command makes the comparison’s default DeepSeek calls too, duplicating paid API work and producing an unexpected second report. Target only the single-model test file here.

})
}
},
TIMEOUT_MS

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Trials share one timeout

This timeout covers the entire test case, while its EVAL_TRIALS runs execute sequentially. The default three trials share one 180-second deadline even though each trial can make multiple model requests. A slow, otherwise valid comparison can therefore time out before its trials finish. The single-model live test uses the same pattern; give each case a budget that accounts for its trials and requests.


function createScriptedCompletion(scenario: AgentToolUseScenario): OpenAICompatCreateCompletion {
let turnIndex = 0
return async () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Scripted eval misses feedback regressions

The scripted completion advances through fixed turns without reading the tool-result messages sent back to the model. If the loop stops including a result or error in the next request, the scripted retry and answer can still occur and the eval can stay green. Assert what the completion receives on later turns so this deterministic suite covers the feedback behavior it claims to test.

for (const [label, modelRuns] of byModel) {
const results = modelRuns.map((run) => run.result)
const passed = results.filter((result) => result.passed).length
const [provider, model] = label.split('/')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Model IDs are truncated

Splitting the label on every / loses part of a model ID that contains a slash, as supported OpenRouter model IDs can. The JSON summary then names the wrong model even though its label and matrix column retain the full ID, making the report harder to use reliably. Preserve the provider and model from the run instead of parsing them from the label.

@@ -0,0 +1,249 @@
import { mkdirSync, writeFileSync } from 'node:fs'
import { dirname } from 'node:path'
import type { AgentToolUseResult, LiveScenarioSummary } from './types'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Relative imports violate app convention

This new file imports ./types, but the Sim import directive requires absolute imports. The same pattern appears in models.ts and harness.ts. Change these to @/evals/agent-tool-use/... imports; this repository requirement must be satisfied before merging.

Context Used: Import patterns for the Sim application (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(evals): compare agent tool-use across models

1 participant