Skip to content

feat(evals): add first agent tool-use evaluation suite #8408

Description

@sudoKrishna

Problem

There is no systematic way to evaluate agent harness behavior. We test the tool loop and the executor with unit/integration tests, but those prove the cases we already thought of. We cannot measure whether a change improves or regresses real agent behavior across tool selection, planning, retrieval, and recovery.

Proposal

Introduce a minimal, deterministic evaluation layer for agent tool use:

  1. Define scenarios declaratively (tools + scripted model turns + tool results + expectations).
  2. Execute them through the real agent tool loop.
  3. Score named checks and basic metrics (pass/fail, tool calls, iterations, latency, tokens).
  4. Run locally with one command and in CI, writing a JSON + Markdown artifact.

First suite: agent tool-use reliability

Scope: the OpenAI-compatible streaming tool loop (apps/sim/providers/openai-compat/streaming-tool-loop.ts), which serves OpenAI, DeepSeek, Groq, Cerebras, and the other compatible providers. The model is scripted, so the suite is deterministic and needs no provider key, while the code under test is the real loop.

Coverage:

Category Cases
tool-selection single relevant tool; correct tool among distractors
planning multi-step dependent chain; parallel independent calls
retrieval retrieved value survives into the final answer
recovery tool error then retry; unknown tool; malformed argument JSON

Suggested starting point

  • Location: apps/sim/evals/agent-tool-use/
  • Reuse the existing Vitest coverage of streaming-tool-loop.ts and @sim/testing mocks.
  • A first implementation exists on a fork and can seed the PR (link once pushed).

Acceptance criteria

  • Define and run at least 5 agent scenarios (first cut has 8)
  • Report pass/fail plus basic metrics (JSON + Markdown)
  • Document how to add a new case
  • Run locally with one command (bun run test:evals from apps/sim)
  • Collected by the normal bun run test so a regression fails CI

Follow-up (out of scope for the first eval)

Run the same scripted model through the full DAGExecutor so agent-block wiring, variable resolution, and the executor retry/fallback policy are measured alongside the loop. The result/scoring shape is entry-point independent so both can share the report.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    featureNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions