Problem
There is no systematic way to evaluate agent harness behavior. We test the tool loop and the executor with unit/integration tests, but those prove the cases we already thought of. We cannot measure whether a change improves or regresses real agent behavior across tool selection, planning, retrieval, and recovery.
Proposal
Introduce a minimal, deterministic evaluation layer for agent tool use:
- Define scenarios declaratively (tools + scripted model turns + tool results + expectations).
- Execute them through the real agent tool loop.
- Score named checks and basic metrics (pass/fail, tool calls, iterations, latency, tokens).
- Run locally with one command and in CI, writing a JSON + Markdown artifact.
First suite: agent tool-use reliability
Scope: the OpenAI-compatible streaming tool loop (apps/sim/providers/openai-compat/streaming-tool-loop.ts), which serves OpenAI, DeepSeek, Groq, Cerebras, and the other compatible providers. The model is scripted, so the suite is deterministic and needs no provider key, while the code under test is the real loop.
Coverage:
| Category |
Cases |
tool-selection |
single relevant tool; correct tool among distractors |
planning |
multi-step dependent chain; parallel independent calls |
retrieval |
retrieved value survives into the final answer |
recovery |
tool error then retry; unknown tool; malformed argument JSON |
Suggested starting point
- Location:
apps/sim/evals/agent-tool-use/
- Reuse the existing Vitest coverage of
streaming-tool-loop.ts and @sim/testing mocks.
- A first implementation exists on a fork and can seed the PR (link once pushed).
Acceptance criteria
Follow-up (out of scope for the first eval)
Run the same scripted model through the full DAGExecutor so agent-block wiring, variable resolution, and the executor retry/fallback policy are measured alongside the loop. The result/scoring shape is entry-point independent so both can share the report.
Problem
There is no systematic way to evaluate agent harness behavior. We test the tool loop and the executor with unit/integration tests, but those prove the cases we already thought of. We cannot measure whether a change improves or regresses real agent behavior across tool selection, planning, retrieval, and recovery.
Proposal
Introduce a minimal, deterministic evaluation layer for agent tool use:
First suite: agent tool-use reliability
Scope: the OpenAI-compatible streaming tool loop (
apps/sim/providers/openai-compat/streaming-tool-loop.ts), which serves OpenAI, DeepSeek, Groq, Cerebras, and the other compatible providers. The model is scripted, so the suite is deterministic and needs no provider key, while the code under test is the real loop.Coverage:
tool-selectionplanningretrievalrecoverySuggested starting point
apps/sim/evals/agent-tool-use/streaming-tool-loop.tsand@sim/testingmocks.Acceptance criteria
bun run test:evalsfromapps/sim)bun run testso a regression fails CIFollow-up (out of scope for the first eval)
Run the same scripted model through the full
DAGExecutorso agent-block wiring, variable resolution, and the executor retry/fallback policy are measured alongside the loop. The result/scoring shape is entry-point independent so both can share the report.