feat(check): parse vapi-checks.yml check configuration - #60
Merged
Merged
Conversation
scott-lowe-vapi
force-pushed
the
refactor/config-free-engine-modules
branch
from
October 1, 2026 23:11
3e0e277 to
0518af1
Compare
scott-lowe-vapi
force-pushed
the
feat/check-config
branch
from
October 1, 2026 23:11
2fb28f4 to
58e2d36
Compare
This was referenced Oct 1, 2026
Contributor
Author
This was referenced Oct 1, 2026
scott-lowe-vapi
marked this pull request as ready for review
October 1, 2026 23:52
vtkovapi
approved these changes
Oct 3, 2026
Contributor
Author
Merge activity
|
scott-lowe-vapi
changed the base branch from
refactor/config-free-engine-modules
to
graphite-base/60
October 3, 2026 06:02
Add the config contract for inline simulation PR checks: a root vapi-checks.yml declaring, per check, the org whose files to read, an optional run org and base URL, the targets (assistants/<id> or squads/<id>), the suites and simulations to run, credential and phone bindings (promotion.yml's shape), extra trigger paths, and run settings (transport, iterations, timeout, tool-mock policy, webhook stripping) with repo-wide defaults. It is a separate file from promotion.yml because single-org customers can't have a valid promotion pipeline. Unknown keys are rejected, including `mode`: checks always build the target from the branch. Nothing reads the config yet; the dry-run CLI that uses it follows. Refs TEST-141 Co-Authored-By: Claude Opus 5.5 <[email protected]>
scott-lowe-vapi
force-pushed
the
feat/check-config
branch
from
October 3, 2026 06:04
58e2d36 to
c9a6467
Compare
scott-lowe-vapi
added a commit
that referenced
this pull request
Oct 3, 2026
…LI (#61) ## Value **V.A.L.U.E. tier:** project — PR 5 of 10 for inline simulation PR checks ([TEST-141](https://linear.app/vapi/issue/TEST-141/gitops-run-simulation-suites-against-pr-changes-inline-as-ci-checks)); this PR adds an offline dry run, and nothing is sent. - **Problem:** to test a PR's changes without deploying them, the check has to turn the branch's files into one inline `POST /eval/simulation/run` body: the target with every tool, handoff and structured output, plus every scenario, judge and personality. It has to assemble that body the way push and the runtime would, or the check tests something other than what ships. - **Who it affects:** gitops users, who can now see exactly what a check would send (`npm run check -- core --dry-run --print-payload`) before any minutes are spent. Every later PR (mock policy, live runs, workflow, promotion gate) builds on this payload. - **What changes:** - **`src/check-payload.ts`** (with `check-payload-assistant.ts` and `check-payload-refs.ts`) builds the body from `orgResourcesRead` and the org states. It's pure, makes no network calls, and collects every problem so one dry run reports them all. - **Tools:** - runtime tool order: `model.tools`, then `toolIds`, then `toolRefs`. This is what `callAssistantsGet` does, and the parity run's likely `endCall` difference came from getting it wrong; - `##` comments stripped, and UUID references resolved through state; - a `toolRefs` pin wins over a duplicate `toolIds` entry, with a warning that the version pin is ignored; - `knowledgeBase` tools are kept by run-org UUID, since the API refuses them inline. - **Squads:** members are inlined (from a file, by UUID, or inline), and handoffs to members switch from ID to member name. - **Assistants:** - hook `do[].toolId` becomes an inline tool; - `artifactPlan.structuredOutputIds`, plus structured outputs that link the assistant through their own `assistant_ids`, are inlined; - linkage and server fields are stripped. - **Simulations:** - suites and simulations are deduplicated into entries with unique names of at most 80 characters; - judges' `structuredOutputId`s are inlined; - local personalities are inlined, and stock ones pass by ID. - **Credentials** are bound by name to the run org's UUIDs using promotion's binding policy (`bind` / `omit`). - **Fails the build, naming the field:** - a missing file; - a handoff leaving the target, or an assistant target that hands off; - legacy `assistantDestinations` by ID; - members without a name, or with the same name; - two inlined tools with the same type and name, which stored would keep but inline would silently drop; - tools referenced by ID inside overrides (under strict mocks); - over chat: audio judges, scenario hooks, or no required text judge; - any reference still a name after the build (the backstop for #31's silent drops); - a body over 4.5 MB. - **Warns** when text mentions `handoff_to_…` but a handoff is auto-named, because generated names differ inline and stored. - **`src/check-cmd.ts`:** `npm run check -- <check>|--all --dry-run [--print-payload [dir]]`. - Offline: no API key needed, and it honours `VAPI_GITOPS_ROOT`. - Exits 0 when every payload builds, and 2 on a usage, config or build error. - Without `--dry-run` it exits 2; live runs land in PR 7. - `package.json` gets a `check` script, and the README and AGENTS.md command tables get `npm run check` rows. - New fixture `tests/fixtures/check-parity/`: the TEST-141 parity squad written as gitops files. - **Deliberately not here:** the fail-closed mock policy (dead servers, default error mocks, the tool-type allowlist), which is the next PR. Until then the payload carries tools as written, and there's no live path. ## Evidence of value **The builder reproduces the payload that scored 15/15 in the parity run.** - **How:** the dental squad from the 2026-10-01 parity experiment was written as gitops files (`.md` assistants with `toolIds`, a handoff tool by `assistantId`, judges by `structuredOutputId`). `npm run check -- core --dry-run --print-payload` was run on it, and the output was diffed against the inline body the experiment sent, rebuilt from `parity.mjs`. - **Result:** the only differences are these. | Path | Experiment sent | Built from files | Why | |---|---|---|---| | `members[*].assistant.model.tools` order | `[lookup_patient, handoff, endCall]`, `[check_availability, endCall]` | `[endCall, lookup_patient, handoff]`, `[endCall, check_availability]` | **Intended:** the runtime order of the stored arm (`model.tools` then `toolIds`). Same tools, byte-for-byte, order aside | | `squad.name`, `personality.name` | `inline`, `caller` | `Bright Smile Dental`, `Dental caller` | Fixture names | | `iterations`, `transport` | 5, (API default) | 1, `vapi.webchat` | `vapi-checks.yml` defaults | - Every prompt, judge, tool mock, tool definition and handoff destination is identical. The handoff by `assistantId: scheduler` came out as `assistantName: "Scheduler"`, exactly what the experiment sent. **Tests:** `npm test` goes from 404 to 430 passing (26 new), and `npm run build` is clean. ## Testing plan - **`tests/check-payload.test.ts`** (21 tests, temp-dir fixtures): - the parity fixture end to end: tool order, member-name handoffs, `.md` prompt, inline judges, entries, transport; - `.md` body as the only system message; - missing `toolIds`, and `toolIds` by UUID with server fields stripped; - the `toolRefs` pin and warning; - `knowledgeBase` kept by UUID, and failing with no UUID; - duplicate tool names; - hook tools; - structured outputs from `artifactPlan` and from `assistant_ids` (by slug and by UUID); - an assistant target that hands off; - squad members by file, UUID and inline; - every squad problem reported at once; - strict vs off overrides; - credentials bound, omitted and missing; - a leftover reference, with free-form `parameters` ignored; - the chat rules, and voice allowing them; - stock and local personalities, name truncation and dedupe; - missing suites, simulations, scenarios, personalities and targets; - the auto-handoff warning; - the 4.5 MB limit. - **`tests/check-cmd.test.ts`** (5 tests): - a dry run with `--print-payload`; - `--all` with one broken check exiting 2; - no config exiting 0; - usage, selection, config and live-mode errors exiting 2; - a missing state file as a warning. - **Not tested:** - **Any live run:** nothing is sent until PR 7, so API acceptance of a built body is shown only for the parity fixture (the experiment's 201). - **Real customer squads:** override merge semantics and integration tool shapes are untested. - Phone-number bindings beyond the unit-tested promotion helper. - EU base URLs. Stacked on #60. Refs TEST-141 🤖 Generated with [Claude Code](https://claude.com/claude-code)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Value
V.A.L.U.E. tier: project — PR 4 of 10 for inline simulation PR checks (TEST-141); this PR adds a config parser with no caller yet.
promotion.ymlcan't hold it:promotionConfigParserequires a pipeline of two or more orgs, and most customers have one.src/check-config.tsparses a rootvapi-checks.ymlinto typed check definitions:org, plus an optionalrunOrgandbaseUrl;targets(assistants/<id>orsquads/<id>, nested IDs allowed);suitesand/orsimulations;bindings(promotion's shape, viapromotionBindingsParse);paths;defaults: transportchat, 1 iteration, 20 min,toolMocks: strict,stripWebhooks: true.vapi-checks.example.ymlis the commented starting point: a single-org check, plus a commented CI-org check.mode:key fails with a message explaining that checks always build from the branch, since a "run what's deployed" mode was deliberately left out.Evidence of value
runOrg=org, credentialsbind/ phonesomit); default and per-check overrides; a CI-org check with EUbaseUrland bindingstests/check-config.test.ts)npm testnpm run buildcleanTesting plan
tests/check-config.test.tscovers:mode, slugs, target shape, extension and..in IDs, duplicates, missing suites/simulations, each setting's range,baseUrl,paths, bindings;checksConfigLoad, which returnsnullwhen novapi-checks.ymlexists (checks are opt-in).baseUrlfallback order (.env.<runOrg>,$VAPI_BASE_URL,api.vapi.ai) is applied at run time in PR 7.Stacked on #59.
Refs TEST-141
🤖 Generated with Claude Code