Skip to content

Fix Copilot macro runner tool identity and result guards - #3099

Open
George Ng (GeorgeNgMsft) wants to merge 1 commit into
georgengmsft-macro-replay-classificationfrom
georgengmsft-macro-runner-tool-access
Open

George Ng (GeorgeNgMsft) wants to merge 1 commit into
georgengmsft-macro-replay-classificationfrom
georgengmsft-macro-runner-tool-access

Conversation

@GeorgeNgMsft

Copy link
Copy Markdown
Contributor

This follow-up to #3097 fixes two coupled problems in agent-guided macro execution: the runner could mistake MCP backend provenance for a callable namespace, and induced result guards described Copilot's event/UI wrapper rather than the response the runner actually receives. The change preserves the original tool identity, permissions, raw evidence, and immutable approved versions.

  • Resolve the exact captured callable before deferred discovery. Document the verified web_search / github-mcp-server/web_search bridge without allowing arbitrary aliases or provider substitution.
  • Capture separately redacted model-facing results from the SDK's result.content, preserving valid JSON primitives and ordinary text as well as the raw event result.
  • Use the model-facing representation for normal guards and prior-step result bindings throughout agent-required procedures, including mixed procedures. Deterministic-only induction continues using raw results.
  • Keep legacy guards intact, warn when model-facing evidence is absent, and require recapture and explicit approval when an older macro needs unobservable fields.
  • Add regression coverage for callable/provenance preservation, redaction, absent versus null/false/zero/empty results, mixed-step references, deterministic compatibility, and runner safeguards.

Local validation before final adversarial review

Actual standalone Copilot CLI worker 1.0.89 executed a benign IANA search, captured it, and ran its normally induced/validated/approved version 2 through the real typeagent:typeagent-macro-runner. The runner inspected the immutable version and executed exactly one live web_search with the original arguments. Execution events report success: true; an independent harness checked all 11 normal inferred type/path guards, including answer and citation paths, against the actual returned model-facing content. No guard pruning, synthetic result, retry, provider substitution, or candidate submission was used in this final test.

Harness boundary: real headless CLI events were streamed through production SessionCapture; unrelated conversation-history sinks were isolated, and the headless CLI final-result marker was adapted to SDK session.idle. This was not an unmodified interactive extension-recording test. The fixture and its persisted macro/trace/handoff were removed after retaining evidence and verifying the approved version was unchanged.

Evidence identifiers: run 36b64f4a-bc0d-4002-8b1e-950aeec44773, runner b1f17ff8-34a8-4554-8820-052e2d04b457, live search call_4pXqFIfIxtxnas16NpdfhPf9.

  • Dependency-aware plugin, macro, and server builds passed.
  • 46 macro tests, 178 plugin unit tests, 14 replay-host tests, and the recording RPC test passed.
  • Formatting and all four PR gates passed against the exact parent base, including tests in lint/complexity checks.
  • Two independent final adversarial reviewers examined the expanded change after full guarded CLI validation; neither reported a substantive issue. Reviewer model identities were not reliably exposed, so model diversity is unverified.
  • Reviewed base: b46ad19a93dd240a6ea169ffa315aa5759e9bf39; reviewed content committed unchanged as 3aab21a035767af6d3573243ce719a2f6093c137.

Compatibility and recovery

The original access refusal was observed in an actual user run, but baseline resolution was nondeterministic: another baseline fixture found the tool and then failed the separate wrapper guards. This PR does not claim a proven permission or subagent-inheritance defect.

The user's existing Consider Stay Home Day version 2 was neither changed nor replayed. Its old wrapper-field guards may still be unverifiable. After updating the server/plugin and starting a fresh Copilot session, record a new interaction, inspect the new draft's inputs and guards, and explicitly approve it. Do not waive old guards or rewrite the approved version. Parameterization/reasoning features are out of scope.

Local deployment and rollback

Only five artifacts were updated: the installed server bundle, two runner profiles, and two extension bundles. Each extension received only the verified capture-module change; every other installed module was preserved byte-for-byte. The rebuilt server bundle differed only in the intended induction logic. Backups and hashes were recorded; the installed daemon was restarted and verified responsive. No routing settings, conversation bindings, Azure identities, or authentication helpers were changed.

Local evidence/rollback directory:
C:\Users\georgeng\.copilot\session-state\af340a10-e91b-4f95-8997-d99fb88a000d\files

Key files: guarded-verification.json, real-capture-definition.json, cli-tool-metadata.json, fixture-cleanup.json, result-deployment.json, extension-module-provenance.json, and runtime-before-results. Original runner-only backups are runner-before-0.agent.md and runner-before-1.agent.md. Restore only receipt-listed files after checking for subsequent updates; stop/start the daemon when restoring its bundle.

This PR targets the inherited parent branch, leaves #3097 unchanged, and is for human review; it has not been merged or self-approved.

Preserve recorded callable identity and backend provenance. Capture redacted model-facing results separately and use them for whole-agent macro guards and result bindings, with explicit legacy recovery guidance.

Co-authored-by: Copilot App <[email protected]>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant