Skip to content

v0.9.8: search cleanup recovery, new search providers, billing and chat reliability - #8494

Merged
waleedlatif1 merged 42 commits into
mainfrom
staging
Oct 1, 2026
Merged

waleedlatif1 merged 42 commits into
mainfrom
staging

Conversation

@icecrasher321

@icecrasher321 icecrasher321 commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

icecrasher321 and others added 30 commits September 30, 2026 09:51
* docs(library): update best-ai-automation-tools-2026

* Pi Babysit: address PR #8465 feedback

---------

Co-authored-by: Sim Pi Agent <[email protected]>
* docs(library): update best-gumloop-alternatives-in-2026

* Pi Babysit: address PR #8466 feedback

---------

Co-authored-by: Sim Pi Agent <[email protected]>
… 2026 (#8467)

* feat(library): Best AI Agent Builders for Slack and CRM Automation in 2026

* Pi Babysit: address PR #8467 feedback

---------

Co-authored-by: Sim Pi Agent <[email protected]>
…ch Does Your Team Need? (#8468)

* feat(library): Marketing Automation Platform vs AI Agent Builder: Which Does Your Team Need?

* Pi Babysit: address PR #8468 feedback

---------

Co-authored-by: Sim Pi Agent <[email protected]>
…r CPU load (#8458)

The four-patch turn-budget test drove 2,400 deltas over a ~300 KB base, about
2.4 s of CPU on an idle machine and 11-15 s at load 150, past the 10 s test
timeout. Pacing already runs on fake timers; the flake was CPU time, not wall
clock. Four 30 s patches at a 500 ms tick still stream ~25 MB uncapped against
the 8 MiB budget, so the test still fails when the cap is removed.
…8462)

* fix(sim-cli): report an embedded SimApiError with its own exit code

The embedded renderer returned 1 for every SimApiError, so an operation
wait that timed out (exit 4) or needs configuration (exit 3) read as a
generic failure in-process while the installed CLI reported the real
status. Return error.exitCode, as the terminal renderer does.

* test(sim-cli): make the embedded wait-timeout test independent of runner speed

The test gave the wait 10 ms of real time, so a slow runner could reach
the deadline before the first status check and exit 4 without printing
the receipt. The clock now advances only when the stubbed status is
served, so the wait times out only after it has a receipt to print.
… file mount (#8456)

* fix(sandbox): withhold workbench certification after an unprovenanced file mount

A persistent chat workbench stays "clean" only while every input it received
was classified secret-free, and the scratch-file read hands its bytes to the
model on that basis. File mounts resolved from platform file objects were
counted only when a provenance source existed, so a mount whose key has no
canonical metadata record (or no principal to bind one) left the machine
certified. Both the `files` parameter and a mount marker in context variables
reach this resolver from model-supplied Function parameters.

The resolver now reports how many mounts had no provenance source. When a
workbench session receives any, the session request carries
`unprovenancedInputs` and the code boundary records the machine as unknown.
Workflow runs have no session and keep their existing absence policy.

* test(sandbox): certify workbench history against real Redis and close the mount bypass

- Replace the certification unit test, which restated the history script in a
  fake Redis, with an integration suite that runs the real code boundary against
  the real script in a disposable Redis.
- Have the route test's mocked mount resolver return a fixed count per test
  rather than restating the counting rule.
- Strip model-supplied `_sandboxFiles` from Copilot Function calls. Only resolved
  inputs may populate it, and a supplied URL mount would skip their provenance.
- Document that public storage contexts always count as unprovenanced mounts.

* test(sandbox): restore only the env the certification suite changed

The suite's cleanup assigned undefined to REDIS_URL when it had been
unset, which stores the string "undefined" for later suites in the
worker, and asked for a Redis client even when the suite was skipped.

* test(sandbox): close the shared Redis client after the certification suite
* fix(mothership): settle Chat runs no controller will finish

A run whose process died before finalize, whose controller was superseded
with no successor, or that was stopped while no controller existed stayed
unfinished forever: its chat marker kept pointing at it, so the chat read
as busy and a reconnect polled a run nothing would ever end.

- Stop now settles the run as cancelled once no controller of its stream
  holds the chat lock, including after it force-releases a controller that
  did not exit in time.
- The stale-execution cron settles leased runs whose stream holds no chat
  lock, has no replay buffer left, and has been idle past the orchestration
  budget (so a reconnect has nothing left to resume), and runs without a
  lease once idle for 24 hours. A run Stop already closed settles as
  cancelled, any other as error; its chat marker is released.
- Each settle is one conditional update on the run row that requires it to
  be unfinished, idle, and still naming the controller that was observed,
  so a finalizing controller or a successor's claim wins or loses against
  it atomically and the run settles exactly once.

* fix(mothership): lock chats before runs when settling orphans, and label them precisely

- Settling now locks the affected chat rows first, in id order, as a
  controller's claim does. Locking the run and then the chat deadlocked
  against a concurrent reconnect claim.
- Each sweep batch fails on its own, the chat markers of a batch clear in
  one statement, and a sweep settles at most 5k rows with a short pause
  between full batches.
- A run settles as cancelled only when its user pressed Stop. A newer turn
  also closes tool admission on older runs, and those now settle as errors.
- Runs without a controller lease keep their last write as their
  completion and retention time and read "never finalized (no controller
  lease)".
- The leased-run grace no longer derives from the orchestration deadline.
  Liveness comes only from the heartbeat-renewed chat lock; the grace and
  the replay TTL only bound how long a reconnect can resume a dead run.

* fix(mothership): never sweep a current headless run

A headless turn has no chat lease and no heartbeat, so its age says nothing
about whether it is still running once runs have no deadline. The sweep's
lease-less rule now applies only to runs admitted before the current
tool-execution protocol: every run the current code admits records the
current version, so after a deploy no such row can be live. A current
headless run is left to its own lifecycle, which always settles it.

The protocol version moves beside the other async-run constants so the
sweep can read it without importing the repository.

* refactor(mothership): require the recorded Stop inside the stopped-run settle

- Settling a stopped run now passes the Stop-row check as the update's own
  guard, so it cannot cancel a run nobody stopped; the separate stopped
  branch is gone and every settle derives cancelled from the Stop row.
- Chat lock ownership is read through getChatStreamLockOwners and trusted
  only when verified, instead of a second Redis read of the same keys.
- Settle transactions use the shared DbTransaction type.

* fix(mothership): fence orphan settlement on the chat lock and resume sweeps where they stopped

- The sweep takes each unowned leased run's chat lock under the run's own
  stream before settling it and releases it after the commit, so a
  reconnect can no longer lock the chat between the ownership check and
  the settle and then lose its claim; a reconnect that meets the fence
  retries.
- A sweep examines at most 10k candidates and settles at most about 5k,
  resuming from a cursor saved in Redis and wrapping to the first run, so
  runs that cannot be settled yet never starve the ones after them.
- Every settled run whose chat marker was released is announced, legacy
  runs included, so an open client stops showing the chat as busy.

* fix(knowledge): record a terminal status for every Slack Assistant run

The Slack Assistant admits its own run row and only ever marked it as an
error, so every completed or stopped Slack turn stayed active. It now
records the terminal status once, after the turn ends, through the shared
run-status update: complete on success, cancelled when its user stopped
it (in Slack or in Sim), and error otherwise.

Every other run-creating path already settles its run: interactive turns
through their controller's finalize, and headless turns that admit their
own run in the lifecycle's own finally.

* fix(knowledge): record the Slack run's status after its turn is saved

- The Slack Assistant now writes its run's one terminal status after its
  outcome and response are persisted, from the final outcome, so a failed
  save ends the run as an error instead of complete.
- A Stop lookup that fails no longer skips that write: the turn is treated
  as not stopped, logged, and settled as an error.
- The orphaned-run suite deletes the sweep cursor before each test and in
  teardown, so no later suite starts from its leftover position.
…under its statement timeout (#8460)

* fix(search): bound retirement pages by mutated rows so they finish under the statement timeout

A retirement page read 25,000 IDs and updated or deleted every target row among them in one
statement. Retiring a document is a non-HOT update that writes every index on `document`, and a
deleted chunk cascades into its projections, so on a KB that dominates the table a page's write
cost, not its scan, outran the two-minute statement timeout and failed the deploy migration.

Each page now mutates at most a row limit of its target rows. A page that reaches the limit
advances the cursor only to its last mutated row, and already-retired documents never spend the
limit. The limit starts at 2,000, halves after a slow page or a statement timeout (the timed-out
page rolls back with its cursor and is retried), and doubles after a fast full page. The completion
rechecks, which walk every captured KB once, run with a 30-minute timeout. The retirement stays
idempotent and resumes from its saved cursor.

* fix(search): shrink only timed-out page mutations and pace Search retirement pages

Only a statement timeout from a page's mutating statement now halves the row limit and retries
the rolled-back page. Any other timeout, such as a completion recheck, fails the run at once
instead of repeating the same statement at every smaller limit.

Pages are timed around the whole call, commit included, so the synchronous-replication wait
counts toward the slow-page threshold. Each page is followed by a pause as long as the page, up
to five seconds, and the row limit is capped at 8,000. Phase changes no longer adjust the limit.

The progress log now carries the phase, cursor and rows mutated, and the migration logs slow-page
halvings, phase changes and the start of the completion recheck.

* test(search): prove retired documents never spend the retirement row limit

A run of already-retired Search documents longer than the row limit must be crossed in one page. The test counts documents-phase statements and fails if the page filter on unretired rows is removed.

* fix(search): scale the retirement scan window with the row limit and time out resumed rechecks like completion

A capped page re-reads its scan from its last mutated row, so a fixed 25,000-ID window re-read
most of the same IDs on every page once the row limit shrank. Each page now reads at most four IDs
per row of its limit, capped at 25,000, and any fast page doubles the limit so sparse stretches
widen the window again.

A retry that finds retirement already complete revalidates every captured KB, as completion does,
so it now runs under the same 30-minute timeout instead of the two-minute page timeout.

* fix(search): never grow a Search retirement page back to a size that timed out, and pause after a timeout

A timed-out page halved the row limit, but one fast page doubled it straight back, so the run
alternated between the size that timed out and half of it, rolling back a full statement-timeout
page each time. The limit now grows only up to half of the smallest size that timed out, and a
timed-out page is followed by the same pause as any other page.
…c execute responses until the log is final (#8473)

* fix(logs): record a run's cost before it reads finished, and hold sync execute responses until the log is final

* fix(logs): stop the ledger-order reader when a completion throws

* fix(logs): keep a completion failure visible when the ledger-order reader also fails

* fix(logs): let the ledger-order reader observe every finished run

* fix(logs): assert ledger-before-terminal ordering deterministically at the ledger lock

* fix(logs): scope the ledger-lock waiter to the execution and surface early completion failures
)

* fix(mothership): keep Chat stream legs alive past the hour without a wall clock

- End the replay GET at its cap without a terminal event, so the client
  re-attaches from its cursor instead of ending a live turn with resume_timeout.
- Replace the per-leg worker SSE wall clock with an idle timeout
  (WORKER_STREAM_IDLE_TIMEOUT_MS, 120 s) that fails a silent leg as a retryable
  interruption; a caller-set timeout still applies.
- Let StreamRetryWindow run without a deadline by default and replenish its
  reachable budget only after five minutes of healthy streaming, never per event.
- Split the tool watchdog, permission wait, client tool wait, and delegation TTL
  onto their own constants and remove ORCHESTRATION_TIMEOUT_MS.
- Refresh the replay buffer TTLs from the chat-lock heartbeat, and report a
  replay gap for a cursor ahead of a buffer whose numbering restarted.
- Document why maxDuration stays on the chat POST and execute routes.

* fix(mothership): renew the Chat reconnect budget after a tail delivers events

A reconnect attempt whose re-attached tail delivered new events now restarts
the retry budget at the base delay, so separate network drops hours apart in
a long turn no longer add up to the ten-attempt exhaustion.

* fix(mothership): refresh the stream byte counter with its buffer, and never revive a closed buffer

- The chat-lock heartbeat now slides the owner byte counter's TTL along with the
  replay buffer's. Otherwise, after a park longer than an hour, the counter expired
  while the ring survived: the ring could grow to about twice its target, and its
  refunds drained the user counter.
- Scheduling a finished stream's cleanup now marks it closed, and the refresh leaves
  a closed stream alone. A heartbeat still in flight at teardown can no longer
  re-extend a buffer whose cleanup was already scheduled.
- Describe the worker idle timeout in terms of network intermediaries, and state
  that the delegation TTL reuses the long-running tool watchdog's cap.

* fix(mothership): bound a stalled worker error body and refill retries only on real progress

- A non-OK worker response's body is read under the same idle bound as the leg, so
  a stalled error body can no longer block the turn forever.
- The reachable retry budget refills only when a leg's delivered events span five
  minutes, from its first event to its latest. A leg that delivered one event and
  then only kept alive has made no progress and no longer refills it.

* fix(mothership): mark a closing buffer before expiring it, and report a capped replay as ended

- scheduleBufferCleanup sets the closed marker before shortening the TTLs, so a
  heartbeat refresh that lands mid-pipeline is either overridden or sees the marker.
- A reconnect that ends at its cap is reported as ended without a terminal, not as
  a client disconnect: the outcome is read before the route closes its own stream.
- The buffer TTL integration test keeps a 5 s TTL against a 12 s park, so a slow
  runner cannot expire the buffer between heartbeats.
… workflow block logs (#8459)

* fix(mothership): keep a tainted mount's verdict across the run_code crossing

A run_code / run_function that mounted a workspace file whose sidecar is
`unknown` latched its per-call registry through the crossing with only
`source-provenance-incomplete` and `inherited-incomplete-source`. Both are in
the registry's absence set, so the withheld result named no guard, and an
output file written from that registry (outputs.files[].path) was recorded as
`unrecorded` absence instead of taint.

The crossing now inherits the mounted registry's own reasons when it latched,
so the refusal names `mounted-file-provenance-unavailable` and writers keep the
taint. The mount refusal also reports the file id. No projection is relaxed:
the result is still withheld, exact mounts still redact, clean mounts are
unchanged.

* improvement(mothership): tell the model why a tool result was withheld

A withheld result reached the model as a bare `{ success: true }`, so the agent
could not tell a tainted input from an oversized payload and retried or
guessed. The withheld result now carries `withheldReason`, chosen from
code-defined wording by the guard that tripped: an input file, table, or
document with unknown secret provenance; provenance that could not be
verified; or content that could not be checked. It rides the existing
`resultWithheld` disclosure beside any effect ids, and the key is reserved so
an id cannot displace it.

No content, reason literal, or origin crosses; an absent registry still
carries nothing.

* improvement(mothership): log the size of a withheld tool result

A `content-refused` withholding said only that a complete registry refused the
payload. It can be refused by its encoded size or by the number of values the
projection walks. Row-shaped payloads reach the 100k-value traversal cap well
before the 16 MiB byte cap, and only while the call has an active secret. The
withheld log lines now report the result's encoded bytes and value count, so a
refusal names which cap it hit. Numbers only; no content is logged.

* fix(mothership): bound run_workflow block-log outputs before the secret projection

A run_workflow result echoes every block's output in `logs`. A run with many
row-shaped block outputs can exceed the projection's 100k-value traversal cap
while staying under its byte cap. Whenever the call had an active secret, the
whole result was then withheld, including the final output and error, and the
model saw a bare success.

Block-log outputs are now bounded to a quarter of each projection cap
(`MAX_CONTENT_NODES` values and the default byte cap). Past the budget, the
bulkiest outputs are replaced, largest first, with a
`logs get <executionId> --trace` pointer, in the same form the oversized-input
compaction already uses. The executionId, status, final output, error, and
`select` values are untouched, and the other three quarters of each cap stay
free for them.

* fix(mothership): bound browser-run workflow logs and keep the pointer length-blind

run_workflow can complete in the browser, and that restoration projected the
raw block logs, so a large browser-run workflow was still withheld. The
block-log compaction now lives in the shared workflow-output module and applies
to both the server handler and the client restoration when no `select` is
given. `select` still reads full values.

The omitted-output pointer no longer reports bytes. They were measured before
secret projection and so disclosed a secret's length. It reports the value
count and says "inspect with logs get" rather than promising the full value.

One `measureModelContent` in the projection module now serves both the
compaction and the withheld-size logging. It stops counting at the value cap
and tolerates values JSON cannot encode. The byte cap is exported beside
`MAX_CONTENT_NODES`.

* fix(mothership): resolve a browser run's select before the secret projection

The browser-run restoration projected the full raw block logs and only then
applied `select`, so a large run was still withheld whenever a selector was
given. The server handler selects first. Both paths now build their log fields
through one shared `presentWorkflowLogsForModel`, from raw logs and before
projection: a `select` resolves against the full logs and replaces them;
otherwise the echoed logs are bounded. Selected values are still projected, so
a selected secret is redacted.

* test(providers): expect the withheld reason on provider tool results

The provider tool path projects through the same withheld-result shape, so an
omitted model result now carries `resultWithheld` and its fixed-wording
`withheldReason`. The expectations still asserted the old empty output.

* fix(mothership): bound lifted outputs, length-blind input markers, and a walk-based measure

- run_block and run_workflow_until_block lift the stopping block's output into
  `output`. That copy is now bounded with the same budget and pointer as a
  block-log output, so one huge block no longer gets the whole response
  withheld.
- The truncated-input marker no longer carries the input's length. It is
  written before secret projection, so the length disclosed a secret's length.
- `measureModelContent` now walks the value by JSON's rules instead of
  serializing it. It stops at the first value, byte, or depth limit it passes
  and reports `exceeded`, measuring a string only while it fits the remaining
  byte budget. An output past a limit, including one nested past the
  projection's depth limit, becomes a pointer instead of voiding the run.
- The browser-run restoration now truncates echoed block inputs as the server
  handler does: both paths build their log fields through
  `presentWorkflowLogsForModel`.

* fix(mothership): keep no raw prefix in a truncated block input

A truncated block input kept its first 200 raw characters, written before
secret projection. A secret straddling that cut left a fragment that
whole-literal redaction cannot match, so part of the secret reached the model.
The marker now keeps nothing of the raw input:
`…[input omitted; inspect with logs get <executionId> --trace]`. Inputs over
the limit are echoed upstream data the caller already has or can fetch, so the
preview carried little.

* fix(mothership): bound run outputs only when the projection walks them against a secret

Without an active secret the model-facing projection passes JSON through under its byte cap
alone, so the block-output budget turned a large lifted run_block output into a pointer for no
reason. Output compaction now applies only when the call's registry makes the projection walk
the result, and a lifted output keeps the final output's share beside the bounded logs.

* fix(mothership): measure a run's whole model-facing result and bound its final output

A fixed share for the lifted output left no room for the log entries, envelope and error the
projection also counts, and a Response block's final output was never bounded, so a run near the
caps was still withheld whole. The assembled result is now measured as the projection will walk
it, and a final output that would push it past a cap is replaced with a pointer, on the server
and browser-run paths alike. An unencodable result is still refused as before.

* fix(mothership): leave an unencodable block output for the projection to refuse

The output bound handles size only. A block output JSON cannot encode cannot be checked, so it is
no longer sized past the budget and replaced with a pointer; the projection refuses it as it did
before this PR.

* fix(mothership): leave a result that fits the caps, and every result without a secret, as staging returns it

Block-log outputs were bounded whenever a secret was active, so a 30k-value or 4.5 MB result that
crossed in full before became pointers. Outputs are now bounded only when the whole result would
pass a projection cap. Without an active secret the server handler keeps its preview input marker
and the browser-run path leaves its logs untouched, in their original position, so a result with
no secret is byte-identical to before; only a walked call gets the marker that keeps nothing.

* test(mothership): pin the unencodable-output guard and the error in the whole-result measure

The unencodable-output test placed the BigInt in the lifted output, which the whole-result walk
reaches first, so it passed without the guard. It now puts a bulky log ahead of the BigInt log.
A new test fails if the error is left out of the measure the result is bounded against.
* feat(search): add Lucidchart and Lucidspark live search

* fix(search): validate Lucid queries and isolate stale candidates
… the replay ring by bytes (#8469)

* chore(mothership): sync the worker protocol for the read-only run replay

Adds StreamReplayRequest and StreamReplayEnd from the worker's contracts
(bun run contracts:sync).

* fix(mothership): re-sync a reconnect the replay ring cannot serve from the worker log

A reconnect whose cursor fell behind the ring, a fresh tab reading a ring that lost
its head, and a cursor ahead of a ring whose numbering restarted now stream the run
from the worker's read-only replay instead of ending the turn with replay_gap or
replaying a partial response.

- The reconnect route opens POST /api/streams/replay with no receipt and forwards it
  under its own cursors from 1, with x-mothership-stream-replay: log so the client
  rebuilds the turn from an empty response. A response on the replay is never handed
  back to the ring, which shares no position with the log.
- A parked replay holds the response open until the run resumes; a capped, stalled,
  or cut replay ends it without a terminal so the client re-attaches.
- A batch read the ring cannot serve returns no ring events.
- A live tail whose ring restarts under it ends without a terminal, so the client
  re-attaches and is re-synced.
- An unknown run keeps the replay_gap terminal; an unreachable worker answers 503
  so the client retries.

* fix(mothership): trim the replay ring by bytes so a long run slides instead of refusing

The append script pruned the oldest events by count only, so any stream averaging
more than ~335 B per event reached the 32 MiB owner budget before 100k events and
the refusal ended the turn. The ring now also trims its oldest events until the
retained bytes fit three quarters of the owner ceiling, refunding exactly what it
drops in the same script. A byte trim never drops a member the write adds, so an
inflated counter still refuses rather than silently discarding the new frame.

* fix(mothership): never rebuild or paint a turn from a ring that lost its head

A byte trim advances the ring past seq 1, and three callers read it from seq 0
assuming the head was there: stream recovery rebuilt a controller's context from
the tail, and both chat snapshot routes painted a truncated turn. One predicate,
startsAtReplayHead, now guards them and the reconnect route's gap check:

- recovery refuses with StreamReplayHeadTrimmedError instead of persisting a
  truncated turn;
- the snapshot routes skip the snapshot, so the client reconnects;
- the reconnect route re-syncs the view from the worker's log (stream) or serves
  no tail events (batch). When recovery was refused, a parked or stalled replay
  ends the view with recovery_unavailable instead of re-attaching forever.

The append script also trims a replayed member that lands below the ring for bytes,
keeping the tail contiguous, and caps byte trimming at 4096 members per append so an
oversized ring catches up over several appends. The budget docs now say the user
counter bounds bytes held, not bytes written per hour.

* fix(mothership): recover a headless ring from an empty context and keep log readers on the log

- Recovery no longer refuses a run whose ring lost its head, which orphaned long
  runs after a Sim deploy. It treats that ring like an expired one: the new
  controller starts from an empty context at the ring's latest seq, and re-attaches
  with an empty receipt. The worker then re-sends the whole response and re-hands its
  parked calls. Usage stays with the worker's per-run settlement; a re-attach under
  the same message identity is never a second run.
- A client re-synced from the log sends source=log from then on, so the ring never
  serves its log cursors, even after it restarts and grows past them.
- A live tail ends without a terminal as soon as its ring loses its head or
  restarts, so it re-attaches and re-syncs instead of reading re-sent text.
- A replay that ends short of the terminal and cap holds its response at least 10 s
  (longer while parked), so a stalled run is not replayed every second.
- replay_end is parsed with a schema tied to the protocol's reasons; an unknown
  reason ends the replay instead of passing as a run event.
- The replay forwarder moves into session/run-replay.ts, the chat snapshot reader is
  shared by both chat routes, and checkForReplayGap is removed.

* fix(mothership): end a held replay as soon as its run resumes or finishes, and check a busy tail's ring less often

- A replay held after a park now ends as soon as Sim sees the run leave its park, and
  any held replay ends as soon as the run reaches a terminal, so an approval no longer
  freezes the view for up to 10 s. A park Sim has not yet marked, a stall and a cut
  connection keep the 10 s floor.
- A live tail checks that its ring can still serve it only after a quiet poll or every
  eighth busy one, instead of two Redis reads on every 250 ms poll.
- The recovery integration test re-sends a replayed go tool and re-hands a Sim call
  the dead controller already ran: it is resumed with its stored result, never run
  again, and the turn keeps one block per tool.

* fix(mothership): fall back to replay_gap when the worker refuses a replay's key

A deployment whose key may not call the worker's replay (401/403) now falls back
to the replay_gap terminal as a missing run (404) does, instead of answering 503
until the client's reconnect budget runs out. A failed buffer TTL refresh during
the chat-lock heartbeat is logged as such, not as a lock-extension failure.

* fix(mothership): bound a silent replay, never skip past a trimmed cursor, and re-sync an expired ring

- The worker replay is bounded like a stream leg: no response headers, or no bytes
  including keepalives, for the idle timeout ends it so the reader re-attaches.
- A ring read that starts after the reader's next event (the ring trimmed its head
  between the gap check and the read) is never delivered; the reader re-attaches and
  re-syncs from the log, in both the live tail and batch reads.
- An empty ring serves only a reader starting from cursor 0; a live run whose buffer
  expired under a reader re-syncs from the log, and a finished one answers its
  terminal since its transcript is persisted.
- A leg that delivers events after a failure starts a fresh 30 s reachable window;
  its three retries still refill only after five minutes of delivered events.
- The two new integration suites close their worker server and restore env even
  when they are skipped.

* test(mothership): drive the run replay liveness tests through the central agent-url mock

* fix(mothership): log a worker's refusal of the run replay, and test the mid-tail trim race

- A 401 or 403 from the worker's replay endpoint is logged with its status, so a
  rotated or wrong worker key is visible instead of every reader silently falling
  back to replay_gap.
- A live tail whose ring trims past its cursor between polls ends without a
  terminal and never delivers the events after the gap.

* fix(mothership): create the log re-sync set outside render

* refactor(mothership): reuse the SSE idle timeout for the run replay, and track the log re-sync as one stream id

- processSSEStream takes an optional idle timeout and passes it to readSSELines, so the
  run replay uses the shared idle bound instead of its own reader wrapper. Callers that
  omit it are unchanged; the replay's header wait keeps its own timer.
- A chat view re-syncs one stream at a time, so the log re-sync flag is the stream's id
  rather than a set.
… outlive their period (#8461)

* fix(billing): enforce the plan usage limit mid-run and bill runs that outlive their period

Unbounded Chat runs need spend enforced inside a run, not only at its edges.

- update-cost answers every callback (200 and duplicate 409) with a top-level
  usageExceeded verdict read through the cached execution usage gate, plus the
  usageUpgrade card payload when exceeded.
- Continuation validation and the lifecycle's continuation admission read the
  original payer's spend through the same gate. A refused validation returns
  402 { code: USAGE_LIMIT_EXCEEDED, error, usageUpgrade }; a blocked account
  returns 402 { code: BILLING_BLOCKED, error } and never gets the usage card.
  A refused lifecycle continuation renders the upgrade card and stops the
  worker run so the next message after an upgrade is not refused as busy.
- Mid-run paths treat an unreadable ledger as unknown and keep the run going;
  admission before a run still fails closed.
- The card's action and copy come from one resolver shared by Sim's synthetic
  card and the verdicts the worker echoes into its durable log, with member-cap
  copy for a member over the cap their organization set.
- A cumulative charge that outlives its Stripe billing period records later
  spend in one row per later period, stamped with the payer's current period
  under a share lock on the subscription row, so a closed period is never
  topped up and the whole run is invoiced exactly once. Threshold settlement
  follows the stamped period.
- Sync the worker's billing contract (usageExceeded, UsageUpgrade,
  UsageLimitRefusal).

* fix(billing): roll forward-moved periods, lock only past period end, skip ended admissions

- Roll a cumulative charge into the payer's current period whenever that
  period starts after the latest row's, so an anchor reset or resync inside
  the old period never tops up a closed period.
- Share-lock the payer's subscription row only once the latest row's period
  has ended; before that no close can be due, and the lock would starve the
  rollover update for a busy payer.
- Mid-run usage checks report unknown (continue) once the run's admitted
  period has ended, instead of judging the old period against its allowance.
- Refuse request keys containing "@" when period rows are in play, so they
  cannot collide with another request's period rows.
- Document the mixed-version and rollback window: code that predates period
  rows can double-count a run's post-rollover spend only between a rollover
  and that period's close (at least an hour); the exposure is cents to dollars.

* fix(billing): judge long runs against the current period and close review gaps

- Mid-run usage checks judge a run that outlived its admitted period against
  the same payer's current period, re-read after a gate read that straddles
  the period end, and continue only when the current period cannot be read.
  Direct-v1 continuations read account usage through the same mid-run rules.
- Share-lock the payer's subscription on every period-aware write, so an early
  period-start move waits for an in-flight top-up.
- Reserve "@" in every cumulative idempotency key; update-cost rejects it.
- Carry a member cap through BillingLimitError to the member card, and send
  every usage-limit refusal, including a worker 402 on the first leg, through
  one handler that also stops the worker run.
- Every validate 402 now carries a declared body: USAGE_LIMIT_EXCEEDED,
  BILLING_BLOCKED, or USAGE_UNAVAILABLE for a new turn refused because usage
  could not be read.
- The update-cost verdict schema ties usageUpgrade to usageExceeded.

* fix(billing): judge mid-run usage against the payer's current period

- Mid-run checks always judge the admitted payer's current subscription
  period (cached for a minute), so an early anchor move or a rollover is
  judged against the period charges now land in; a straddling read is judged
  again against the next period.
- A direct-v1 continuation checks the payer saved in its account decision,
  against that payer's current period, never a payer chosen from the actor's
  current memberships.
- An unreadable usage read no longer reports a spent-limit message; new turns
  refused for it get neutral copy.
- Dispatch-time refusals pass the verdict scope, so a member over the cap
  their organization set gets the member card.

* chore(billing): sync the worker usage refusal contract

* fix(billing): reload an ended cached period, keep org payers org-scoped, keep blocked accounts blocked

* fix(billing): cache admitted direct-v1 continuation verdicts and bound the callback's standing read

- A direct-v1 continuation re-read the payer's full period ledger on every resume leg. Its
  admitted verdict is now served for the execution gate's TTL, with concurrent misses coalesced,
  like the attributed path; a refusal or an unreadable ledger is always read again, and the read
  is skipped when billing is off.
- The cost callback waits at most 1 s on the payer's standing, well inside the worker's 5 s
  callback timeout, and answers not exceeded past it; the abandoned read still caches its
  admission.
- The straddling-period test now reaches the re-judge branch.

* fix(billing): roll a direct-v1 run's spend into the payer's current Stripe period and judge its standing mid-run

- A direct-v1 account decision now carries the payer's subscription from admission, and a cost
  callback rolls later spend into that subscription's current period exactly as an attributed
  run does, so a run that outlives its period never tops up a closed one. Only a Stripe period
  rolls; reporting windows and free payers keep their frozen period.
- A direct-v1 cost callback reports the admitted payer's standing, and the direct gate checks
  the actor and payer for a block before their spend, so a blocked account is never paused with
  the upgrade card. The account block check moves to billing core and is shared with continuation
  validation.
- The attributed mid-run gate returns early when billing is off.
- Tests: a direct run across a rollover against real PostgreSQL, per-period token shares, and the
  usage card replay asserted on parsed segments.

* test(mothership): pin a usage refusal as non-retryable under the stream retry window

* test(billing): pin the Stripe-only rollover gate and the decision's subscription ID parsing

- A payer whose period is not a Stripe period is never rolled or share-locked.
- An account decision refuses a payer subscription ID that is not a non-empty string.
- Documents that a subscription replaced mid-run can only under-enforce the limit.

* fix(billing): judge a non-Stripe run against the period its charges land in

A reporting-window or default-period payer's charges never roll forward, so after its admitted
window ends the mid-run verdict judged an empty new window and under-enforced the limit. Both
the attributed and direct-v1 verdicts now judge the current period only for a Stripe payer,
matching the cost callback's rollover gate, and the admitted period otherwise, even after it
ends. Stripe payers keep being judged against their current period.

* test(billing): assert observable verdicts and responses instead of mock calls

- The account block, continuation delegation, cache, billing-off and rollover tests assert the
  verdict or HTTP response. Where behaviour depends on an input, the fake answers by that input,
  as the real ledger and settlement do.
- Pins against real PostgreSQL that a reporting run's top-ups after its window ends are counted in
  that window, where its request was first charged, and a later run's charges in the next.

* refactor(billing): drop unconsumed refusal plumbing and the straddle retry, and make the reporting-window test deterministic

- A new turn's 402 is empty again and the stream no longer parses 402 bodies: the worker replaces
  any validation 402 body with its own message and only polls continuation, so the new-turn codes,
  USAGE_UNAVAILABLE and the server-side refusal reasons had no consumer.
- The mid-run verdict reads its current period once; a period that ends during the read is left
  to the next callback.
- The cumulative-usage '@' check left the ledger; the cost callback already refuses such keys.
- The reporting-window integration test derives its boundary from the first row's created_at and
  waits on the database clock.
…#8305)

* feat(dashboards): add table-backed dashboard resources behind rollout flag

* refactor(files): separate discovery from storage context

* fix(charts): keep pie labels readable and separate neutral colors

* feat(dashboards): compute percentages from row conditions

* refactor(dashboards): store dashboards as workspace files

Dashboards are now ordinary workspace files, handled like Sim pages, instead
of a separate resource. Creating or uploading `<Name>.dashboard` drops the
suffix and stamps `text/x-sim-dashboard`; the type is sticky across content
writes. The file viewer renders it live behind the `dashboards` flag, and the
public share viewer shows a workspace-only notice.

- Remove the dashboard resource: sidebar page, API routes, hooks, contracts,
  application layer, Mothership dashboards/dashboard_folders tools, resource
  tags, and the per-turn dashboardsEnabled payload.
- Revert the file discovery column (0385) and drop the dashboard folder
  resource enum (0384); dashboards never shipped, so no backfill.
- Chat panel decides previewability and tab/picker icons from the file type,
  not the name, so extensionless dashboards render and get the chart icon.
- Renderer: authored left label columns are kept intact, horizontal bar
  frames grow with row count, and hovered rows get a label-and-bar highlight.
- Simplify the create-dashboard skill around one validated example.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(dashboards): address review findings on chart layout and precision

- Size horizontal bar `.chart` previews by category count like dashboard panels.
- Keep the ECharts label column for percentage bar widths, resolve percentage
  grid insets for the row highlight, and keep the time axis on the queried range.
- Show small readout values with significant digits instead of rounding to 0.
- Pass the dashboard's timezone-adjusted today to the range calendar.
- Decide the Chat panel's Markdown mode from the file record.
- Replace mock-call assertions in the EChartsView tests with DOM behavior.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(charts): size grouped bar rows and keep the row highlight through resizes

Unstacked bar series sit side by side in a category row, so grouped charts keep
the ECharts label column and their rows fit every bar slot. The row highlight
redraws the active row after each render, so a resize moves it with the plot.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(charts): keep authored tooltip arrays and drop the v2 chat mode override

Chart tooltip defaults now apply to every entry of an authored `tooltip`
array instead of replacing it with one object. Restore the v2 chat payload
to staging: the hard-coded `mode: 'agent'` came from the removed dashboard
resource work and would overwrite a resumed chat's saved mode.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(charts): reserve the authored bar gap between grouped bars

Grouped horizontal bar rows now size their gap from the series `barGap` the way
ECharts does (pixels, a percentage of bar width, default 20%, overlap for
negative gaps) instead of a fixed 4px.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(charts): format index-encoded pie tooltips and test dashboard refresh through the UI

Pie tooltips encoded by dimension index now use the custom formatter, resolving
the measure through the encode indexes ECharts passes it. The dashboard preview
test drops the mocked controls and callback-driven cases and clicks the real
Refresh button to check which analytics queries are invalidated; the range and
timezone math stays covered by the time tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(charts): format bar tooltip values with each series' own value axis

Multi-axis bar charts now read the tooltip unit from the axis each series is
plotted on, so a secondary axis no longer inherits the first axis's format.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* test(dashboards): drop mock-call and rendered-text assertions

Removes assertions on mocked collaborators and rendered text from the
dashboard, chart and analytics tests per the repository testing rules, keeping
the observable checks (status codes, thrown errors, DOM roles, computed
summaries, emitted zoom ranges). Deletes the animated-number test, which had
no behavior left to assert.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* test(dashboards): drop redundant mock resets and default environment pragmas

The shared Vitest config already resets mocks and stubbed globals before each
test and defaults to the node environment.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* refactor(dashboards): move authoring guidance to Mothership like Sim Pages

Drop the create-dashboard built-in skill and its rollout gating in the skill
lists. Mothership now learns the dashboard format from a sim-dashboards
reference in its own research-and-deliverables skill, the pattern Sim Pages
use; Sim workspace built-ins stay Agent-block documents. The dashboards flag
still gates the viewer and table analytics.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* refactor(dashboards): give DashboardFeatureGate a props interface

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* feat(dashboards): report dashboard parse errors on file writes

The v2 file create, replace and edit responses now carry `diagnostics` for a
dashboard file: its parse errors, or an empty list. Writes are never blocked,
like the page lint; table columns and queries are still checked when panels
render. Parse errors are reported as `path: message` lines, and a block with
no recognized kind names the allowed kinds and unknown keys instead of Zod's
union dump.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(dashboards): split diagnostics, switch panel query modes, return create revision

- One diagnostics entry per parse error instead of a newline-joined string.
- A panel choosing aggregate or columns drops the other mode's inherited
  dashboard fields, so shared defaults serve table and aggregate panels.
- The v2 create response returns the revision it produced, like the replace and
  edit responses.
- The analytics route test uses the shared createMockRequest helper.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(dashboards): drop inherited ordering on query-mode switches and cap diagnostics size

A panel switching away from the dashboard's query mode no longer inherits its
sort or limit, which name aggregate aliases or top-N groups in one mode and
rows in the other. Write diagnostics check the 128 KB source limit before
decoding the payload.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* revert(dashboards): restore dashboards as a separate workspace resource

Reverts "store dashboards as workspace files" and "move authoring guidance to
Mothership like Sim Pages", restoring the dashboard resource: its page,
folders, APIs, resource tags, Mothership dashboards/dashboard_folders tools,
per-turn availability, the create-dashboard built-in skill, and file
discovery. Removes the write-time diagnostics that only applied to plain
files. Keeps the chart renderer fixes, readable parse errors, panel query-mode
handling, the create-response revision, and the test cleanups. The authoring
skill keeps its simplified form with the default-formatting note. Migrations
are regenerated on top of staging in the next commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(dashboards): renumber resource migrations after staging and use central test mocks

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(dashboards): list by folder, canonical folder paths, named props

The dashboard browser now asks for one folder's dashboards instead of
filtering the first 500 across the workspace, folder paths use the shared
segment encoder, and the README names the renumbered migrations.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(dashboards): address review round on the restored resource

Fork payloads accept file discovery, dashboard queries are keyed and
invalidated per workspace, dashboard contexts open tabs and resolve in
organization chats, archived dashboard folders keep their paths, and the
folder dialog skips no-op moves and hides the edited subtree.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(dashboards): cap dashboard folders and reject folder name collisions

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* fix(dashboards): name the dashboards page props

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* feat(mothership): send entitlements instead of dashboardsEnabled

Restores the entitlement evaluator registry removed in v1.0.0 and sends
the granted list with every Mothership turn; dashboards is the first
entitlement. Sim still enforces every gated operation.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* refactor(dashboards): give agent dashboard commands the files folder syntax

Dashboards take folder paths like files: create/list --folder, move --to
(/ is the root), rename and set-content as separate actions, and folder
create/move/delete by positional path with --recursive for non-empty
deletes. The UI routes keep folder ids.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* feat(dashboards): one Sim-built dashboard per workspace

A workspace now has a single dashboard that Sim builds. The Dashboards page
renders it directly, or an empty state until the first save. Folders, the
list, create, rename and move are removed; the agent reads and saves it with
dashboards get / set, and replacing existing content needs its revision. A
partial unique index enforces one live dashboard per workspace, and generic
file writes can no longer produce the dashboard content type.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* feat(dashboards): put Dashboard under New chat in the sidebar

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* feat(dashboards): headerless dashboard page with a grid-matched icon

The dashboard page renders only the dashboard (or its empty state), so the
header, its Delete action, and the delete route, hook, and use case go away.
A new EMCN Dashboard icon is drawn on the shared sidebar icon grid and
replaces the analytics ChartColumn on dashboard surfaces.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV

* refactor(dashboards): drop file discovery for a content-type exclusion

With one dashboard per workspace, the listed/unlisted discovery column, its
migration, legacy-writer trigger, and the upload backfill script are no longer
needed. File listings, search, and workflow file pickers exclude the dashboard
content type instead, and the one-dashboard index is now migration 0392.

* refactor(dashboards): store dashboards in their own table

A dashboard is now a row in a dashboard table with its own id, keyed to its
workspace by a unique index (dropping it later allows several). Reads and
saves go through a small repository with a revision check, so the file-side
guards, exclusions, fixed file name, and dashboard file previews are gone.
Chat contexts reference dashboards by dashboardId, and saves are audited as
dashboard events.

* fix(dashboards): accept dashboard chat contexts and bound revisions

The chat request schema now admits dashboard contexts, revisions outside the
integer column range are rejected as validation errors, resource tabs only
fetch the dashboard when one is open, and the authoring guide notes that
range-pinned panels do not drag-zoom.

* fix(dashboards): named heading sizes and no redundant dashboard name lookups

The dashboard name is fixed, so resource tabs and chat chips no longer fetch
the dashboard to title it; dashboard headings use named text sizes.

* refactor(dashboards): keep the authoring reference in Mothership's skill

Mothership's create-dashboard skill now carries the dashboard syntax like
every other worker skill, so the builtin-create-dashboard workspace skill,
its source, and its flag gating are removed.

* fix(dashboards): server loading fallback and muted gate text

* chore(db): drop leftover script-migration test fixture change

---------

Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
* fix(desktop): return source connections to the app

* fix(desktop): preserve enrollment mode and completion state
* feat(search): add Zoom and Google Meet providers

* fix(search): bind meeting pagination and preserve read scopes
* fix(sandbox): preserve workbench output provenance

* fix(sandbox): preserve boundary compatibility
* fix(async-jobs): retain accepted run IDs for immediate lookup

* fix(async-jobs): finish cancellation scans after root errors
… before its first charge (#8480)

* fix(billing): keep billing a request whose period start moved forward before its first charge

A Stripe anchor reset between admission and a request's first cost callback
stamps row 0 with the reset period. The ledger binding only accepted a later
period starting at or after the admitted period's end, so every later callback
was refused with a billing-context mismatch and its spend went unbilled. The
binding now accepts the same forward-only rule the roll uses: a start later
than the admitted one. An earlier period is still refused.

Also bound the upgrade-card subscription read in update-cost under the same
1 s standing deadline as the verdict read.

* fix(billing): keep an exceeded verdict when the upgrade-card read is slow

The shared deadline discarded an exceeded verdict when the card lookup ran
past the budget. The verdict read keeps the deadline; the card lookup now
falls back to the plan-upgrade card past the same deadline.
… every tick (#8479)

* fix(mothership): close abandoned tool meters once instead of alarming every tick

A tool meter row (cost unknown) stays open when the process that owned the
tool ends mid-execution, and nothing ever closed it. The replay tick counted
those rows and logged "Service usage requires reconciliation" at ERROR on
every tick in every process, forever. Its 5-minute threshold also flagged
tools that were still legitimately running.

The replay tick now closes meters older than twice the longest tool watchdog,
keeping a pricing failure's error or recording that the tool never finished,
and logs each closed meter once with its stream, tool call and reason. Known
spend is unaffected: it is saved and delivered as separate receipts.

* fix(mothership): keep a closed tool meter final against a late tool completion
…8484)

* fix(mothership): answer a retried task wake whose turn already ran

The worker retries a task wake under the same run ID until its own run
appears. When sim ended that turn without reaching the worker (a
usage-limit refusal), every retry reopened the turn, hit the unique
stream-id constraint, and failed behind a generic message, so the worker
retried forever.

Under the chat lock, a wake whose run ID already has a sim run now
releases the lock and answers not-found, which the worker treats as a
refusal and dismisses the notification. An in-flight turn still holds the
lock and answers busy. The headless run-record catch-all now logs the
underlying insert error.

* fix(mothership): release the wake's chat lock when the run lookup fails

The retried-wake check reads copilot_runs after taking the chat lock. If that read threw, the lock stayed held until its TTL because the wake turn that releases it never started. Release with the exact lease on any throw after the acquire.

* fix(mothership): release the wake's own chat lease when the run lookup fails

* test(mothership): stub the chat lease getter in the wake unit test mock
…#8478)

* fix(mothership): explain a Chat turn the worker ends without a reason

When the worker rebuilds an ended run from its log (a resume or reattach
that reaches a run that already ended, for example at its deadline), it
sends an error terminal with no error event. Sim then fell back to the
generic "An unexpected error occurred while processing the response."
The turn now says the run had already ended and can be continued by
sending a message. A reason the worker reports, a replay refusal, and a
Stop all still take precedence.

Also:
- Log the Go stream's error text under errorMessage/detail so it no longer
  overwrites the log line's own message (stream.ts, buffer.ts).
- Rename STREAM_TIMEOUT_MS to CHAT_RUN_DEADLINE_MS and document it as the
  worker's default run deadline, now only the base of USAGE_SETTLE_MS.
- Update the byte-budget doc: a reader behind the ring trim is re-synced
  from the worker log; replay_gap is only the fallback.

* fix(mothership): keep the ended-run message surface-neutral and below a replay refusal

The fallback is shared by Chat, workflow execute and inbox, so it no longer
tells the reader to send a message. A replay refusal now wins by guard rather
than by spread order, with a test that covers the combination.

* refactor(mothership): drop the unreachable refusal guard and reuse the stream-abort test helper
…led execution (#8481)

* fix(mothership): report a browser-claimed workflow tool from its settled execution

When a browser claims a Chat workflow tool, the execute route runs the
workflow and keeps running it after the browser detaches, but only the
browser's confirmation completed the tool call. A tab that closed, lost its
network, or dropped its pagehide beacon left the Chat turn waiting for the
full client wait while the worker swept the call.

The execute route now records the bound execution's structural completion
itself once it settles, guarded on the call still running under that
execution's claim, so a browser report or background detach that lands first
is kept and nothing is delivered twice.

* fix(mothership): report queued async runs and pre-log failures of a browser-claimed workflow tool

* refactor(mothership): settle a browser-claimed workflow tool from its log status alone

* test(mothership): prove a losing settlement report publishes no confirmation
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

⚠️ Cross-repo companion check

One or more companion PRs aren't merged into main yet (aggregated across the feature PRs in this release). Merging this without them will leave copilot and sim out of sync — merge them in lockstep.

  • ⚠️ simstudioai/mothership#556 — merged into staging (this PR targets main) — feat(search): align Lucid live search contracts
  • ⚠️ simstudioai/mothership#531 — merged into staging (this PR targets main) — feat(dashboards): entitlements and the workspace dashboard commands
  • ⚠️ simstudioai/mothership#558 — merged into staging (this PR targets main) — feat(search): add Zoom and Google Meet companion contracts

* feat(analytics): add Freebuff Ads conversion tracking

* fix(analytics): drop the Freebuff click id when marketing consent is withdrawn
@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 3/5

[High risk] Adds new search providers and billing endpoints with broad API surface changes.

The PR should not merge until mid-run usage enforcement and charges arriving after terminal subscription settlement are addressed.

Findings

  1. P1 Slow reads bypass usage limits ▶
  2. P1 Closed periods receive later charges ▶

Summary

This release adds live Search providers and workspace dashboards while changing desktop connection recovery, chat replay, billing, and execution reliability.

  • Mid-run billing can fail to stop an over-limit run when standing reads repeatedly exceed the callback budget.
  • Post-terminal subscription callbacks can add spend after the period’s final settlement.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Run cost callback] --> B[Record cumulative usage]
  B --> C{Subscription period advanced?}
  C -- Yes --> D[Write new period row]
  C -- No --> E[Top up latest row]
  D --> F[Read usage standing]
  E --> F
  F --> G{Read finishes within 1 second?}
  G -- Yes, exceeded --> H[Pause run]
  G -- No --> I[Answer not exceeded]
Loading

Reviews (1) · Last reviewed commit: "fix(mothership): report a browser-claime..."

Comment thread apps/sim/app/api/billing/update-cost/route.ts Outdated
Comment thread apps/sim/lib/billing/core/usage-log.ts
* fix(integrations): complete account connections in place

* fix(integrations): preserve active account authorizations
… executing (#8482)

* fix(mothership): never sweep a Chat run while one of its Sim tools is executing

The orphaned-run sweep settled a leased run once it had gone an hour without a
status write and no controller held its chat lock. A Sim tool call writes
nothing to its run while it executes; only its execution lease heartbeat shows
it is alive. So a long tool call (a workflow run can take well over an hour)
could have its run settled as interrupted while the worker still held the run,
closing tool admission under it. The default one-hour run deadline masked this.

The sweep now skips a leased run with an unsettled, unrevoked tool execution
whose lease has not expired, both when it selects candidates and in the
guarded update. It also locks the run rows before that update, so a tool
admission (which locks its run row) either commits a lease the update then
sees, or finds the run settled and is refused.

* fix(mothership): lock a run's unsettled tool executions before the sweep settles it

A lease heartbeat writes only the tool row, so one that passed its expiry
check before the lease ran out could commit after the sweep read the old
lease and settled the run. The sweep now locks those tool rows after the
run rows, so the guarded update sees a committed renewal and a later
heartbeat finds the lease expired. Lock-holder tests release on failure.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review completed against the latest diff

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
Tip: instead of fixing issues one by one fix them all with cubic

Re-trigger cubic

Comment thread apps/sim/lib/execution/remote-sandbox/index.ts
Comment thread apps/sim/lib/mothership/request/application/controls.ts
Comment thread apps/sim/lib/mothership/request/session/abort.ts
Comment thread apps/sim/app/api/workflows/[id]/execute/route.ts
Comment thread apps/sim/lib/mothership/request/application/recover-stream.ts Outdated
Comment thread apps/sim/lib/mothership/chat/payload.ts
Comment thread apps/sim/lib/api/contracts/dashboards.ts Outdated
Comment thread apps/sim/lib/dashboards/repository.integration.ts Outdated
Comment thread apps/sim/.env.example
Comment thread apps/sim/lib/mothership/resources/types.ts
@waleedlatif1

Copy link
Copy Markdown
Collaborator

Search-provider review update: the four inline findings belonging to the provider/rollout changes have been individually assessed. Three corrections are in #8496; the organization-less workspace finding was resolved because managed connected accounts already require a canonical organization. The three valid release threads remain open until that fix reaches staging.

The companion warning also applies to these providers: mothership#556 (Lucid) and mothership#558 (Zoom/Google Meet) are included in the open worker release https://github.com/simstudioai/mothership/pull/562 but have not reached worker main. The Sim and worker releases need to stay coordinated.

* fix(integrations): close connection recovery edge cases

* fix(integrations): verify persisted connection outcomes
…adout, isolate repository tests (#8499)

* fix(dashboards): carry dashboard chat context ids, label the chart readout, isolate repository tests

- User message contexts keep a dashboard mention's dashboardId, both in the
  optimistic message and when a persisted message is reopened, matching the
  context the server stores.
- The time-series readout row has role="group", so its aria-label is exposed to
  assistive technology.
- The revision test in the dashboard repository suite seeds its own workspace
  instead of depending on the previous test's row.

* fix(dashboards): name the chart readout group only when it has values
…d model the empty new-turn 402 (#8502)

* fix(billing): build the mid-run usage card from the admitted payer and model the empty new-turn 402

A direct-v1 run's mid-run verdict reads the payer saved in its account
decision, but the upgrade card was resolved from the actor's current
subscription, so a payer/actor mismatch picked the wrong action and copy.
The exceeded account verdict now carries the payer and subscription it
already read, and update-cost and the validate continuation pass it to
resolveUsageUpgradePayload instead of a second lookup.

The validate contract declared every 402 as a JSON refusal, while a new
turn's 402 has no body; the 402 schema now allows the empty body. The
wire is unchanged.

* fix(billing): drop the unreachable card-read deadline from the usage upgrade card

Every exceeded verdict that reaches update-cost now carries its payer, so the
deadline-bounded actor subscription lookup could no longer run. Remove the
parameter, its call-site argument, the stale TSDoc, and the test that passed
without exercising it.

* test(copilot): pin the new-turn usage refusal to an empty 402
…loor the replay TTL (#8501)

* fix(mothership): number recovered turns past an unreadable ring and floor the replay TTL

A recovered run whose replay ring read back empty while its seq counter
survived (every retained entry unreadable) resumed numbering at 0, writing
over the old range and moving the counter backwards past readers' cursors.
Resume from the counter whenever no event was recovered.

COPILOT_STREAM_TTL_SECONDS below the 20 s chat-lock heartbeat let an idle
live buffer expire between refreshes. Floor it at 60 s.

* test(mothership): leave a margin for Redis TTL rounding in the live-buffer floor check
* fix(search): recheck Zoom approval and bound MCP serialization

* fix(search): preserve byte-only MCP response limits

* fix(search): reject inherited JSON serializers
…act bounds (#8495)

* fix(mothership): order dashboards between chats and tables in resource menus

* chore(rules): drop the resource-menu sidebar-mirroring rule

* fix(dashboards): scope entitlements, keep mention ids, and tighten contracts

---------

Co-authored-by: Waleed Latif <[email protected]>
…8503)

* fix(sandbox): redact temporary session credentials in model output

* fix(sandbox): unify output handling and avoid response copies

* fix(sandbox): scan session output without recursion
…ready invoiced (#8505)

* fix(billing): refuse charges into a period the terminal settlement already invoiced

A subscription's deletion settles its terminal period at once (claim, overage,
final invoice, bookkeeping), but recordCumulativeUsage kept topping up that
period's row for a still-running request because the subscription's period
never rolls after deletion. That spend was never billed.

The terminal claim now advances the close marker to the period's end under its
FOR UPDATE lock, and recordCumulativeUsage reads the marker in its existing FOR
SHARE read: a charge whose target period ends at or before the marker throws
CumulativeUsagePeriodClosedError, which update-cost answers with the existing
non-retryable BILLING_PERIOD_ELAPSED outcome, so the worker quarantines the leg
for reconciliation instead of the spend disappearing into an invoiced period.

No migration: the marker is an existing column, and every reader treats a
marker at or past periodStart as current, so v0.9.6 behaves unchanged.

* test(billing): cover a terminal claim overlapping an in-flight charge

Both lock orders against real PostgreSQL: a claim waits for an in-flight
charge so the final sum includes it, and a charge that waited on an
in-flight claim is refused.
@waleedlatif1
waleedlatif1 merged commit 2a7f85d into main Oct 1, 2026
73 checks passed

This branch was successfully deployed

1 active deployment
Preview — c57c9058 Deployed Oct 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

requires-mothership-merge Has a companion PR on the mothership/copilot side — merge in lockstep

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants