v0.9.8: search cleanup recovery, new search providers, billing and chat reliability - #8494
Conversation
* docs(library): update best-ai-automation-tools-2026 * Pi Babysit: address PR #8465 feedback --------- Co-authored-by: Sim Pi Agent <[email protected]>
* docs(library): update best-gumloop-alternatives-in-2026 * Pi Babysit: address PR #8466 feedback --------- Co-authored-by: Sim Pi Agent <[email protected]>
… 2026 (#8467) * feat(library): Best AI Agent Builders for Slack and CRM Automation in 2026 * Pi Babysit: address PR #8467 feedback --------- Co-authored-by: Sim Pi Agent <[email protected]>
…ch Does Your Team Need? (#8468) * feat(library): Marketing Automation Platform vs AI Agent Builder: Which Does Your Team Need? * Pi Babysit: address PR #8468 feedback --------- Co-authored-by: Sim Pi Agent <[email protected]>
…r CPU load (#8458) The four-patch turn-budget test drove 2,400 deltas over a ~300 KB base, about 2.4 s of CPU on an idle machine and 11-15 s at load 150, past the 10 s test timeout. Pacing already runs on fake timers; the flake was CPU time, not wall clock. Four 30 s patches at a 500 ms tick still stream ~25 MB uncapped against the 8 MiB budget, so the test still fails when the cap is removed.
…8462) * fix(sim-cli): report an embedded SimApiError with its own exit code The embedded renderer returned 1 for every SimApiError, so an operation wait that timed out (exit 4) or needs configuration (exit 3) read as a generic failure in-process while the installed CLI reported the real status. Return error.exitCode, as the terminal renderer does. * test(sim-cli): make the embedded wait-timeout test independent of runner speed The test gave the wait 10 ms of real time, so a slow runner could reach the deadline before the first status check and exit 4 without printing the receipt. The clock now advances only when the stubbed status is served, so the wait times out only after it has a receipt to print.
… file mount (#8456) * fix(sandbox): withhold workbench certification after an unprovenanced file mount A persistent chat workbench stays "clean" only while every input it received was classified secret-free, and the scratch-file read hands its bytes to the model on that basis. File mounts resolved from platform file objects were counted only when a provenance source existed, so a mount whose key has no canonical metadata record (or no principal to bind one) left the machine certified. Both the `files` parameter and a mount marker in context variables reach this resolver from model-supplied Function parameters. The resolver now reports how many mounts had no provenance source. When a workbench session receives any, the session request carries `unprovenancedInputs` and the code boundary records the machine as unknown. Workflow runs have no session and keep their existing absence policy. * test(sandbox): certify workbench history against real Redis and close the mount bypass - Replace the certification unit test, which restated the history script in a fake Redis, with an integration suite that runs the real code boundary against the real script in a disposable Redis. - Have the route test's mocked mount resolver return a fixed count per test rather than restating the counting rule. - Strip model-supplied `_sandboxFiles` from Copilot Function calls. Only resolved inputs may populate it, and a supplied URL mount would skip their provenance. - Document that public storage contexts always count as unprovenanced mounts. * test(sandbox): restore only the env the certification suite changed The suite's cleanup assigned undefined to REDIS_URL when it had been unset, which stores the string "undefined" for later suites in the worker, and asked for a Redis client even when the suite was skipped. * test(sandbox): close the shared Redis client after the certification suite
* fix(mothership): settle Chat runs no controller will finish A run whose process died before finalize, whose controller was superseded with no successor, or that was stopped while no controller existed stayed unfinished forever: its chat marker kept pointing at it, so the chat read as busy and a reconnect polled a run nothing would ever end. - Stop now settles the run as cancelled once no controller of its stream holds the chat lock, including after it force-releases a controller that did not exit in time. - The stale-execution cron settles leased runs whose stream holds no chat lock, has no replay buffer left, and has been idle past the orchestration budget (so a reconnect has nothing left to resume), and runs without a lease once idle for 24 hours. A run Stop already closed settles as cancelled, any other as error; its chat marker is released. - Each settle is one conditional update on the run row that requires it to be unfinished, idle, and still naming the controller that was observed, so a finalizing controller or a successor's claim wins or loses against it atomically and the run settles exactly once. * fix(mothership): lock chats before runs when settling orphans, and label them precisely - Settling now locks the affected chat rows first, in id order, as a controller's claim does. Locking the run and then the chat deadlocked against a concurrent reconnect claim. - Each sweep batch fails on its own, the chat markers of a batch clear in one statement, and a sweep settles at most 5k rows with a short pause between full batches. - A run settles as cancelled only when its user pressed Stop. A newer turn also closes tool admission on older runs, and those now settle as errors. - Runs without a controller lease keep their last write as their completion and retention time and read "never finalized (no controller lease)". - The leased-run grace no longer derives from the orchestration deadline. Liveness comes only from the heartbeat-renewed chat lock; the grace and the replay TTL only bound how long a reconnect can resume a dead run. * fix(mothership): never sweep a current headless run A headless turn has no chat lease and no heartbeat, so its age says nothing about whether it is still running once runs have no deadline. The sweep's lease-less rule now applies only to runs admitted before the current tool-execution protocol: every run the current code admits records the current version, so after a deploy no such row can be live. A current headless run is left to its own lifecycle, which always settles it. The protocol version moves beside the other async-run constants so the sweep can read it without importing the repository. * refactor(mothership): require the recorded Stop inside the stopped-run settle - Settling a stopped run now passes the Stop-row check as the update's own guard, so it cannot cancel a run nobody stopped; the separate stopped branch is gone and every settle derives cancelled from the Stop row. - Chat lock ownership is read through getChatStreamLockOwners and trusted only when verified, instead of a second Redis read of the same keys. - Settle transactions use the shared DbTransaction type. * fix(mothership): fence orphan settlement on the chat lock and resume sweeps where they stopped - The sweep takes each unowned leased run's chat lock under the run's own stream before settling it and releases it after the commit, so a reconnect can no longer lock the chat between the ownership check and the settle and then lose its claim; a reconnect that meets the fence retries. - A sweep examines at most 10k candidates and settles at most about 5k, resuming from a cursor saved in Redis and wrapping to the first run, so runs that cannot be settled yet never starve the ones after them. - Every settled run whose chat marker was released is announced, legacy runs included, so an open client stops showing the chat as busy. * fix(knowledge): record a terminal status for every Slack Assistant run The Slack Assistant admits its own run row and only ever marked it as an error, so every completed or stopped Slack turn stayed active. It now records the terminal status once, after the turn ends, through the shared run-status update: complete on success, cancelled when its user stopped it (in Slack or in Sim), and error otherwise. Every other run-creating path already settles its run: interactive turns through their controller's finalize, and headless turns that admit their own run in the lifecycle's own finally. * fix(knowledge): record the Slack run's status after its turn is saved - The Slack Assistant now writes its run's one terminal status after its outcome and response are persisted, from the final outcome, so a failed save ends the run as an error instead of complete. - A Stop lookup that fails no longer skips that write: the turn is treated as not stopped, logged, and settled as an error. - The orphaned-run suite deletes the sweep cursor before each test and in teardown, so no later suite starts from its leftover position.
…under its statement timeout (#8460) * fix(search): bound retirement pages by mutated rows so they finish under the statement timeout A retirement page read 25,000 IDs and updated or deleted every target row among them in one statement. Retiring a document is a non-HOT update that writes every index on `document`, and a deleted chunk cascades into its projections, so on a KB that dominates the table a page's write cost, not its scan, outran the two-minute statement timeout and failed the deploy migration. Each page now mutates at most a row limit of its target rows. A page that reaches the limit advances the cursor only to its last mutated row, and already-retired documents never spend the limit. The limit starts at 2,000, halves after a slow page or a statement timeout (the timed-out page rolls back with its cursor and is retried), and doubles after a fast full page. The completion rechecks, which walk every captured KB once, run with a 30-minute timeout. The retirement stays idempotent and resumes from its saved cursor. * fix(search): shrink only timed-out page mutations and pace Search retirement pages Only a statement timeout from a page's mutating statement now halves the row limit and retries the rolled-back page. Any other timeout, such as a completion recheck, fails the run at once instead of repeating the same statement at every smaller limit. Pages are timed around the whole call, commit included, so the synchronous-replication wait counts toward the slow-page threshold. Each page is followed by a pause as long as the page, up to five seconds, and the row limit is capped at 8,000. Phase changes no longer adjust the limit. The progress log now carries the phase, cursor and rows mutated, and the migration logs slow-page halvings, phase changes and the start of the completion recheck. * test(search): prove retired documents never spend the retirement row limit A run of already-retired Search documents longer than the row limit must be crossed in one page. The test counts documents-phase statements and fails if the page filter on unretired rows is removed. * fix(search): scale the retirement scan window with the row limit and time out resumed rechecks like completion A capped page re-reads its scan from its last mutated row, so a fixed 25,000-ID window re-read most of the same IDs on every page once the row limit shrank. Each page now reads at most four IDs per row of its limit, capped at 25,000, and any fast page doubles the limit so sparse stretches widen the window again. A retry that finds retirement already complete revalidates every captured KB, as completion does, so it now runs under the same 30-minute timeout instead of the two-minute page timeout. * fix(search): never grow a Search retirement page back to a size that timed out, and pause after a timeout A timed-out page halved the row limit, but one fast page doubled it straight back, so the run alternated between the size that timed out and half of it, rolling back a full statement-timeout page each time. The limit now grows only up to half of the smallest size that timed out, and a timed-out page is followed by the same pause as any other page.
…c execute responses until the log is final (#8473) * fix(logs): record a run's cost before it reads finished, and hold sync execute responses until the log is final * fix(logs): stop the ledger-order reader when a completion throws * fix(logs): keep a completion failure visible when the ledger-order reader also fails * fix(logs): let the ledger-order reader observe every finished run * fix(logs): assert ledger-before-terminal ordering deterministically at the ledger lock * fix(logs): scope the ledger-lock waiter to the execution and surface early completion failures
) * fix(mothership): keep Chat stream legs alive past the hour without a wall clock - End the replay GET at its cap without a terminal event, so the client re-attaches from its cursor instead of ending a live turn with resume_timeout. - Replace the per-leg worker SSE wall clock with an idle timeout (WORKER_STREAM_IDLE_TIMEOUT_MS, 120 s) that fails a silent leg as a retryable interruption; a caller-set timeout still applies. - Let StreamRetryWindow run without a deadline by default and replenish its reachable budget only after five minutes of healthy streaming, never per event. - Split the tool watchdog, permission wait, client tool wait, and delegation TTL onto their own constants and remove ORCHESTRATION_TIMEOUT_MS. - Refresh the replay buffer TTLs from the chat-lock heartbeat, and report a replay gap for a cursor ahead of a buffer whose numbering restarted. - Document why maxDuration stays on the chat POST and execute routes. * fix(mothership): renew the Chat reconnect budget after a tail delivers events A reconnect attempt whose re-attached tail delivered new events now restarts the retry budget at the base delay, so separate network drops hours apart in a long turn no longer add up to the ten-attempt exhaustion. * fix(mothership): refresh the stream byte counter with its buffer, and never revive a closed buffer - The chat-lock heartbeat now slides the owner byte counter's TTL along with the replay buffer's. Otherwise, after a park longer than an hour, the counter expired while the ring survived: the ring could grow to about twice its target, and its refunds drained the user counter. - Scheduling a finished stream's cleanup now marks it closed, and the refresh leaves a closed stream alone. A heartbeat still in flight at teardown can no longer re-extend a buffer whose cleanup was already scheduled. - Describe the worker idle timeout in terms of network intermediaries, and state that the delegation TTL reuses the long-running tool watchdog's cap. * fix(mothership): bound a stalled worker error body and refill retries only on real progress - A non-OK worker response's body is read under the same idle bound as the leg, so a stalled error body can no longer block the turn forever. - The reachable retry budget refills only when a leg's delivered events span five minutes, from its first event to its latest. A leg that delivered one event and then only kept alive has made no progress and no longer refills it. * fix(mothership): mark a closing buffer before expiring it, and report a capped replay as ended - scheduleBufferCleanup sets the closed marker before shortening the TTLs, so a heartbeat refresh that lands mid-pipeline is either overridden or sees the marker. - A reconnect that ends at its cap is reported as ended without a terminal, not as a client disconnect: the outcome is read before the route closes its own stream. - The buffer TTL integration test keeps a 5 s TTL against a 12 s park, so a slow runner cannot expire the buffer between heartbeats.
… workflow block logs (#8459) * fix(mothership): keep a tainted mount's verdict across the run_code crossing A run_code / run_function that mounted a workspace file whose sidecar is `unknown` latched its per-call registry through the crossing with only `source-provenance-incomplete` and `inherited-incomplete-source`. Both are in the registry's absence set, so the withheld result named no guard, and an output file written from that registry (outputs.files[].path) was recorded as `unrecorded` absence instead of taint. The crossing now inherits the mounted registry's own reasons when it latched, so the refusal names `mounted-file-provenance-unavailable` and writers keep the taint. The mount refusal also reports the file id. No projection is relaxed: the result is still withheld, exact mounts still redact, clean mounts are unchanged. * improvement(mothership): tell the model why a tool result was withheld A withheld result reached the model as a bare `{ success: true }`, so the agent could not tell a tainted input from an oversized payload and retried or guessed. The withheld result now carries `withheldReason`, chosen from code-defined wording by the guard that tripped: an input file, table, or document with unknown secret provenance; provenance that could not be verified; or content that could not be checked. It rides the existing `resultWithheld` disclosure beside any effect ids, and the key is reserved so an id cannot displace it. No content, reason literal, or origin crosses; an absent registry still carries nothing. * improvement(mothership): log the size of a withheld tool result A `content-refused` withholding said only that a complete registry refused the payload. It can be refused by its encoded size or by the number of values the projection walks. Row-shaped payloads reach the 100k-value traversal cap well before the 16 MiB byte cap, and only while the call has an active secret. The withheld log lines now report the result's encoded bytes and value count, so a refusal names which cap it hit. Numbers only; no content is logged. * fix(mothership): bound run_workflow block-log outputs before the secret projection A run_workflow result echoes every block's output in `logs`. A run with many row-shaped block outputs can exceed the projection's 100k-value traversal cap while staying under its byte cap. Whenever the call had an active secret, the whole result was then withheld, including the final output and error, and the model saw a bare success. Block-log outputs are now bounded to a quarter of each projection cap (`MAX_CONTENT_NODES` values and the default byte cap). Past the budget, the bulkiest outputs are replaced, largest first, with a `logs get <executionId> --trace` pointer, in the same form the oversized-input compaction already uses. The executionId, status, final output, error, and `select` values are untouched, and the other three quarters of each cap stay free for them. * fix(mothership): bound browser-run workflow logs and keep the pointer length-blind run_workflow can complete in the browser, and that restoration projected the raw block logs, so a large browser-run workflow was still withheld. The block-log compaction now lives in the shared workflow-output module and applies to both the server handler and the client restoration when no `select` is given. `select` still reads full values. The omitted-output pointer no longer reports bytes. They were measured before secret projection and so disclosed a secret's length. It reports the value count and says "inspect with logs get" rather than promising the full value. One `measureModelContent` in the projection module now serves both the compaction and the withheld-size logging. It stops counting at the value cap and tolerates values JSON cannot encode. The byte cap is exported beside `MAX_CONTENT_NODES`. * fix(mothership): resolve a browser run's select before the secret projection The browser-run restoration projected the full raw block logs and only then applied `select`, so a large run was still withheld whenever a selector was given. The server handler selects first. Both paths now build their log fields through one shared `presentWorkflowLogsForModel`, from raw logs and before projection: a `select` resolves against the full logs and replaces them; otherwise the echoed logs are bounded. Selected values are still projected, so a selected secret is redacted. * test(providers): expect the withheld reason on provider tool results The provider tool path projects through the same withheld-result shape, so an omitted model result now carries `resultWithheld` and its fixed-wording `withheldReason`. The expectations still asserted the old empty output. * fix(mothership): bound lifted outputs, length-blind input markers, and a walk-based measure - run_block and run_workflow_until_block lift the stopping block's output into `output`. That copy is now bounded with the same budget and pointer as a block-log output, so one huge block no longer gets the whole response withheld. - The truncated-input marker no longer carries the input's length. It is written before secret projection, so the length disclosed a secret's length. - `measureModelContent` now walks the value by JSON's rules instead of serializing it. It stops at the first value, byte, or depth limit it passes and reports `exceeded`, measuring a string only while it fits the remaining byte budget. An output past a limit, including one nested past the projection's depth limit, becomes a pointer instead of voiding the run. - The browser-run restoration now truncates echoed block inputs as the server handler does: both paths build their log fields through `presentWorkflowLogsForModel`. * fix(mothership): keep no raw prefix in a truncated block input A truncated block input kept its first 200 raw characters, written before secret projection. A secret straddling that cut left a fragment that whole-literal redaction cannot match, so part of the secret reached the model. The marker now keeps nothing of the raw input: `…[input omitted; inspect with logs get <executionId> --trace]`. Inputs over the limit are echoed upstream data the caller already has or can fetch, so the preview carried little. * fix(mothership): bound run outputs only when the projection walks them against a secret Without an active secret the model-facing projection passes JSON through under its byte cap alone, so the block-output budget turned a large lifted run_block output into a pointer for no reason. Output compaction now applies only when the call's registry makes the projection walk the result, and a lifted output keeps the final output's share beside the bounded logs. * fix(mothership): measure a run's whole model-facing result and bound its final output A fixed share for the lifted output left no room for the log entries, envelope and error the projection also counts, and a Response block's final output was never bounded, so a run near the caps was still withheld whole. The assembled result is now measured as the projection will walk it, and a final output that would push it past a cap is replaced with a pointer, on the server and browser-run paths alike. An unencodable result is still refused as before. * fix(mothership): leave an unencodable block output for the projection to refuse The output bound handles size only. A block output JSON cannot encode cannot be checked, so it is no longer sized past the budget and replaced with a pointer; the projection refuses it as it did before this PR. * fix(mothership): leave a result that fits the caps, and every result without a secret, as staging returns it Block-log outputs were bounded whenever a secret was active, so a 30k-value or 4.5 MB result that crossed in full before became pointers. Outputs are now bounded only when the whole result would pass a projection cap. Without an active secret the server handler keeps its preview input marker and the browser-run path leaves its logs untouched, in their original position, so a result with no secret is byte-identical to before; only a walked call gets the marker that keeps nothing. * test(mothership): pin the unencodable-output guard and the error in the whole-result measure The unencodable-output test placed the BigInt in the lifted output, which the whole-result walk reaches first, so it passed without the guard. It now puts a bulky log ahead of the BigInt log. A new test fails if the error is left out of the measure the result is bounded against.
* feat(search): add Lucidchart and Lucidspark live search * fix(search): validate Lucid queries and isolate stale candidates
… the replay ring by bytes (#8469) * chore(mothership): sync the worker protocol for the read-only run replay Adds StreamReplayRequest and StreamReplayEnd from the worker's contracts (bun run contracts:sync). * fix(mothership): re-sync a reconnect the replay ring cannot serve from the worker log A reconnect whose cursor fell behind the ring, a fresh tab reading a ring that lost its head, and a cursor ahead of a ring whose numbering restarted now stream the run from the worker's read-only replay instead of ending the turn with replay_gap or replaying a partial response. - The reconnect route opens POST /api/streams/replay with no receipt and forwards it under its own cursors from 1, with x-mothership-stream-replay: log so the client rebuilds the turn from an empty response. A response on the replay is never handed back to the ring, which shares no position with the log. - A parked replay holds the response open until the run resumes; a capped, stalled, or cut replay ends it without a terminal so the client re-attaches. - A batch read the ring cannot serve returns no ring events. - A live tail whose ring restarts under it ends without a terminal, so the client re-attaches and is re-synced. - An unknown run keeps the replay_gap terminal; an unreachable worker answers 503 so the client retries. * fix(mothership): trim the replay ring by bytes so a long run slides instead of refusing The append script pruned the oldest events by count only, so any stream averaging more than ~335 B per event reached the 32 MiB owner budget before 100k events and the refusal ended the turn. The ring now also trims its oldest events until the retained bytes fit three quarters of the owner ceiling, refunding exactly what it drops in the same script. A byte trim never drops a member the write adds, so an inflated counter still refuses rather than silently discarding the new frame. * fix(mothership): never rebuild or paint a turn from a ring that lost its head A byte trim advances the ring past seq 1, and three callers read it from seq 0 assuming the head was there: stream recovery rebuilt a controller's context from the tail, and both chat snapshot routes painted a truncated turn. One predicate, startsAtReplayHead, now guards them and the reconnect route's gap check: - recovery refuses with StreamReplayHeadTrimmedError instead of persisting a truncated turn; - the snapshot routes skip the snapshot, so the client reconnects; - the reconnect route re-syncs the view from the worker's log (stream) or serves no tail events (batch). When recovery was refused, a parked or stalled replay ends the view with recovery_unavailable instead of re-attaching forever. The append script also trims a replayed member that lands below the ring for bytes, keeping the tail contiguous, and caps byte trimming at 4096 members per append so an oversized ring catches up over several appends. The budget docs now say the user counter bounds bytes held, not bytes written per hour. * fix(mothership): recover a headless ring from an empty context and keep log readers on the log - Recovery no longer refuses a run whose ring lost its head, which orphaned long runs after a Sim deploy. It treats that ring like an expired one: the new controller starts from an empty context at the ring's latest seq, and re-attaches with an empty receipt. The worker then re-sends the whole response and re-hands its parked calls. Usage stays with the worker's per-run settlement; a re-attach under the same message identity is never a second run. - A client re-synced from the log sends source=log from then on, so the ring never serves its log cursors, even after it restarts and grows past them. - A live tail ends without a terminal as soon as its ring loses its head or restarts, so it re-attaches and re-syncs instead of reading re-sent text. - A replay that ends short of the terminal and cap holds its response at least 10 s (longer while parked), so a stalled run is not replayed every second. - replay_end is parsed with a schema tied to the protocol's reasons; an unknown reason ends the replay instead of passing as a run event. - The replay forwarder moves into session/run-replay.ts, the chat snapshot reader is shared by both chat routes, and checkForReplayGap is removed. * fix(mothership): end a held replay as soon as its run resumes or finishes, and check a busy tail's ring less often - A replay held after a park now ends as soon as Sim sees the run leave its park, and any held replay ends as soon as the run reaches a terminal, so an approval no longer freezes the view for up to 10 s. A park Sim has not yet marked, a stall and a cut connection keep the 10 s floor. - A live tail checks that its ring can still serve it only after a quiet poll or every eighth busy one, instead of two Redis reads on every 250 ms poll. - The recovery integration test re-sends a replayed go tool and re-hands a Sim call the dead controller already ran: it is resumed with its stored result, never run again, and the turn keeps one block per tool. * fix(mothership): fall back to replay_gap when the worker refuses a replay's key A deployment whose key may not call the worker's replay (401/403) now falls back to the replay_gap terminal as a missing run (404) does, instead of answering 503 until the client's reconnect budget runs out. A failed buffer TTL refresh during the chat-lock heartbeat is logged as such, not as a lock-extension failure. * fix(mothership): bound a silent replay, never skip past a trimmed cursor, and re-sync an expired ring - The worker replay is bounded like a stream leg: no response headers, or no bytes including keepalives, for the idle timeout ends it so the reader re-attaches. - A ring read that starts after the reader's next event (the ring trimmed its head between the gap check and the read) is never delivered; the reader re-attaches and re-syncs from the log, in both the live tail and batch reads. - An empty ring serves only a reader starting from cursor 0; a live run whose buffer expired under a reader re-syncs from the log, and a finished one answers its terminal since its transcript is persisted. - A leg that delivers events after a failure starts a fresh 30 s reachable window; its three retries still refill only after five minutes of delivered events. - The two new integration suites close their worker server and restore env even when they are skipped. * test(mothership): drive the run replay liveness tests through the central agent-url mock * fix(mothership): log a worker's refusal of the run replay, and test the mid-tail trim race - A 401 or 403 from the worker's replay endpoint is logged with its status, so a rotated or wrong worker key is visible instead of every reader silently falling back to replay_gap. - A live tail whose ring trims past its cursor between polls ends without a terminal and never delivers the events after the gap. * fix(mothership): create the log re-sync set outside render * refactor(mothership): reuse the SSE idle timeout for the run replay, and track the log re-sync as one stream id - processSSEStream takes an optional idle timeout and passes it to readSSELines, so the run replay uses the shared idle bound instead of its own reader wrapper. Callers that omit it are unchanged; the replay's header wait keeps its own timer. - A chat view re-syncs one stream at a time, so the log re-sync flag is the stream's id rather than a set.
… outlive their period (#8461) * fix(billing): enforce the plan usage limit mid-run and bill runs that outlive their period Unbounded Chat runs need spend enforced inside a run, not only at its edges. - update-cost answers every callback (200 and duplicate 409) with a top-level usageExceeded verdict read through the cached execution usage gate, plus the usageUpgrade card payload when exceeded. - Continuation validation and the lifecycle's continuation admission read the original payer's spend through the same gate. A refused validation returns 402 { code: USAGE_LIMIT_EXCEEDED, error, usageUpgrade }; a blocked account returns 402 { code: BILLING_BLOCKED, error } and never gets the usage card. A refused lifecycle continuation renders the upgrade card and stops the worker run so the next message after an upgrade is not refused as busy. - Mid-run paths treat an unreadable ledger as unknown and keep the run going; admission before a run still fails closed. - The card's action and copy come from one resolver shared by Sim's synthetic card and the verdicts the worker echoes into its durable log, with member-cap copy for a member over the cap their organization set. - A cumulative charge that outlives its Stripe billing period records later spend in one row per later period, stamped with the payer's current period under a share lock on the subscription row, so a closed period is never topped up and the whole run is invoiced exactly once. Threshold settlement follows the stamped period. - Sync the worker's billing contract (usageExceeded, UsageUpgrade, UsageLimitRefusal). * fix(billing): roll forward-moved periods, lock only past period end, skip ended admissions - Roll a cumulative charge into the payer's current period whenever that period starts after the latest row's, so an anchor reset or resync inside the old period never tops up a closed period. - Share-lock the payer's subscription row only once the latest row's period has ended; before that no close can be due, and the lock would starve the rollover update for a busy payer. - Mid-run usage checks report unknown (continue) once the run's admitted period has ended, instead of judging the old period against its allowance. - Refuse request keys containing "@" when period rows are in play, so they cannot collide with another request's period rows. - Document the mixed-version and rollback window: code that predates period rows can double-count a run's post-rollover spend only between a rollover and that period's close (at least an hour); the exposure is cents to dollars. * fix(billing): judge long runs against the current period and close review gaps - Mid-run usage checks judge a run that outlived its admitted period against the same payer's current period, re-read after a gate read that straddles the period end, and continue only when the current period cannot be read. Direct-v1 continuations read account usage through the same mid-run rules. - Share-lock the payer's subscription on every period-aware write, so an early period-start move waits for an in-flight top-up. - Reserve "@" in every cumulative idempotency key; update-cost rejects it. - Carry a member cap through BillingLimitError to the member card, and send every usage-limit refusal, including a worker 402 on the first leg, through one handler that also stops the worker run. - Every validate 402 now carries a declared body: USAGE_LIMIT_EXCEEDED, BILLING_BLOCKED, or USAGE_UNAVAILABLE for a new turn refused because usage could not be read. - The update-cost verdict schema ties usageUpgrade to usageExceeded. * fix(billing): judge mid-run usage against the payer's current period - Mid-run checks always judge the admitted payer's current subscription period (cached for a minute), so an early anchor move or a rollover is judged against the period charges now land in; a straddling read is judged again against the next period. - A direct-v1 continuation checks the payer saved in its account decision, against that payer's current period, never a payer chosen from the actor's current memberships. - An unreadable usage read no longer reports a spent-limit message; new turns refused for it get neutral copy. - Dispatch-time refusals pass the verdict scope, so a member over the cap their organization set gets the member card. * chore(billing): sync the worker usage refusal contract * fix(billing): reload an ended cached period, keep org payers org-scoped, keep blocked accounts blocked * fix(billing): cache admitted direct-v1 continuation verdicts and bound the callback's standing read - A direct-v1 continuation re-read the payer's full period ledger on every resume leg. Its admitted verdict is now served for the execution gate's TTL, with concurrent misses coalesced, like the attributed path; a refusal or an unreadable ledger is always read again, and the read is skipped when billing is off. - The cost callback waits at most 1 s on the payer's standing, well inside the worker's 5 s callback timeout, and answers not exceeded past it; the abandoned read still caches its admission. - The straddling-period test now reaches the re-judge branch. * fix(billing): roll a direct-v1 run's spend into the payer's current Stripe period and judge its standing mid-run - A direct-v1 account decision now carries the payer's subscription from admission, and a cost callback rolls later spend into that subscription's current period exactly as an attributed run does, so a run that outlives its period never tops up a closed one. Only a Stripe period rolls; reporting windows and free payers keep their frozen period. - A direct-v1 cost callback reports the admitted payer's standing, and the direct gate checks the actor and payer for a block before their spend, so a blocked account is never paused with the upgrade card. The account block check moves to billing core and is shared with continuation validation. - The attributed mid-run gate returns early when billing is off. - Tests: a direct run across a rollover against real PostgreSQL, per-period token shares, and the usage card replay asserted on parsed segments. * test(mothership): pin a usage refusal as non-retryable under the stream retry window * test(billing): pin the Stripe-only rollover gate and the decision's subscription ID parsing - A payer whose period is not a Stripe period is never rolled or share-locked. - An account decision refuses a payer subscription ID that is not a non-empty string. - Documents that a subscription replaced mid-run can only under-enforce the limit. * fix(billing): judge a non-Stripe run against the period its charges land in A reporting-window or default-period payer's charges never roll forward, so after its admitted window ends the mid-run verdict judged an empty new window and under-enforced the limit. Both the attributed and direct-v1 verdicts now judge the current period only for a Stripe payer, matching the cost callback's rollover gate, and the admitted period otherwise, even after it ends. Stripe payers keep being judged against their current period. * test(billing): assert observable verdicts and responses instead of mock calls - The account block, continuation delegation, cache, billing-off and rollover tests assert the verdict or HTTP response. Where behaviour depends on an input, the fake answers by that input, as the real ledger and settlement do. - Pins against real PostgreSQL that a reporting run's top-ups after its window ends are counted in that window, where its request was first charged, and a later run's charges in the next. * refactor(billing): drop unconsumed refusal plumbing and the straddle retry, and make the reporting-window test deterministic - A new turn's 402 is empty again and the stream no longer parses 402 bodies: the worker replaces any validation 402 body with its own message and only polls continuation, so the new-turn codes, USAGE_UNAVAILABLE and the server-side refusal reasons had no consumer. - The mid-run verdict reads its current period once; a period that ends during the read is left to the next callback. - The cumulative-usage '@' check left the ledger; the cost callback already refuses such keys. - The reporting-window integration test derives its boundary from the first row's created_at and waits on the database clock.
…#8305) * feat(dashboards): add table-backed dashboard resources behind rollout flag * refactor(files): separate discovery from storage context * fix(charts): keep pie labels readable and separate neutral colors * feat(dashboards): compute percentages from row conditions * refactor(dashboards): store dashboards as workspace files Dashboards are now ordinary workspace files, handled like Sim pages, instead of a separate resource. Creating or uploading `<Name>.dashboard` drops the suffix and stamps `text/x-sim-dashboard`; the type is sticky across content writes. The file viewer renders it live behind the `dashboards` flag, and the public share viewer shows a workspace-only notice. - Remove the dashboard resource: sidebar page, API routes, hooks, contracts, application layer, Mothership dashboards/dashboard_folders tools, resource tags, and the per-turn dashboardsEnabled payload. - Revert the file discovery column (0385) and drop the dashboard folder resource enum (0384); dashboards never shipped, so no backfill. - Chat panel decides previewability and tab/picker icons from the file type, not the name, so extensionless dashboards render and get the chart icon. - Renderer: authored left label columns are kept intact, horizontal bar frames grow with row count, and hovered rows get a label-and-bar highlight. - Simplify the create-dashboard skill around one validated example. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(dashboards): address review findings on chart layout and precision - Size horizontal bar `.chart` previews by category count like dashboard panels. - Keep the ECharts label column for percentage bar widths, resolve percentage grid insets for the row highlight, and keep the time axis on the queried range. - Show small readout values with significant digits instead of rounding to 0. - Pass the dashboard's timezone-adjusted today to the range calendar. - Decide the Chat panel's Markdown mode from the file record. - Replace mock-call assertions in the EChartsView tests with DOM behavior. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(charts): size grouped bar rows and keep the row highlight through resizes Unstacked bar series sit side by side in a category row, so grouped charts keep the ECharts label column and their rows fit every bar slot. The row highlight redraws the active row after each render, so a resize moves it with the plot. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(charts): keep authored tooltip arrays and drop the v2 chat mode override Chart tooltip defaults now apply to every entry of an authored `tooltip` array instead of replacing it with one object. Restore the v2 chat payload to staging: the hard-coded `mode: 'agent'` came from the removed dashboard resource work and would overwrite a resumed chat's saved mode. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(charts): reserve the authored bar gap between grouped bars Grouped horizontal bar rows now size their gap from the series `barGap` the way ECharts does (pixels, a percentage of bar width, default 20%, overlap for negative gaps) instead of a fixed 4px. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(charts): format index-encoded pie tooltips and test dashboard refresh through the UI Pie tooltips encoded by dimension index now use the custom formatter, resolving the measure through the encode indexes ECharts passes it. The dashboard preview test drops the mocked controls and callback-driven cases and clicks the real Refresh button to check which analytics queries are invalidated; the range and timezone math stays covered by the time tests. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(charts): format bar tooltip values with each series' own value axis Multi-axis bar charts now read the tooltip unit from the axis each series is plotted on, so a secondary axis no longer inherits the first axis's format. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * test(dashboards): drop mock-call and rendered-text assertions Removes assertions on mocked collaborators and rendered text from the dashboard, chart and analytics tests per the repository testing rules, keeping the observable checks (status codes, thrown errors, DOM roles, computed summaries, emitted zoom ranges). Deletes the animated-number test, which had no behavior left to assert. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * test(dashboards): drop redundant mock resets and default environment pragmas The shared Vitest config already resets mocks and stubbed globals before each test and defaults to the node environment. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * refactor(dashboards): move authoring guidance to Mothership like Sim Pages Drop the create-dashboard built-in skill and its rollout gating in the skill lists. Mothership now learns the dashboard format from a sim-dashboards reference in its own research-and-deliverables skill, the pattern Sim Pages use; Sim workspace built-ins stay Agent-block documents. The dashboards flag still gates the viewer and table analytics. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * refactor(dashboards): give DashboardFeatureGate a props interface Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * feat(dashboards): report dashboard parse errors on file writes The v2 file create, replace and edit responses now carry `diagnostics` for a dashboard file: its parse errors, or an empty list. Writes are never blocked, like the page lint; table columns and queries are still checked when panels render. Parse errors are reported as `path: message` lines, and a block with no recognized kind names the allowed kinds and unknown keys instead of Zod's union dump. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(dashboards): split diagnostics, switch panel query modes, return create revision - One diagnostics entry per parse error instead of a newline-joined string. - A panel choosing aggregate or columns drops the other mode's inherited dashboard fields, so shared defaults serve table and aggregate panels. - The v2 create response returns the revision it produced, like the replace and edit responses. - The analytics route test uses the shared createMockRequest helper. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(dashboards): drop inherited ordering on query-mode switches and cap diagnostics size A panel switching away from the dashboard's query mode no longer inherits its sort or limit, which name aggregate aliases or top-N groups in one mode and rows in the other. Write diagnostics check the 128 KB source limit before decoding the payload. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * revert(dashboards): restore dashboards as a separate workspace resource Reverts "store dashboards as workspace files" and "move authoring guidance to Mothership like Sim Pages", restoring the dashboard resource: its page, folders, APIs, resource tags, Mothership dashboards/dashboard_folders tools, per-turn availability, the create-dashboard built-in skill, and file discovery. Removes the write-time diagnostics that only applied to plain files. Keeps the chart renderer fixes, readable parse errors, panel query-mode handling, the create-response revision, and the test cleanups. The authoring skill keeps its simplified form with the default-formatting note. Migrations are regenerated on top of staging in the next commit. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(dashboards): renumber resource migrations after staging and use central test mocks Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(dashboards): list by folder, canonical folder paths, named props The dashboard browser now asks for one folder's dashboards instead of filtering the first 500 across the workspace, folder paths use the shared segment encoder, and the README names the renumbered migrations. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(dashboards): address review round on the restored resource Fork payloads accept file discovery, dashboard queries are keyed and invalidated per workspace, dashboard contexts open tabs and resolve in organization chats, archived dashboard folders keep their paths, and the folder dialog skips no-op moves and hides the edited subtree. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(dashboards): cap dashboard folders and reject folder name collisions Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * fix(dashboards): name the dashboards page props Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * feat(mothership): send entitlements instead of dashboardsEnabled Restores the entitlement evaluator registry removed in v1.0.0 and sends the granted list with every Mothership turn; dashboards is the first entitlement. Sim still enforces every gated operation. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * refactor(dashboards): give agent dashboard commands the files folder syntax Dashboards take folder paths like files: create/list --folder, move --to (/ is the root), rename and set-content as separate actions, and folder create/move/delete by positional path with --recursive for non-empty deletes. The UI routes keep folder ids. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * feat(dashboards): one Sim-built dashboard per workspace A workspace now has a single dashboard that Sim builds. The Dashboards page renders it directly, or an empty state until the first save. Folders, the list, create, rename and move are removed; the agent reads and saves it with dashboards get / set, and replacing existing content needs its revision. A partial unique index enforces one live dashboard per workspace, and generic file writes can no longer produce the dashboard content type. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * feat(dashboards): put Dashboard under New chat in the sidebar Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * feat(dashboards): headerless dashboard page with a grid-matched icon The dashboard page renders only the dashboard (or its empty state), so the header, its Delete action, and the delete route, hook, and use case go away. A new EMCN Dashboard icon is drawn on the shared sidebar icon grid and replaces the analytics ChartColumn on dashboard surfaces. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01S6aTRnkiu7PxYZPNPXYEMV * refactor(dashboards): drop file discovery for a content-type exclusion With one dashboard per workspace, the listed/unlisted discovery column, its migration, legacy-writer trigger, and the upload backfill script are no longer needed. File listings, search, and workflow file pickers exclude the dashboard content type instead, and the one-dashboard index is now migration 0392. * refactor(dashboards): store dashboards in their own table A dashboard is now a row in a dashboard table with its own id, keyed to its workspace by a unique index (dropping it later allows several). Reads and saves go through a small repository with a revision check, so the file-side guards, exclusions, fixed file name, and dashboard file previews are gone. Chat contexts reference dashboards by dashboardId, and saves are audited as dashboard events. * fix(dashboards): accept dashboard chat contexts and bound revisions The chat request schema now admits dashboard contexts, revisions outside the integer column range are rejected as validation errors, resource tabs only fetch the dashboard when one is open, and the authoring guide notes that range-pinned panels do not drag-zoom. * fix(dashboards): named heading sizes and no redundant dashboard name lookups The dashboard name is fixed, so resource tabs and chat chips no longer fetch the dashboard to title it; dashboard headings use named text sizes. * refactor(dashboards): keep the authoring reference in Mothership's skill Mothership's create-dashboard skill now carries the dashboard syntax like every other worker skill, so the builtin-create-dashboard workspace skill, its source, and its flag gating are removed. * fix(dashboards): server loading fallback and muted gate text * chore(db): drop leftover script-migration test fixture change --------- Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
* fix(desktop): return source connections to the app * fix(desktop): preserve enrollment mode and completion state
* feat(search): add Zoom and Google Meet providers * fix(search): bind meeting pagination and preserve read scopes
* fix(sandbox): preserve workbench output provenance * fix(sandbox): preserve boundary compatibility
* fix(async-jobs): retain accepted run IDs for immediate lookup * fix(async-jobs): finish cancellation scans after root errors
… before its first charge (#8480) * fix(billing): keep billing a request whose period start moved forward before its first charge A Stripe anchor reset between admission and a request's first cost callback stamps row 0 with the reset period. The ledger binding only accepted a later period starting at or after the admitted period's end, so every later callback was refused with a billing-context mismatch and its spend went unbilled. The binding now accepts the same forward-only rule the roll uses: a start later than the admitted one. An earlier period is still refused. Also bound the upgrade-card subscription read in update-cost under the same 1 s standing deadline as the verdict read. * fix(billing): keep an exceeded verdict when the upgrade-card read is slow The shared deadline discarded an exceeded verdict when the card lookup ran past the budget. The verdict read keeps the deadline; the card lookup now falls back to the plan-upgrade card past the same deadline.
… every tick (#8479) * fix(mothership): close abandoned tool meters once instead of alarming every tick A tool meter row (cost unknown) stays open when the process that owned the tool ends mid-execution, and nothing ever closed it. The replay tick counted those rows and logged "Service usage requires reconciliation" at ERROR on every tick in every process, forever. Its 5-minute threshold also flagged tools that were still legitimately running. The replay tick now closes meters older than twice the longest tool watchdog, keeping a pricing failure's error or recording that the tool never finished, and logs each closed meter once with its stream, tool call and reason. Known spend is unaffected: it is saved and delivered as separate receipts. * fix(mothership): keep a closed tool meter final against a late tool completion
…8484) * fix(mothership): answer a retried task wake whose turn already ran The worker retries a task wake under the same run ID until its own run appears. When sim ended that turn without reaching the worker (a usage-limit refusal), every retry reopened the turn, hit the unique stream-id constraint, and failed behind a generic message, so the worker retried forever. Under the chat lock, a wake whose run ID already has a sim run now releases the lock and answers not-found, which the worker treats as a refusal and dismisses the notification. An in-flight turn still holds the lock and answers busy. The headless run-record catch-all now logs the underlying insert error. * fix(mothership): release the wake's chat lock when the run lookup fails The retried-wake check reads copilot_runs after taking the chat lock. If that read threw, the lock stayed held until its TTL because the wake turn that releases it never started. Release with the exact lease on any throw after the acquire. * fix(mothership): release the wake's own chat lease when the run lookup fails * test(mothership): stub the chat lease getter in the wake unit test mock
…#8478) * fix(mothership): explain a Chat turn the worker ends without a reason When the worker rebuilds an ended run from its log (a resume or reattach that reaches a run that already ended, for example at its deadline), it sends an error terminal with no error event. Sim then fell back to the generic "An unexpected error occurred while processing the response." The turn now says the run had already ended and can be continued by sending a message. A reason the worker reports, a replay refusal, and a Stop all still take precedence. Also: - Log the Go stream's error text under errorMessage/detail so it no longer overwrites the log line's own message (stream.ts, buffer.ts). - Rename STREAM_TIMEOUT_MS to CHAT_RUN_DEADLINE_MS and document it as the worker's default run deadline, now only the base of USAGE_SETTLE_MS. - Update the byte-budget doc: a reader behind the ring trim is re-synced from the worker log; replay_gap is only the fallback. * fix(mothership): keep the ended-run message surface-neutral and below a replay refusal The fallback is shared by Chat, workflow execute and inbox, so it no longer tells the reader to send a message. A replay refusal now wins by guard rather than by spread order, with a test that covers the combination. * refactor(mothership): drop the unreachable refusal guard and reuse the stream-abort test helper
…led execution (#8481) * fix(mothership): report a browser-claimed workflow tool from its settled execution When a browser claims a Chat workflow tool, the execute route runs the workflow and keeps running it after the browser detaches, but only the browser's confirmation completed the tool call. A tab that closed, lost its network, or dropped its pagehide beacon left the Chat turn waiting for the full client wait while the worker swept the call. The execute route now records the bound execution's structural completion itself once it settles, guarded on the call still running under that execution's claim, so a browser report or background detach that lands first is kept and nothing is delivered twice. * fix(mothership): report queued async runs and pre-log failures of a browser-claimed workflow tool * refactor(mothership): settle a browser-claimed workflow tool from its log status alone * test(mothership): prove a losing settlement report publishes no confirmation
|
* feat(analytics): add Freebuff Ads conversion tracking * fix(analytics): drop the Freebuff click id when marketing consent is withdrawn
|
* fix(integrations): complete account connections in place * fix(integrations): preserve active account authorizations
… executing (#8482) * fix(mothership): never sweep a Chat run while one of its Sim tools is executing The orphaned-run sweep settled a leased run once it had gone an hour without a status write and no controller held its chat lock. A Sim tool call writes nothing to its run while it executes; only its execution lease heartbeat shows it is alive. So a long tool call (a workflow run can take well over an hour) could have its run settled as interrupted while the worker still held the run, closing tool admission under it. The default one-hour run deadline masked this. The sweep now skips a leased run with an unsettled, unrevoked tool execution whose lease has not expired, both when it selects candidates and in the guarded update. It also locks the run rows before that update, so a tool admission (which locks its run row) either commits a lease the update then sees, or finds the run settled and is refused. * fix(mothership): lock a run's unsettled tool executions before the sweep settles it A lease heartbeat writes only the tool row, so one that passed its expiry check before the lease ran out could commit after the sweep read the old lease and settled the run. The sweep now locks those tool rows after the run rows, so the guarded update sees a committed renewal and a later heartbeat finds the lease expired. Lock-holder tests release on failure.
There was a problem hiding this comment.
Review completed against the latest diff
Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
Tip: instead of fixing issues one by one fix them all with cubic
Re-trigger cubic
|
Search-provider review update: the four inline findings belonging to the provider/rollout changes have been individually assessed. Three corrections are in #8496; the organization-less workspace finding was resolved because managed connected accounts already require a canonical organization. The three valid release threads remain open until that fix reaches staging. The companion warning also applies to these providers: mothership#556 (Lucid) and mothership#558 (Zoom/Google Meet) are included in the open worker release https://github.com/simstudioai/mothership/pull/562 but have not reached worker main. The Sim and worker releases need to stay coordinated. |
* fix(integrations): close connection recovery edge cases * fix(integrations): verify persisted connection outcomes
…adout, isolate repository tests (#8499) * fix(dashboards): carry dashboard chat context ids, label the chart readout, isolate repository tests - User message contexts keep a dashboard mention's dashboardId, both in the optimistic message and when a persisted message is reopened, matching the context the server stores. - The time-series readout row has role="group", so its aria-label is exposed to assistive technology. - The revision test in the dashboard repository suite seeds its own workspace instead of depending on the previous test's row. * fix(dashboards): name the chart readout group only when it has values
…d model the empty new-turn 402 (#8502) * fix(billing): build the mid-run usage card from the admitted payer and model the empty new-turn 402 A direct-v1 run's mid-run verdict reads the payer saved in its account decision, but the upgrade card was resolved from the actor's current subscription, so a payer/actor mismatch picked the wrong action and copy. The exceeded account verdict now carries the payer and subscription it already read, and update-cost and the validate continuation pass it to resolveUsageUpgradePayload instead of a second lookup. The validate contract declared every 402 as a JSON refusal, while a new turn's 402 has no body; the 402 schema now allows the empty body. The wire is unchanged. * fix(billing): drop the unreachable card-read deadline from the usage upgrade card Every exceeded verdict that reaches update-cost now carries its payer, so the deadline-bounded actor subscription lookup could no longer run. Remove the parameter, its call-site argument, the stale TSDoc, and the test that passed without exercising it. * test(copilot): pin the new-turn usage refusal to an empty 402
…loor the replay TTL (#8501) * fix(mothership): number recovered turns past an unreadable ring and floor the replay TTL A recovered run whose replay ring read back empty while its seq counter survived (every retained entry unreadable) resumed numbering at 0, writing over the old range and moving the counter backwards past readers' cursors. Resume from the counter whenever no event was recovered. COPILOT_STREAM_TTL_SECONDS below the 20 s chat-lock heartbeat let an idle live buffer expire between refreshes. Floor it at 60 s. * test(mothership): leave a margin for Redis TTL rounding in the live-buffer floor check
* fix(search): recheck Zoom approval and bound MCP serialization * fix(search): preserve byte-only MCP response limits * fix(search): reject inherited JSON serializers
…act bounds (#8495) * fix(mothership): order dashboards between chats and tables in resource menus * chore(rules): drop the resource-menu sidebar-mirroring rule * fix(dashboards): scope entitlements, keep mention ids, and tighten contracts --------- Co-authored-by: Waleed Latif <[email protected]>
…8503) * fix(sandbox): redact temporary session credentials in model output * fix(sandbox): unify output handling and avoid response copies * fix(sandbox): scan session output without recursion
…ready invoiced (#8505) * fix(billing): refuse charges into a period the terminal settlement already invoiced A subscription's deletion settles its terminal period at once (claim, overage, final invoice, bookkeeping), but recordCumulativeUsage kept topping up that period's row for a still-running request because the subscription's period never rolls after deletion. That spend was never billed. The terminal claim now advances the close marker to the period's end under its FOR UPDATE lock, and recordCumulativeUsage reads the marker in its existing FOR SHARE read: a charge whose target period ends at or before the marker throws CumulativeUsagePeriodClosedError, which update-cost answers with the existing non-retryable BILLING_PERIOD_ELAPSED outcome, so the worker quarantines the leg for reconciliation instead of the spend disappearing into an invoiced period. No migration: the marker is an existing column, and every reader treats a marker at or past periodStart as current, so v0.9.6 behaves unchanged. * test(billing): cover a terminal claim overlapping an in-flight charge Both lock orders against real PostgreSQL: a claim waits for an in-flight charge so the final sum includes it, and a charge that waited on an in-flight claim is refused.
Uh oh!
There was an error while loading. Please reload this page.