Repository navigation
fix(client): reset connection_failures when a store that gave up is started again - #139
Merged
Merged
Conversation
…tarted again `_rearm_waiters` reset the internal `_failures` count but not `diagnostics.connection_failures`, so a restarted store reported the old run's failures until its new run failed or succeeded: after two recoverable failures and a 401, `start()` left `failed` None with `connection_failures == 2` and nothing failed in the new run. That counter is the outage signal now that recoverable failures retry indefinitely, so a healthy restart read as a store already riding out an outage. The TypeScript SDK resets it in `start()`. `test_a_restart_resets_the_failure_count` could not see this: it reads the counter only after the new run's first failure, which overwrites the stale value. The new test holds the new run's first request open and reads the counter right after `start()`; it fails with `2 == 0` when the reset is reverted. Spec: launchdarkly/ai-sdks-monorepo#40 (TESTING.md §3.25). Co-Authored-By: Claude Opus 5.5 <[email protected]>
jeffdupont
approved these changes
Oct 5, 2026
The docstring listed a give-up as ending delivery with no hint that start() undoes it. Matches TESTING.md §3.25 (ai-sdks-monorepo#40). Co-Authored-By: Claude Opus 5.5 <[email protected]>
1 task done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
After a store gives up and you call
start()again,diagnostics.connection_failureskept the old run's count. So a healthy restarted store looked like it was already in an outage. This is a one-line fix, and it matches the TypeScript SDK. Found in review of https://github.com/launchdarkly/ai-sdks-monorepo/pull/40#pullrequestreview-5419525312.Changes
_rearm_waitersresetsdiagnostics.connection_failures, alongside the internal_failurescount it already reset. Before this, two recoverable failures and then a 401, followed bystart(), gavefailed is Noneandconnection_failures == 2, with nothing failed in the new run. TypeScript resets the counter instart()(js-ai-sdk#107).test_a_restart_reports_no_failures_before_the_new_run_has_any. It holds the new run's first request open and reads the counter right afterstart(). The existingtest_a_restart_resets_the_failure_countcouldn't catch this, because it reads the counter only after the new run fails, and that failure overwrites the stale value. The new test fails with2 == 0when the reset is reverted.agents.mdnow say the counter resets on restart.wait_for_skillsdocstring now says a store that gave up waits again oncestart()runs, and onlycloseis final. Same change for TypeScript: docs(client): say a restarted store's waitForSkills waits again js-ai-sdk#113.Spec: TESTING.md §3.25 in https://github.com/launchdarkly/ai-sdks-monorepo/pull/40.
Test plan
make lint,make format-check,make typecheckmake test: 2201 passed, 11 skippedTestWatchSkillsOverTheTransport::test_a_revocation_prunes_without_a_restartfailed once in my runs. It also fails onxie/agent-skillswithout this change (1 in 25 runs), so it's an existing flaky test.🤖 Generated with Claude Code