Skip to content

improvement(search): run Search retirement as a throttled background migration - #8521

Closed
waleedlatif1 wants to merge 1 commit into
stagingfrom
improvement/search-retirement-background
Closed

waleedlatif1 wants to merge 1 commit into
stagingfrom
improvement/search-retirement-background

Conversation

@waleedlatif1

Copy link
Copy Markdown
Collaborator

Summary

  • Takes 0029_retire_all_search_embeddings off the deploy's critical path. The deploy slice advances the retirement for at most one minute, never starts index maintenance, and throws ScriptMigrationDeferred. The release then switches over, and 0029 stays unrecorded until the cleanup is actually finished.

  • Adds .github/workflows/search-retirement.yml as the background runner. It runs an hourly 50-minute retirement slice against staging and production, using the same secrets and migration role as migrations.yml. A daily off-peak slice (03:07 UTC) is the only scheduled run allowed to start REINDEX CONCURRENTLY and VACUUM. The first slice that finishes both journals 0029 and its superseded names. Every slice resumes the existing cursor and maintenance checkpoints, with no reset.

  • Throttles on database health, not just page time. After each page it pauses for the longest of:

    • a duty-cycle pause (≤50% busy);
    • a WAL pause that holds the whole database's WAL rate, measured with pg_current_wal_lsn() and counting every writer, under 10% of max_wal_size per checkpoint_timeout, so checkpoints stay time-triggered;
    • an exponential back-off (15 s doubling to 5 min) that also halves the page whenever its commit takes over 1 s. The commit is timed apart from the statements, so it isolates the synchronous-replication and fsync wait.

    It also pauses before a page when pg_stat_replication lag exceeds 10 s.

  • Session advisory lock: only one runner retires pages at a time. A deploy that finds the background runner active defers at once.

  • Maintenance is budgeted. No rebuild or vacuum starts after the slice's budget, and one already started finishes, since a cancelled concurrent rebuild starts over. Vacuums run with autovacuum's 2 ms vacuum_cost_delay instead of the unthrottled manual default.

  • One structured Search retirement batch log per page:

    • rows, page and commit time, WAL bytes;
    • the pause and the signal that set it (duty / wal / backoff / slow_commit / slow_page / phase_change / replica_lag);
    • the row limit.
  • Rewrites the runbook (search-embedding-retirement.md): execution model, pacing, observing, and pausing (gh workflow disable search-retirement.yml plus cancelling the in-flight run, which rolls back with its cursor).

Why

A bulk delete run inside the deploy pushed WAL far past max_wal_size. Checkpoints ran back to back, and each one re-logs a full page image on the first touch. fsync and synchronous-replication commit waits stretched, and unrelated app writes hit lock and statement timeouts. Page-time adaptation alone did not stop this: each page stayed under its own timeout while the cluster degraded. The app does not depend on this cleanup finishing (documents are already fenced), so it belongs outside the deploy, paced on the database's own pressure.

Research

Chosen pattern: GitLab's (deploy enqueues, background worker executes, health throttling, finalize when done), using the infrastructure Sim already has:

  • ScriptMigrationDeferred plus the script_migrations journal for "unfinished, retry later";
  • the existing cursor and progress tables for resumability;
  • the GitHub Actions migration job's role and secrets as the worker.

The cron routes and Trigger.dev tasks run as the app role. That role cannot rebuild or vacuum these tables and cannot read replication lag, so using them would have meant new credentials.

Type of Change

  • Improvement

Testing

  • Extended 0027_retire_search_embeddings.integration.ts (real Postgres + pgvector). Each new test went red with its guard reverted:
    • deploy slice returns within budget plus one page, defers unrecorded, and defers immediately while another runner holds the lock;
    • background run resumes strictly past the saved cursor, completes, and journals;
    • a slow commit (an injected deferred trigger sleeps at commit) halves the page and backs off exponentially, then recovers;
    • the paced WAL rate stays within the budget;
    • maintenance never starts in the deploy slice, stops after one checkpointed rebuild past its budget, and resumes without redoing it.
  • Updated the legacy-checkpoint cases: the deploy now retires and defers maintenance, and the background run rebuilds and journals.
  • Smoke-ran the CLI entry against a scratch database. A budgeted slice deferred; a --maintenance slice rebuilt, vacuumed and journaled.
  • bun run test:integration (packages/db: 131 passed; apps/sim: 1274 passed), packages/db unit tests, bun run type-check, bun run lint, bun run check:audits, docs-manifest:check.

Checklist

  • Code follows project style guidelines
  • Self-reviewed my changes
  • Tests added/updated and passing (new tests pass the test-audit authoring gate)
  • No new warnings introduced
  • I confirm that I have read and agree to the terms outlined in the Contributor License Agreement (CLA)

…migration

The deploy advances the retirement for one minute and defers; a scheduled
workflow advances it between deploys in resumable slices paced on commit
latency, WAL rate and replication lag, starts index maintenance only in an
off-peak window, and journals 0029 once everything is finished.
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@vercel

vercel Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
docs Ready Ready Preview Oct 1, 2026 4:30pm UTC

Request Review

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 2/5

[Critical risk] Adds background job to manage database search index retirement.

The PR should not merge until the deploy slice’s unbounded preparation and validation, and the WAL-admission gap, are addressed.

Findings

  1. P1 Initial snapshot exceeds deploy budget ▶
  2. P1 Completed retirement repeats lengthy validation ▶
  3. P1 WAL pressure can outlast pause ▶
  4. P2 Phase changes skip commit backoff ▶

Summary

This PR moves Search-embedding retirement into resumable deploy and scheduled background slices, adds database-health pacing, and reserves index maintenance for off-peak runs.

  • The deploy budget does not cover initial target preparation or potentially lengthy completion validation.
  • WAL-rate admission and slow-commit handling have gaps under database pressure.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Deploy[Deploy: 60-second slice] --> Retirement[Retirement cursor and pages]
  Schedule[Hourly background slice] --> Retirement
  Retirement -->|pages remain| Deferred[Defer without journal]
  Retirement -->|complete| Maintenance{Maintenance allowed?}
  Maintenance -->|no| Deferred
  Maintenance -->|off-peak or manual| Rebuild[Reindex and vacuum checkpoints]
  Rebuild -->|complete| Journal[Journal 0029 and superseded names]
Loading

Reviews (1) · Last reviewed commit: "improvement(search): run Search retireme..."

return 'deferred'
}
try {
if (!(await prepareTargets(sql))) return 'complete'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Initial snapshot exceeds deploy budget

If the first deploy has many Search knowledge bases, prepareTargets scans and records all of them before the deadline is checked. The migration can remain on the deploy path well beyond its promised one-minute slice without starting a retirement page. The initial snapshot needs a budget or a resumable checkpoint.

Knowledge Base Used: Database schema and migrations

const pause = (ms: number) => sleep(Math.max(0, Math.min(ms, deadline - performance.now())))

for (;;) {
if (performance.now() >= deadline) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Completed retirement repeats lengthy validation

When retirement is saved as done but maintenance is still pending, each deploy repeats the full target-marker validation. The deadline is checked only before that call, which has a 30-minute statement timeout. A deploy intended to defer after one minute can therefore wait for the recheck on every release until maintenance journals 0029.

Knowledge Base Used: Database schema and migrations

Comment on lines +301 to +307
const dutyPauseMs = pageMs * throttle.dutyRatio
/** Long enough that the WAL written during the page, spread over page and pause, fits the budget. */
const walPauseMs = walBytes / walBytesPerMs - pageMs
const strainPauseMs = state === 'slow_commit' ? backoffMs() : 0
const pauseMs = Math.max(
0,
Math.min(throttle.maxPauseMs, Math.max(dutyPauseMs, walPauseMs, strainPauseMs))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 WAL pressure can outlast pause

If a page and other database writers produce enough WAL to require more than five minutes of pacing, this code caps the pause anyway. It also does not count WAL written during the pause before starting the next page. Retirement can thus add more WAL while the database remains above the configured rate budget.

Comment on lines +279 to +287
if (page.transition) {
/** A phase change may include a full recheck, which says nothing about page cost. */
state = 'phase_change'
logger.info('Search retirement phase changed', {
transition: page.transition,
batches,
mutated,
})
} else if (page.commitMs > throttle.slowCommitMs) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Phase changes skip commit backoff

A phase-change page whose commit takes longer than slowCommitMs enters this branch before the slow-commit check. It neither reduces the next page size nor applies the commit back-off. That leaves a gap in the pressure response precisely when a slow commit is observed; phase changes can be exempted from page-duration sizing without ignoring commit latency.

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

Superseded: the retirement moves out of the deploy entirely and runs as an operator-run, throttled maintenance command. A new PR replaces this one.

@waleedlatif1
waleedlatif1 deleted the improvement/search-retirement-background branch October 1, 2026 16:58

This branch was successfully deployed

1 active deployment
Preview — 8db39d7e Deployed Oct 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant