Skip to content

Requeue near the tail in O(offset) instead of scanning the queue with LINSERT - #436

Merged
tekmaven merged 1 commit into
mainfrom
rh/requeue-without-linsert
Oct 2, 2026
Merged

tekmaven merged 1 commit into
mainfrom
rh/requeue-without-linsert

Conversation

@tekmaven

@tekmaven tekmaven commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Why

requeue.lua and reserve.lua put a requeued test back offset + 1 entries from the tail of the queue (the end RPOP reserves from) with LINSERT BEFORE <pivot> (requeue.lua, reserve.lua). LINSERT finds its pivot by scanning from the head, and the pivot sits next to the tail, so every requeue walks the whole queue on Redis's single command thread. Replicas repeat the scan.

On large queues that adds up. In our CI, a burst of requeues on a queue of several hundred thousand tests kept a Redis main thread ~80% busy in LINSERT for about 15 minutes, at ~46 ms per call. Other builds sharing that Redis hit the client's 2-second read timeout. One build's leader timed out in push_batch, and its workers then failed with LostMaster.

What

This replaces the pivot lookup with a tail rotation, which is O(offset) instead of O(queue length):

  1. LRANGE queue -(offset + 1) -1 reads the entries that should stay ahead of the test.
  2. LTRIM queue 0 -(offset + 2) drops them.
  3. RPUSH the requeued entry, then RPUSH them back in order.

Queues of at most offset + 1 entries, and offsets of zero or less, still LPUSH to the head. The list never becomes empty during the rotation, so the queue key keeps its TTL.

KEYS and ARGV are unchanged, so neither the Ruby nor the Python client changes. The helper is duplicated in both scripts because the Python client loads scripts as-is and doesn't resolve -- @include.

Behavior

The new scripts give the same results as the LINSERT version. I checked this by running both against Redis 8.6.3 over 2,450 cases:

  • offsets −1 to 45;
  • queue lengths 0–60, 100 and 1,000;
  • both requeue.lua, and reserve.lua deferring 1, 2 or 4 of its own requeued tests.

Return values, queue contents, requeued-by, running and the queue's TTL matched in every case.

Two edge cases differ, by design:

  • Duplicate pivot value. If the pivot's value also appears closer to the head (a duplicate queue entry), LINSERT inserted before that first copy. That was often the head, which is the back of the line. The test now lands at the requested offset.
  • Offsets below −1. These now push to the head, where LINSERT inserted at index -offset - 1 from the head. Offsets are positive in practice; the default is 42.

Benchmarks

Setup: local Redis 8.6.3 on an Apple-silicon laptop, ~250-byte entries, requeue_offset 4. Times are Redis-reported time per EVALSHA (INFO commandstats).

Queue entries LINSERT (before) Tail rotation (after)
10,000 125 µs 14 µs
100,000 1.2 ms 13 µs
650,000 9.9 ms 14 µs
650,000, pivot among 2,000 earlier requeues, offset 42 41 ms 27 µs
  • Default offset: with the default of 42, the new script takes ~28 µs per call.
  • Effect on other clients: while one client requeued back to back on the 650,000-entry queue, a second client's PING latency went from p50 10.1 ms / p99 13.7 ms to p50 0.13 ms / p99 0.49 ms. Throughput went from 81 to 4,127 requeues/s.

Tests

Four new tests in ruby/test/ci/queue/redis_test.rb, written against the public API (poll, requeue, to_a):

  • A requeued test runs again after the next offset + 1 tests.
  • With only offset + 1 tests left, it runs last and the queue keeps its TTL.
  • With a zero offset, it runs last.
  • With another worker registered, a worker that pops its own requeued test puts it behind the next offset + 1 tests.

All four pass against the old LINSERT scripts. Each one fails against at least one deliberately broken copy of the new code:

  • the zero-offset check removed;
  • the length check off by one;
  • LTRIM keeping one entry too many;
  • LRANGE off by one;
  • the requeued entry dropped;
  • reserve.lua always pushing to the head.

Results locally on Redis 8.6.3:

  • Ruby suite (Ruby 4.0.7): 327 tests, 0 failures.
  • Python suite (3.13): 35 passed.
  • End-to-end Python check: a test requeued with offset 4 on a 50-test queue runs again after the next 5.

Rollout

There's no version bump in this PR. Workers on old and new versions can share a queue: the list layout is unchanged, and both versions put requeued tests in the same place.

… LINSERT

requeue.lua and reserve.lua put a requeued test back `offset` + 1 entries
from the tail with `LINSERT BEFORE <pivot>`. LINSERT finds the pivot by
scanning from the head, so every requeue costs O(queue length) on Redis's
single command thread: about 10 ms per call on a 650k-entry queue, and about
40 ms once earlier requeues have piled up near the tail. A burst of requeues
on a large queue can then block every other build that shares the Redis.

Instead, read the `offset` + 1 tail entries with LRANGE, drop them with
LTRIM, push the requeued entry, and push them back. That gives the same list
in O(offset). Queues of at most `offset` + 1 entries, and offsets of zero or
less, still push to the head.

One intended difference: when the pivot's value also appears nearer the
head (a duplicate entry), LINSERT inserted before that first copy, often
sending the test to the back of the queue. It now lands at the offset.
@tekmaven
tekmaven marked this pull request as ready for review October 2, 2026 02:21
@tekmaven
tekmaven merged commit 209316c into main Oct 2, 2026
34 checks passed

This branch was successfully deployed

1 active deployment
rubygems — 471e2abe Deployed Oct 2, 2026 by shopify-shipit[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants