Skip to content

Support free-threaded Python builds - #92

Merged
codingjoe merged 9 commits into
mainfrom
codingjoe-free-threading-support
Oct 8, 2026
Merged

codingjoe merged 9 commits into
mainfrom
codingjoe-free-threading-support

Conversation

@codingjoe

Copy link
Copy Markdown
Owner

Free-threaded Python lets worker threads run truly in parallel, so one process reaches the throughput of many processes with a single Redis connection pool and one copy of your application state. Threadmill could not deliver that, and could not tell whether a build delivers it at all: a C extension that does not declare free-threading support silently re-enables the GIL for the whole process.

  • Add --threads to the worker pool. Every thread keeps its own idle polling state, and the worker warns when a free-threaded interpreter runs with the GIL re-enabled.
  • Keep one process per core with a single thread as the defaults, and document when threads beat processes and when they do not.
  • Add benchmarks/test_scaling.py, which measures threads against processes on both interpreters, and benchmarks/test_parallelism.py, which measures the same workload against celery and dramatiq. On a free-threaded build threads reach the throughput of processes. On a GIL build they gain nothing.
  • Chart the free-threading result next to the other queues. Every queue now runs four threads where its worker supports them, and the free-threading bar is measured on 3.14t, so the chart compares interpreters rather than pool sizes.
  • Run the suite on 3.14t in CI.

On Python 3.14's free-threaded build one process running many threads
reaches the throughput of many processes running one thread each, while
using a single Redis connection pool and one copy of the application
state.

- Give each worker thread its own idle backoff and rotation offset. The
  previous instance-wide state was read and written by every thread in
  the process, so threads reset and doubled each other's delay on top of
  racing on the counter.
- Warn at startup when a free-threaded interpreter runs with the GIL
  enabled. A C extension that has not declared free-threading support,
  hiredis being the common one, re-enables the GIL process-wide without
  any other signal, so the parallelism is silently lost.
- Add a CPU-bound scaling benchmark comparing process and thread
  parallelism. It shows threads give exactly 1.00x on a GIL build and
  2.83x on a free-threaded build.
- Run the test suite on 3.14t in CI and declare the free-threading
  classifiers.
The scaling benchmark shows that one process with four threads reaches
the same throughput on a free-threaded build as four processes with one
thread each, so the existing defaults already reach full parallelism on
3.14t. Auto-tuning them would buy a smaller memory and connection
footprint rather than throughput, at the cost of confining a crash or a
task recycling to a single process.

Record that reasoning next to the numbers that produced it, and tell
users which way to move and what they give up when they do.
Resolves four conflicts:

- threadmill/backends/redis.py: main added _parse_lease_started_at and a
  batch acquire that returns lease tokens. Kept both beside the per-thread
  IdleBackoff, and routed the rotation argument through that per-thread
  state, because main's version still read the instance attributes this
  branch removed.
- tests/backends/test_redis.py: kept main's batch acquire tests and its
  sent_args[-2]/[-1] argument positions, retargeted to _idle_backoff, and
  dropped the duplicated import of the backend module.
- tests/test_executor.py: dropped the create_task_error and retry_delay
  tests, which main moved onto ThreadmillTaskBackend and relocated to
  tests/backends/test_base.py.
- README.md: kept the single-process example, which the surrounding prose
  introduces.

Also carries three fixes found while validating the merge:

- The scaling benchmark verifies every task succeeded, so a fast drain
  cannot be a lost task. That verification caught exactly that.
- The scaling benchmark drains its own queue. A worker process left behind
  by another worktree sharing the same Redis instance stole tasks and
  reported a drain that never did the work.
- test_run__executes_model_task_in_spawned_worker waited a fixed three
  seconds for a spawned worker, which lost the race under coverage and
  failed two runs in three. It now polls for the result instead.

Measured after the merge, worker start cost included, on 16 CPU-bound tasks
of about a second each:

  build        1p-1t   1p-4t   4p-1t   4p-4t
  3.14 (GIL)   16.59s  16.56s   6.05s  17.08s
  3.14t        16.58s   6.04s   6.05s   6.05s

Threads give nothing on a GIL build, and cost once a pool holds more worker
threads in total than the queue holds tasks. A process requests a prefetch
batch sized by its thread count, so one process can take the whole queue
while its neighbours idle.
The queue comparison measures queue overhead with a trivial echo task, and
the thread scaling benchmark measures threadmill alone, so neither answers
whether threadmill's free-threading parallelism is unusual.

Runs the same CPU-bound workload through every framework that can spread
work across threads, one thread and four on both interpreters, 16 tasks of
about a second each. Measured on this machine:

  build        framework   1 thread  4 threads  ratio
  3.14  (GIL)  threadmill    17.08s     17.07s  1.00x
  3.14  (GIL)  celery        15.85s     15.38s  1.03x
  3.14  (GIL)  dramatiq      15.25s     15.24s  1.00x
  3.14t        threadmill    17.15s      6.07s  2.83x
  3.14t        celery        15.85s      4.87s  3.26x
  3.14t        dramatiq      14.99s      4.62s  3.25x

No framework gains anything from threads on a GIL build, which is the
control that says the measurement is sound. All of them gain real
parallelism on a free-threaded build.

threadmill's lower ratio is its start cost, not weaker parallelism. Its
process pool costs about two seconds to start, against a fraction of a
second for the other two, and that cost sits in both drains. Subtracting
it, four threads run the work 3.7x faster, against 3.5x for dramatiq and
3.8x for celery. Its four-thread and four-process drains are also equal,
6.07s against 6.08s, so threads substitute for processes.

Three fixes the measurements forced:

- The workload drains its own queue on every framework. A worker process
  left behind by another worktree sharing the same Redis instance consumed
  the Celery queue and made a drain wait for a task it had already run.
- Completion is counted by the tasks themselves rather than by a sentinel
  consumed last. With several threads a free thread takes the sentinel
  before its predecessors finish, which would report a drain that never
  did the work.
- One drain per configuration. The tasks are queued once, so a second
  round finds an empty queue and measures only the start cost, and a
  median over rounds averages work against no work.
The chart showed one Threadmill bar, so the headline feature was invisible
next to the other queues.

- Add a Threadmill (free threading) bar, measured with four threads in one
  process: 59,697 tasks per second against 11,963 for one thread.
- Capitalise the label as Threadmill.
- List the bar only where the running interpreter really is free-threaded
  with the GIL off. Elsewhere the drain would only repeat the
  single-threaded rate, which would read as a feature that found no
  speedup.
- Stack the footnote over two lines and note the four-thread configuration.

The bar is a free-threading result rather than four threads overlapping IO.
On a GIL build four threads reach 9,970 tasks per second against 7,497 for
one, a gain of 1.33x, while on a free-threaded build the same four threads
reach 58,306 against 11,881, a gain of 4.91x.
Main added a huey arm to the queue comparison and regenerated both chart
SVGs, which this branch had also changed, so four files conflicted.

- benchmarks/chart.py: kept the two-line footnote that this branch added,
  because the footnote with huey no longer fits the canvas on one line, and
  took main's huey wording for the second line. Dropped a stray closing
  parenthesis left by the conflict.
- docs/images/backend-comparison-*.svg: regenerated from a fresh benchmark
  run rather than merged by hand. They are generated files, and the merged
  chart has to rank huey and the free-threading bar together.
- README.md: took the alt text from the regenerated chart.
- benchmarks/test_backends.py merged on its own, with huey and the
  free-threading queue both present.

Regenerated chart, 60,000 trivial tasks per Threadmill queue and 20,000 for
the rest, one worker process each, on a free-threaded interpreter:

  Threadmill (free threading)   60,259/s
  Threadmill                    12,021/s
  dramatiq                       6,967/s
  huey                           5,389/s
  celery                         2,186/s
  django-tasks-db                2,056/s
  django-tasks-rq                   75/s

Huey lands fourth, behind dramatiq, matching the ranking in main's own run.

Validation: all 16 pre-commit hooks pass, the chart regenerates byte for
byte from the recorded benchmark JSON, both themes draw seven bars with no
overflow, the free-threading bar stays out of a GIL build's chart, and the
suite passes on both interpreters (259 each).
- Drop the note above the warning call. The order of the two lines says it.
- Inline the interpreter checks. They are single expressions, nothing reused
  them, and the benchmarks had to import them to report the build.
- Warn in two sentences instead of three, without naming a C extension. The
  cost the GIL inflicts is still in the docstring, where it explains the check.
- Drive the warning in the tests through the stdlib attributes it reads, so
  the tests cover the absent private API again.
- Drop the comment above the free-threaded CI job.
IdleBackoff sat next to threadmill.retry.ExponentialBackoff, which retries
failed tasks rather than pacing idle polls. Nothing is expected to import it,
so name it _IdleBackoff and say in its docstring what sets it apart.
The chart compared four threads on a free-threaded build against one thread
everywhere else, and needed a footnote to explain the difference away.

- Run four threads wherever the worker supports threads: celery on its thread
  pool, dramatiq and huey by flag, threadmill by threads. Every prefetch window
  stays at READ_AHEAD messages.
- Keep one thread for django-tasks-db and django-tasks-rq, which ship
  single-threaded workers and cannot be told otherwise.
- Take the free-threading row from a 3.14t run and leave the rest of the chart
  on the GIL build, so both Threadmill bars are the same pool on two
  interpreters. chart.py accepts the second run as another argument.
- Drop the footnote, along with the theme colour only it used.
- Deepen the threadmill queues to 120,000 tasks. At 60,000 the marginal drain
  of the free-threading row is one to two seconds, so the one-second
  quantization of the worker start and stop decided the reported rate.
- Re-measure and update the README alt text.
@codingjoe
codingjoe merged commit d8270af into main Oct 8, 2026
5 checks passed
@codingjoe
codingjoe deleted the codingjoe-free-threading-support branch October 8, 2026 15:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant