Repository navigation
Support free-threaded Python builds - #92
Merged
Merged
Conversation
On Python 3.14's free-threaded build one process running many threads reaches the throughput of many processes running one thread each, while using a single Redis connection pool and one copy of the application state. - Give each worker thread its own idle backoff and rotation offset. The previous instance-wide state was read and written by every thread in the process, so threads reset and doubled each other's delay on top of racing on the counter. - Warn at startup when a free-threaded interpreter runs with the GIL enabled. A C extension that has not declared free-threading support, hiredis being the common one, re-enables the GIL process-wide without any other signal, so the parallelism is silently lost. - Add a CPU-bound scaling benchmark comparing process and thread parallelism. It shows threads give exactly 1.00x on a GIL build and 2.83x on a free-threaded build. - Run the test suite on 3.14t in CI and declare the free-threading classifiers.
The scaling benchmark shows that one process with four threads reaches the same throughput on a free-threaded build as four processes with one thread each, so the existing defaults already reach full parallelism on 3.14t. Auto-tuning them would buy a smaller memory and connection footprint rather than throughput, at the cost of confining a crash or a task recycling to a single process. Record that reasoning next to the numbers that produced it, and tell users which way to move and what they give up when they do.
Resolves four conflicts: - threadmill/backends/redis.py: main added _parse_lease_started_at and a batch acquire that returns lease tokens. Kept both beside the per-thread IdleBackoff, and routed the rotation argument through that per-thread state, because main's version still read the instance attributes this branch removed. - tests/backends/test_redis.py: kept main's batch acquire tests and its sent_args[-2]/[-1] argument positions, retargeted to _idle_backoff, and dropped the duplicated import of the backend module. - tests/test_executor.py: dropped the create_task_error and retry_delay tests, which main moved onto ThreadmillTaskBackend and relocated to tests/backends/test_base.py. - README.md: kept the single-process example, which the surrounding prose introduces. Also carries three fixes found while validating the merge: - The scaling benchmark verifies every task succeeded, so a fast drain cannot be a lost task. That verification caught exactly that. - The scaling benchmark drains its own queue. A worker process left behind by another worktree sharing the same Redis instance stole tasks and reported a drain that never did the work. - test_run__executes_model_task_in_spawned_worker waited a fixed three seconds for a spawned worker, which lost the race under coverage and failed two runs in three. It now polls for the result instead. Measured after the merge, worker start cost included, on 16 CPU-bound tasks of about a second each: build 1p-1t 1p-4t 4p-1t 4p-4t 3.14 (GIL) 16.59s 16.56s 6.05s 17.08s 3.14t 16.58s 6.04s 6.05s 6.05s Threads give nothing on a GIL build, and cost once a pool holds more worker threads in total than the queue holds tasks. A process requests a prefetch batch sized by its thread count, so one process can take the whole queue while its neighbours idle.
The queue comparison measures queue overhead with a trivial echo task, and the thread scaling benchmark measures threadmill alone, so neither answers whether threadmill's free-threading parallelism is unusual. Runs the same CPU-bound workload through every framework that can spread work across threads, one thread and four on both interpreters, 16 tasks of about a second each. Measured on this machine: build framework 1 thread 4 threads ratio 3.14 (GIL) threadmill 17.08s 17.07s 1.00x 3.14 (GIL) celery 15.85s 15.38s 1.03x 3.14 (GIL) dramatiq 15.25s 15.24s 1.00x 3.14t threadmill 17.15s 6.07s 2.83x 3.14t celery 15.85s 4.87s 3.26x 3.14t dramatiq 14.99s 4.62s 3.25x No framework gains anything from threads on a GIL build, which is the control that says the measurement is sound. All of them gain real parallelism on a free-threaded build. threadmill's lower ratio is its start cost, not weaker parallelism. Its process pool costs about two seconds to start, against a fraction of a second for the other two, and that cost sits in both drains. Subtracting it, four threads run the work 3.7x faster, against 3.5x for dramatiq and 3.8x for celery. Its four-thread and four-process drains are also equal, 6.07s against 6.08s, so threads substitute for processes. Three fixes the measurements forced: - The workload drains its own queue on every framework. A worker process left behind by another worktree sharing the same Redis instance consumed the Celery queue and made a drain wait for a task it had already run. - Completion is counted by the tasks themselves rather than by a sentinel consumed last. With several threads a free thread takes the sentinel before its predecessors finish, which would report a drain that never did the work. - One drain per configuration. The tasks are queued once, so a second round finds an empty queue and measures only the start cost, and a median over rounds averages work against no work.
The chart showed one Threadmill bar, so the headline feature was invisible next to the other queues. - Add a Threadmill (free threading) bar, measured with four threads in one process: 59,697 tasks per second against 11,963 for one thread. - Capitalise the label as Threadmill. - List the bar only where the running interpreter really is free-threaded with the GIL off. Elsewhere the drain would only repeat the single-threaded rate, which would read as a feature that found no speedup. - Stack the footnote over two lines and note the four-thread configuration. The bar is a free-threading result rather than four threads overlapping IO. On a GIL build four threads reach 9,970 tasks per second against 7,497 for one, a gain of 1.33x, while on a free-threaded build the same four threads reach 58,306 against 11,881, a gain of 4.91x.
Main added a huey arm to the queue comparison and regenerated both chart SVGs, which this branch had also changed, so four files conflicted. - benchmarks/chart.py: kept the two-line footnote that this branch added, because the footnote with huey no longer fits the canvas on one line, and took main's huey wording for the second line. Dropped a stray closing parenthesis left by the conflict. - docs/images/backend-comparison-*.svg: regenerated from a fresh benchmark run rather than merged by hand. They are generated files, and the merged chart has to rank huey and the free-threading bar together. - README.md: took the alt text from the regenerated chart. - benchmarks/test_backends.py merged on its own, with huey and the free-threading queue both present. Regenerated chart, 60,000 trivial tasks per Threadmill queue and 20,000 for the rest, one worker process each, on a free-threaded interpreter: Threadmill (free threading) 60,259/s Threadmill 12,021/s dramatiq 6,967/s huey 5,389/s celery 2,186/s django-tasks-db 2,056/s django-tasks-rq 75/s Huey lands fourth, behind dramatiq, matching the ranking in main's own run. Validation: all 16 pre-commit hooks pass, the chart regenerates byte for byte from the recorded benchmark JSON, both themes draw seven bars with no overflow, the free-threading bar stays out of a GIL build's chart, and the suite passes on both interpreters (259 each).
- Drop the note above the warning call. The order of the two lines says it. - Inline the interpreter checks. They are single expressions, nothing reused them, and the benchmarks had to import them to report the build. - Warn in two sentences instead of three, without naming a C extension. The cost the GIL inflicts is still in the docstring, where it explains the check. - Drive the warning in the tests through the stdlib attributes it reads, so the tests cover the absent private API again. - Drop the comment above the free-threaded CI job.
IdleBackoff sat next to threadmill.retry.ExponentialBackoff, which retries failed tasks rather than pacing idle polls. Nothing is expected to import it, so name it _IdleBackoff and say in its docstring what sets it apart.
The chart compared four threads on a free-threaded build against one thread everywhere else, and needed a footnote to explain the difference away. - Run four threads wherever the worker supports threads: celery on its thread pool, dramatiq and huey by flag, threadmill by threads. Every prefetch window stays at READ_AHEAD messages. - Keep one thread for django-tasks-db and django-tasks-rq, which ship single-threaded workers and cannot be told otherwise. - Take the free-threading row from a 3.14t run and leave the rest of the chart on the GIL build, so both Threadmill bars are the same pool on two interpreters. chart.py accepts the second run as another argument. - Drop the footnote, along with the theme colour only it used. - Deepen the threadmill queues to 120,000 tasks. At 60,000 the marginal drain of the free-threading row is one to two seconds, so the one-second quantization of the worker start and stop decided the reported rate. - Re-measure and update the README alt text.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Free-threaded Python lets worker threads run truly in parallel, so one process reaches the throughput of many processes with a single Redis connection pool and one copy of your application state. Threadmill could not deliver that, and could not tell whether a build delivers it at all: a C extension that does not declare free-threading support silently re-enables the GIL for the whole process.
--threadsto the worker pool. Every thread keeps its own idle polling state, and the worker warns when a free-threaded interpreter runs with the GIL re-enabled.benchmarks/test_scaling.py, which measures threads against processes on both interpreters, andbenchmarks/test_parallelism.py, which measures the same workload against celery and dramatiq. On a free-threaded build threads reach the throughput of processes. On a GIL build they gain nothing.