Skip to content

gh-158897: Optimize io.BufferedReader.readline() by calling memchr() - #158944

Open
vstinner wants to merge 4 commits into
python:mainfrom
vstinner:readline_memchr
Open

vstinner wants to merge 4 commits into
python:mainfrom
vstinner:readline_memchr

Conversation

@vstinner

@vstinner vstinner commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Optimize io.BufferedReader.readline() when lines are close to the buffer size (128 kB by default) or longer than the buffer size. Replace the C loop searching for the newline byte in the buffer with a memchr() call which is more efficient.

…chr()

Optimize io.BufferedReader.readline() when lines are close to the
buffer size (128 kB by default) or longer than the buffer size.
Replace the C loop searching for the newline byte in the buffer with
a memchr() call which is more efficient.
Comment thread Modules/_io/bufferedio.c Outdated
@vstinner

vstinner commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

I ran #158897 (comment) benchmark on Linux (Fedora 44) with CPU isolation.

$ git switch main
$ make
$ ./python bench.py -o main.json -p70
$ git switch readline_memchr 
$ make
$ ./python bench.py -o memchr.json -p70 -v 

Results: Mean +- std dev: [main] 1.78 ms +- 0.14 ms -> [memchr] 417 us +- 34 us: 4.26x faster

This change makes this specific benchmark 4.26x faster!

But well, in practice, text files with lines around 128 kB or longer than 128 kB should be rare.

I had some issues to get reliable results, so I ran the benchmark with -p70 to run more processes than the default (20).

@vstinner

vstinner commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

@leandrodamascena: Would you be able to check if this change solves the performance regression that you spotted in #158897 (comment)? (Can you run your benchmarks on a patched Python?) Tell me if you need help.

@vstinner

vstinner commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

The benchmark #158897 (comment) is a microbenchmark which seems to depend heavily on low-level things like CPU cache misses or CPU branch prediction.

After I fixed my off-by-one error, I reran the benchmark to be sure that the result remains the same, and the measure on the main branch jumped from 1.78 ms +- 0.14 ms to 1.87 ms +- 0.18 ms (same C code, just recompiled).

Anyway, the updated result: Mean +- std dev: [main] 1.87 ms +- 0.18 ms -> [memchr] 330 us +- 30 us: 5.66x faster. It's now 5.7x faster...

I'm not sure how to get more reliable results. I suppose that building Python with PGO would help to get more efficient machine code most of the time.

@leandrodamascena

Copy link
Copy Markdown

Thanks @vstinner! I built the latest version of this PR (with the off-by-one fix) into my CPython OCI images and tested it on AWS Lambda x86_64 with PGO+LTO, comparing it against its base with 25 paired samples per scenario:

  • Mixed lines: 3.981 ms → 0.946 ms (-76.4%, CI -77.9% to -74.5%)
  • Lines close to the buffer size: 247.7 µs → 197.2 µs (-20.5%, CI -21.5% to -19.7%)
  • Long lines: 3.648 ms → 0.784 ms (-78.6%, CI -79.2% to -78.5%)

So yes, this fixes the regression in my Lambda workload. Since main was about 43% slower than before the regression in the mixed case, this should leave it well ahead of where it was before.

Catching regressions early, before they reach a release, is exactly what I hope my workload can do, so getting a fix this quickly is really satisfying. Thanks a lot for looking into this!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting core review extension-modules C modules in the Modules dir performance Performance or resource usage

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants