Repository navigation
Conversation
…chr() Optimize io.BufferedReader.readline() when lines are close to the buffer size (128 kB by default) or longer than the buffer size. Replace the C loop searching for the newline byte in the buffer with a memchr() call which is more efficient.
|
I ran #158897 (comment) benchmark on Linux (Fedora 44) with CPU isolation. Results: This change makes this specific benchmark 4.26x faster! But well, in practice, text files with lines around 128 kB or longer than 128 kB should be rare. I had some issues to get reliable results, so I ran the benchmark with |
|
@leandrodamascena: Would you be able to check if this change solves the performance regression that you spotted in #158897 (comment)? (Can you run your benchmarks on a patched Python?) Tell me if you need help. |
|
The benchmark #158897 (comment) is a microbenchmark which seems to depend heavily on low-level things like CPU cache misses or CPU branch prediction. After I fixed my off-by-one error, I reran the benchmark to be sure that the result remains the same, and the measure on the main branch jumped from Anyway, the updated result: I'm not sure how to get more reliable results. I suppose that building Python with PGO would help to get more efficient machine code most of the time. |
|
Thanks @vstinner! I built the latest version of this PR (with the off-by-one fix) into my CPython OCI images and tested it on AWS Lambda x86_64 with PGO+LTO, comparing it against its base with 25 paired samples per scenario:
So yes, this fixes the regression in my Lambda workload. Since main was about 43% slower than before the regression in the mixed case, this should leave it well ahead of where it was before. Catching regressions early, before they reach a release, is exactly what I hope my workload can do, so getting a fix this quickly is really satisfying. Thanks a lot for looking into this! |
Optimize io.BufferedReader.readline() when lines are close to the buffer size (128 kB by default) or longer than the buffer size. Replace the C loop searching for the newline byte in the buffer with a memchr() call which is more efficient.
BufferedReader.readline()after switching toPyBytesWriter#158897