Skip to content

gh-158585: Optimize bytes object creation - #158664

Open
vstinner wants to merge 6 commits into
python:mainfrom
vstinner:writer_create
Open

vstinner wants to merge 6 commits into
python:mainfrom
vstinner:writer_create

Conversation

@vstinner

@vstinner vstinner commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Add bytes_alloc() helper function. bytes_alloc() and _PyBytes_FromStringAndSize() have less conditional branches than previous code, and so should be a little bit faster.

  • Rename existing bytes_alloc() to bytes_type_alloc().
  • Replace _PyBytes_FromSize(calloc=1) with _PyBytes_FromSizeZero().
  • Add set_ob_shash_unsafe(): similar to set_ob_shash() but don't use an atomic operation on Free Threading. Use this new function on newly allocated bytes objects and in bytes_resize_inplace().
  • Replace PyBytes_FromStringAndSize(NULL, size) with bytes_alloc(size).
  • Use bytes_alloc() in PyBytes_FromString() and _PyBytes_Repeat().
  • Add _PyBytes_FromStringAndSize(): similar to PyBytes_FromStringAndSize(), but str must not be NULL.
  • Replace PyBytes_FromStringAndSize() with _PyBytes_FromStringAndSize().
  • Test that PyBytes_FromString() and PyBytes_FromStringAndSize() return singletons for 0 or 1 bytes.

@vstinner

vstinner commented Oct 3, 2026

Copy link
Copy Markdown
Member Author

Results of PR gh-158665 benchmark: Mean +- std dev: [ref] 34.8 ns +- 0.4 ns -> [change] 32.2 ns +- 0.5 ns: 1.08x faster.

The change makes PyBytesWriter_Finish() 1.08x faster, it saves 2.6 nanoseconds.

@vstinner

vstinner commented Oct 3, 2026 •

Copy link
Copy Markdown
Member Author

Oh. A side effect of this change is that b'abc' * 0 now returns the empty string singleton. I added a test for that.

@vstinner

vstinner commented Oct 3, 2026

Copy link
Copy Markdown
Member Author

I also ran a benchmark on PyBytes_FromStringAndSize("abc", 3) and PyBytes_FromString("abc"): there is no significant different on performance.

        PyObject *bytes = PyBytes_FromStringAndSize("abc", 3);
        if (bytes == NULL) {
            return NULL;
        }
        Py_DECREF(bytes);

and:

        PyObject *bytes = PyBytes_FromString("abc");
        if (bytes == NULL) {
            return NULL;
        }
        Py_DECREF(bytes);

Add bytes_alloc() helper function. bytes_alloc() and
_PyBytes_FromStringAndSize() have less conditional branches than
previous code, and so should be a little bit faster.

* Rename existing bytes_alloc() to bytes_type_alloc().
* Replace _PyBytes_FromSize(calloc=1) with _PyBytes_FromSizeZero().
* Add set_ob_shash_unsafe(): similar to set_ob_shash() but don't use
  an atomic operation on Free Threading. Use this new function
  on newly allocated bytes objects and in bytes_resize_inplace().
* Replace PyBytes_FromStringAndSize(NULL, size) with
  bytes_alloc(size).
* Use bytes_alloc() in PyBytes_FromString() and _PyBytes_Repeat().
* Add _PyBytes_FromStringAndSize(): similar to
  PyBytes_FromStringAndSize(), but str must not be NULL.
* Replace PyBytes_FromStringAndSize() with
  _PyBytes_FromStringAndSize().
* Test that PyBytes_FromString() and PyBytes_FromStringAndSize()
  return singletons for 0 or 1 bytes.
Previously, it creates an empty bytes string which became the empty
string singleton.
@vstinner

vstinner commented Oct 3, 2026

Copy link
Copy Markdown
Member Author

Ooops, I pushed "Add static inline _Py_NewReferenceInline() function" change by mistake. This time, I used git push --force to remove this commit. It shouldn't be part of this PR.

@vstinner
vstinner enabled auto-merge (squash) October 3, 2026 20:54
@eendebakpt

Copy link
Copy Markdown
Contributor

@vstinner Auto merge?

@eendebakpt

Copy link
Copy Markdown
Contributor

There are some corner cases that hit an assert in debug builds. Reproducer:

import sys

class B(bytes):
    pass

def rt(x):          # keep the compiler from constant-folding
    return x

empty = rt(b'')
print(sys.version)

cases = {
    "B() + B()":                  lambda: B() + B(),
    "B() + bytearray()":          lambda: B() + bytearray(),
    "B() + memoryview(b'')":      lambda: B() + memoryview(b''),
    "b'' + B()  (control)":       lambda: empty + B(),
}
for name, fn in cases.items():
    r = fn()                      # debug build: aborts here with
                                  # "Assertion failed: size >= 1 ... bytesobject.c"
    print(f"{name:28} -> {r!r:6} is b'': {r is empty}")

Note: returning the singleton more often is a nice-to-have, if we return another empty bytes in some cases that would be fine for me.

@vstinner
vstinner disabled auto-merge October 3, 2026 20:59
@vstinner

vstinner commented Oct 3, 2026

Copy link
Copy Markdown
Member Author

@vstinner Auto merge?

I consider that the PR is ready to be merged, so yes, I enabled auto merged. After I saw your comment with a regression, I disabled auto merge.

There are some corner cases that hit an assert in debug builds. Reproducer: (...)

Oh, well spotted! I misread if (va.len == 0 && PyBytes_CheckExact(b)) { in _PyBytes_Concat(), I missed the PyBytes_CheckExact() part of the condition.

It seems like this case wasn't tested at all by the Python test suite. I fixed bug and I added tests (test_concat_cpython).

Note: returning the singleton more often is a nice-to-have, if we return another empty bytes in some cases that would be fine for me.

The PR only changes bytes * int to return the empty string singleton if the number is zero. It's more a side effect of the implementation rather than a deliberate optimization :-)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants