Feature or enhancement
Currently in _BINARY_OP_SUBSCR_USTR_INT specialisation we look up just ascii tables. This results in specialisation fails when a character is latin1 encoded which might be common in many European texts.
A small extension to support latin1 would enable the specialisation to support these characters such as
PyObject *res_o = (c < 128)
? (PyObject*)&_Py_SINGLETON(strings).ascii[c]
: (PyObject*)&_Py_SINGLETON(strings).latin1[c - 128];
Scanning some french text results in a 14% speedup with the test done on MacOS. This avoids any fallback to the generic path for the latin characters.
def scan(text):
i = 0
n = len(text)
while i < n:
c = text[i] # the only str[int] site
i += 1
ASCII = "The quick brown fox jumps over the lazy dog. " * 20
FRENCH = "Le garçon préfère les crêpes à la fraîcheur du matin. " * 20
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response
Linked PRs
Feature or enhancement
Currently in
_BINARY_OP_SUBSCR_USTR_INTspecialisation we look up just ascii tables. This results in specialisation fails when a character islatin1encoded which might be common in many European texts.A small extension to support latin1 would enable the specialisation to support these characters such as
Scanning some french text results in a 14% speedup with the test done on MacOS. This avoids any fallback to the generic path for the latin characters.
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response
Linked PRs