Bug report
Bug description:
Documented behaviour: HTMLParser documentation, constructor: "If convert_charrefs is True (the default), all character references (except the ones in script/style elements) are automatically converted to the corresponding Unicode characters." HTML character-reference conversion maps out-of-range numeric values to U+FFFD.
Expected: handle_data receives U+FFFD, with feed() and close() completing without exceptions.
Actual: ValueError: integer string conversion exceeds the 4300-digit limit for the 4301-digit reference.
import html.parser
import re
s = '&#' + '9' * 4301 + ';'
digits = s[2:-1]
if not re.fullmatch(r'&#[0-9]+;', s):
print('REFUTATION REJECTED: invalid decimal character reference')
else:
d = digits.lstrip('0') or '0'
expected = ('data', '\ufffd' if (len(d), d) > (7, '1114111') else chr(int(d)))
class Parser(html.parser.HTMLParser):
def handle_data(self, data):
chunks.append(data)
chunks = []
try:
p = Parser(convert_charrefs=True)
p.feed(s)
p.close()
actual = ('data', ''.join(chunks))
except Exception as e:
actual = ('exception', type(e).__name__, str(e))
if actual != expected:
print('REFUTATION CONFIRMED:', repr(s), 'actual:', actual, 'expected:', expected)
else:
print('REFUTATION REJECTED: actual matches documented expectation')
Output on Python 3.14.6 (Windows-11-10.0.26220-SP0), standard library html.parser:
9999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999999;' actual: ('exception', 'ValueError', 'Exceeds the limit (4300 digits) for integer string conversion: value has 4301 digits; use sys.set_int_max_str_digits() to increase the limit') expected: ('data', '�')
This report was found and written by an automated property-testing tool I run (bugforge). The reproducer above was executed and its output is pasted unedited; no person reviewed the report before it was filed. The search script is in https://github.com/augusto-rehfeldt/bugforge-results/tree/main/html.parser-20261003-193928-c4
CPython versions tested on:
3.14
Operating systems tested on:
Windows
Bug report
Bug description:
Documented behaviour: HTMLParser documentation, constructor: "If convert_charrefs is True (the default), all character references (except the ones in script/style elements) are automatically converted to the corresponding Unicode characters." HTML character-reference conversion maps out-of-range numeric values to U+FFFD.
Expected: handle_data receives U+FFFD, with feed() and close() completing without exceptions.
Actual: ValueError: integer string conversion exceeds the 4300-digit limit for the 4301-digit reference.
Output on Python 3.14.6 (Windows-11-10.0.26220-SP0), standard library
html.parser:This report was found and written by an automated property-testing tool I run (bugforge). The reproducer above was executed and its output is pasted unedited; no person reviewed the report before it was filed. The search script is in https://github.com/augusto-rehfeldt/bugforge-results/tree/main/html.parser-20261003-193928-c4
CPython versions tested on:
3.14
Operating systems tested on:
Windows