Bug report
Bug description:
Documented behaviour: HTMLParser documentation, constructor: "If convert_charrefs is True (the default), all character references (except the ones in script/style elements) are automatically converted to the corresponding Unicode characters."
Expected: handle_starttag('div', [('title', 'A')]) without an exception
Actual: ValueError: integer string conversion exceeds the 4300-digit limit; handle_starttag is not called.
import html.parser
import re
import sys
if not hasattr(sys, "set_int_max_str_digits"):
print("REFUTATION REJECTED: integer-string conversion limit unavailable")
else:
sys.set_int_max_str_digits(4300)
text = '<div title="&#' + '0' * 4300 + '65;">'
match = re.fullmatch(r'<div title="&#([0-9]+);">', text)
if not match or match[1].lstrip("0") != "65":
print("REFUTATION REJECTED: input is not a valid decimal reference to U+0041")
else:
expected = ("OK", [("div", [("title", chr(65))])])
events = []
class Parser(html.parser.HTMLParser):
def handle_starttag(self, tag, attrs):
events.append((tag, attrs))
try:
Parser(convert_charrefs=True).feed(text)
actual = ("OK", events)
except Exception as e:
actual = ("EXCEPTION", type(e).__name__, str(e))
if actual != expected:
print("REFUTATION CONFIRMED:", repr(text),
"actual:", actual, "expected:", expected)
else:
print("REFUTATION REJECTED: actual matches documented expectation")
Output on Python 3.14.6 (Windows-11-10.0.26220-SP0), standard library html.parser:
0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000065;">' actual: ('EXCEPTION', 'ValueError', 'Exceeds the limit (4300 digits) for integer string conversion: value has 4302 digits; use sys.set_int_max_str_digits() to increase the limit') expected: ('OK', [('div', [('title', 'A')])])
This report was found and written by an automated property-testing tool I run (bugforge). The reproducer above was executed and its output is pasted unedited; no person reviewed the report before it was filed. The search script is in https://github.com/augusto-rehfeldt/bugforge-results/tree/main/html.parser-20261003-132957-c1
CPython versions tested on:
3.14
Operating systems tested on:
Windows
Bug report
Bug description:
Documented behaviour: HTMLParser documentation, constructor: "If convert_charrefs is True (the default), all character references (except the ones in script/style elements) are automatically converted to the corresponding Unicode characters."
Expected: handle_starttag('div', [('title', 'A')]) without an exception
Actual: ValueError: integer string conversion exceeds the 4300-digit limit; handle_starttag is not called.
Output on Python 3.14.6 (Windows-11-10.0.26220-SP0), standard library
html.parser:This report was found and written by an automated property-testing tool I run (bugforge). The reproducer above was executed and its output is pasted unedited; no person reviewed the report before it was filed. The search script is in https://github.com/augusto-rehfeldt/bugforge-results/tree/main/html.parser-20261003-132957-c1
CPython versions tested on:
3.14
Operating systems tested on:
Windows