Skip to content

Read interactive input files as UTF-8 - #433

Open
Mohammed Alkindi (MohammedAlkindi) wants to merge 2 commits into
microsoft:mainfrom
MohammedAlkindi:fix/interactive-input-utf8
Open

Mohammed Alkindi (MohammedAlkindi) wants to merge 2 commits into
microsoft:mainfrom
MohammedAlkindi:fix/interactive-input-utf8

Conversation

@MohammedAlkindi

@MohammedAlkindi Mohammed Alkindi (MohammedAlkindi) commented Sep 15, 2026 •

Copy link
Copy Markdown

process_requests opens the input file with no encoding, so it follows locale.getpreferredencoding(False). On a Windows box with a cp1252 code page that fails two ways. Measured on CPython 3.13.13, getpreferredencoding = cp1252:

UTF-8 file, Japanese  -> UnicodeDecodeError: 'charmap' codec can't decode byte 0x81
UTF-8 file, "café"    -> 'café au lait'   (no exception, handed to the request handler)

The quiet one is the worse of the two, since that mojibake goes into the model prompt.

The TypeScript twin of this function reads the same files with fs.readFileSync(inputFileName).toString(), which is UTF-8 on every platform, so this also makes the two ports agree rather than picking an encoding arbitrarily.

One behaviour change worth stating: an input file actually encoded as cp1252, which currently reads on Windows, now raises. That converts a silently-wrong path into either correct output or a loud error. All ten tracked input*.txt files are ASCII, so nothing shipped is affected.

The test simulates a non-UTF-8 default rather than reading the host locale, so it exercises the same path on Linux CI. It fails on the unmodified source with the UnicodeDecodeError above.

Same file and function as #248.

Bare open() follows locale.getpreferredencoding(False), so on a Windows box with a cp1252 code page an input file either raises UnicodeDecodeError or, when its characters happen to map into cp1252, is read as mojibake with no error and passed to the request handler.

The TypeScript twin reads the same files with fs.readFileSync().toString(), which is UTF-8 on every platform, so this also makes the two ports agree.

@robgruen robgruen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe add a café regression test to explicitly cover silent corruption.

The existing non-ASCII case raises UnicodeDecodeError under a cp1252
default encoding, so it only pins the loud failure. "café" decodes to
"café" instead of raising, which is the case a non-UTF-8 read
corrupts without any error.
@MohammedAlkindi

Copy link
Copy Markdown
Author

Added in 3e61b99: a new test writes a café line as UTF-8 under a cp1252 default encoding and asserts it reads back unchanged. cp1252 decodes those bytes without raising, so that is the case where the old read corrupted input silently.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants