Skip to content

Fix ParserError on subsecond precision beyond nanoseconds - #1020

Open
afonsojanu wants to merge 1 commit into
python-pendulum:masterfrom
afonsojanu:fix/iso8601-subsecond-beyond-nanoseconds
Open

afonsojanu wants to merge 1 commit into
python-pendulum:masterfrom
afonsojanu:fix/iso8601-subsecond-beyond-nanoseconds

Conversation

@afonsojanu

Copy link
Copy Markdown

Summary

pendulum.parse() raises ParserError for an otherwise valid-looking timestamp when the fractional-seconds part has more than 9 digits, e.g.:

>>> pendulum.parse("2001-01-01T12:34:56.1234567890Z")
Traceback (most recent call last):
  ...
pendulum.parsing.exceptions.ParserError: Unable to parse string [2001-01-01T12:34:56.1234567890Z]

datetime.fromisoformat() has no such limit and just truncates to microseconds:

>>> datetime.fromisoformat("2001-01-01T12:34:56.1234567890+00:00")
datetime.datetime(2001, 1, 1, 12, 34, 56, 123456, tzinfo=datetime.timezone.utc)

Reproduces with the "common" datetime format too (this path has no Rust equivalent, so it's hit regardless of the compiled extension):

>>> pendulum.parse("2016/10/06 12:34:56.1234567890")
pendulum.parsing.exceptions.ParserError: Unable to parse string [2016/10/06 12:34:56.1234567890]

Root cause

Both the pure-Python ISO 8601 parser (pendulum/parsing/iso8601.py, used as a fallback when the compiled extension isn't available, e.g. PENDULUM_EXTENSIONS=0) and the "common" format parser (pendulum/parsing/__init__.py, always pure Python) match the subsecond group with \d{1,9}. A 10th digit makes the whole regex fail to match, so the entire string is rejected.

Both call sites already truncate the captured group to 6 characters a few lines below before converting it to microseconds:

subsecond = m.group("subsecond")[:6]
microsecond = int(f"{subsecond:0<6}")

so the {1,9} cap wasn't protecting anything downstream, it was just rejecting valid input before it ever got there.

Fix

Widened both regexes from \d{1,9} to \d+. The existing truncate-to-6-and-pad logic already does the right thing for any number of digits, so no other changes were needed.

Tests

  • Added test_rfc_3339_extended_beyond_nanoseconds and test_common_format_extended_beyond_nanoseconds to tests/parsing/test_parsing.py, covering the public parse() API for both the ISO 8601 and "common" format paths.
  • Added test_parse_iso8601_subsecond_beyond_nanoseconds_pure_python to tests/parsing/test_parse_iso8601.py, importing pendulum.parsing.iso8601.parse_iso8601 directly so the pure-Python parser is exercised even when the compiled extension is installed (the Rust parser doesn't have this bug since it isn't regex-based).

Confirmed all three new tests fail with ParserError against the unpatched regexes and pass after the fix. Ran the full test suite (pytest tests/, with the Rust extension built via maturin develop --release): 1850 passed, 3 skipped (pre-existing, unrelated to this change), no new failures or warnings.

Closes #935.

The pure-Python datetime parsers (used as a fallback when the compiled
Rust extension is unavailable, e.g. PENDULUM_EXTENSIONS=0, and for the
"common" format that has no Rust equivalent) capped the fractional
seconds group at 9 digits (\\d{1,9}). A 10th digit made the whole
string fail to match, raising ParserError instead of parsing it.

The value is already truncated to 6 digits (microseconds) a few lines
below, so the cap served no purpose beyond rejecting otherwise valid
input. Python's datetime.fromisoformat() has no such limit and just
truncates, which is the behavior restored here by widening both
regexes to \\d+.

Fixes ParserError on inputs like:
  pendulum.parse("2001-01-01T12:34:56.1234567890Z")
  pendulum.parse("2016/10/06 12:34:56.1234567890")

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

parse fails with subsecond decimals with more than 9 digits

1 participant