Skip to content

Avoid redundant normalization of parsed wire headers - #207

Open
Johnny-Kao wants to merge 2 commits into
python-hyper:masterfrom
Johnny-Kao:perf/fuse-wire-header-normalization
Open

Johnny-Kao wants to merge 2 commits into
python-hyper:masterfrom
Johnny-Kao:perf/fuse-wire-header-normalization

Conversation

@Johnny-Kao

Copy link
Copy Markdown

Summary

  • build canonical Headers directly while decoding header lines from the wire
  • share Content-Length / Transfer-Encoding semantic normalization with the existing public normalization path
  • preserve the existing syntax-validation-before-semantic-validation ordering

Why

The receive path currently validates each wire header line, materializes a temporary (name, value) list, and then passes that list through normalize_and_validate(..., _parsed=True), which walks the headers again to lowercase names, normalize Content-Length / Transfer-Encoding, and construct the internal three-tuple representation.

This change keeps the existing trust boundary and protocol checks, but removes the intermediate pair representation and redundant traversal. The public Request / Response construction path is unchanged.

Correctness / safety

The parser still validates every header line before applying Content-Length / Transfer-Encoding semantic checks. A regression test locks in that ordering.

Local validation:

  • 79 passed
  • Black 23.3.0: pass
  • isort 5.12.0: pass
  • mypy 1.8.0 strict config: pass
  • git diff --check: pass
  • 50,008 request-level differential cases matched the previous parser exactly for accept/reject behavior, error hint/message, request metadata, raw header casing/order, and canonical header values

The change does not alter obs-fold behavior; it is independent of #198.

Performance

Interleaved A/B measurements in the same Python process were used to reduce scheduler/frequency noise.

On the local Apple Silicon environment:

  • repository realistic server GET cycle: ~3.3% lower median time (27.45 us -> 26.54 us)
  • request receive / Connection.next_event() path: ~5.6% lower median time
  • response header reader: ~11.7% lower median time
  • response-parser traced peak memory in the test workload: ~11.4 KB -> ~6.3 KB

The performance claim is intentionally limited to these parsing/benchmark workloads rather than h11 overall.

@Johnny-Kao

Copy link
Copy Markdown
Author

First time contributing to h11 — this was a really interesting codebase to work on. The implementation is impressively solid, especially around the protocol and validation boundaries. Happy to adjust anything if needed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant