rational M4: the chains and the coverage matrix, measured - #59
Merged
Merged
Conversation
chain.h: basic_chain<S, R...> over DspTap's chain<> (its flush landed as tap/DspTap#52; the submodule pin follows), constructing each stage at the profile relaxed by the stage's design divisor — its lower rate over the chain's lowest, the largest 2^a 3^b at or below it, or the quotient itself below 1 (the stage at bridge's rate under a 48-family r_min, 147/160) — computed at compile time; chain<R...> / chain_q15 / chain_q31; macs_per_output() exact; the 20 named multi-stage chains of the matrix. design.h: profile::relaxed(divisor) and the relaxation tables, the pinned m per band at every divisor the matrix uses, found by the same search as the M2 pins and verified against the shipping designer by test_design. The named ratios move from the umbrella to ratio.h (chain.h needs them). tools/coverage/matrix.py replaces the harris estimates with the numpy search on the 16384-point grid per (profile, band, divisor), costs every stage as stage.h trims its rows, and re-runs the plan's selection rules over all 182 pairs; it writes tests/coverage/matrix_rows.h (every row's stages, divisors and exact MACs / latency per profile) and prints the plan's tables, the relaxation tables, the rows that changed and the rows where the fewest-stages rule bit. 35 rows changed chain against the estimates (the plan's S1 ledger): a relaxed half-band saturates at m = 7 while a relaxed third-band stage is m = 5, so third-band, 4th-band and mixed stages win where the transition is wide. test_matrix.cpp builds every row in double (the cross rows through the bridge adapter in tests/support, 4.2), pins MACs per output and latency exactly at all four profiles, checks the within-family rows are the named chains, measures each row against its promise with a tone battery at exact bins (every image and alias candidate at or below f_pass is at or below -A; worst -71.3 dB at economy, -121.6 at transparent; passband within 0.009 dB), and checks accounting, the impulse and flush. Nothing is skipped: bridge's rate scale (2.2) had landed, so the k > 0 rows are measured too. test_chain.cpp: a chain equals its stages in sequence bit for bit in every format, sequenced-upfirdn vectors for seven chains, accounting from every position of every named chain, chunking, reset, flush as zero padding, latency, MACs, DC, channels. Plan v0.7: section 3 regenerated from the measured counts, the design divisor in 3.1, the fewest-stages list, the C ABI's 28 named chains (the union v0.2 counted as 34), M4 done; READMEs and CLAUDE files follow. The matrix suite and the transparent relaxation search are host suites, excluded on the QEMU legs; the public header count is 6. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
The relaxation tables' 70 dB search joins the transparent one in the bare-metal exclusions (design math covered on every host; under qemu-system-arm's soft double it alone takes minutes), and the chains' accounting sweep covers one superblock of positions with probes a superblock long instead of two (every phase of every stage is still visited; a quarter of the copies). Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
test_matrix.cpp is a host measurement (182 chains in double through a tone battery, bridge composed in) and bare_metal_main.cpp already excludes it by name, but its 182 instantiations alone overflow the M55 image's CODE region by 53 KB at link time, so the TU is only part of the hosted and Hexagon builds. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Every row of the coverage matrix is a distinct chain type, and with the measurement and streaming code templated per row the TU took eight minutes to compile on a host (and the Hexagon cross job longer). The rows are now built once into a type-erased handle (process, accounting, flush, latency, MACs, clone), so the tone battery, the streaming checks and the structural pins are written and instantiated once and a row's template code is its factory alone: the same tests, 129 s to compile. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
… on main The 70 dB relaxation-table search (3 s on a host) ran past the Hexagon job's 20-minute budget under qemu-hexagon, where the design grid's double arithmetic runs some hundreds of times slower (the coarse-grid comparison takes 187 s there); both searches are host suites now, as they already were on the Arm legs. The submodule pin moves from tap/DspTap#52's branch commit to its merge on DspTap main (537ab30, the identical tree). Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
The regex edit did not apply in ad0bccf (only the submodule pin moved); this is the exclusion of both relaxation-table searches on Hexagon. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
Milestone M4 of
rational/PLAN.md:chain.h(basic_chain<S, R...>over DspTap'schain<>, the 20 named multi-stage chains), the committed coverage-matrix generatortools/coverage/matrix.pyre-run on the pinned lengths,design.h's relaxation tables (profile::relaxed), andtests/test_matrix.cpppinning and measuring all 182 rows of section 3. The plan goes to v0.7 with section 3 regenerated from measured counts. The DspTap pin moves to tap/DspTap#52 (chain<>::flush).Why
A chain of stages has one profile, the chain's: f_pass = p · r_min at its lowest rate, so each stage is designed at p over its design divisor, its lower rate over r_min (the plan's 3.1: stages further up a chain are shorter).
basic_chaincomputes that divisor per stage at compile time (the largest 2^a 3^b at or below the quotient, exact within a family; the quotient itself below 1, which is the one stage at bridge's rate under a 48-family r_min, 147/160) and constructs each stage at the profile relaxed by it, with pinned lengths per divisor so construction never searches. The generator costs every stage the same way and asstage.htrims its rows, so the matrix's MACs per output and latency are the shipping chains' numbers, exact.Re-running the plan's selection rules on the pinned lengths changed 35 of the 182 rows against the harris-estimate tables (risk S1's ledger, in the plan's section 8): a relaxed half-band saturates at m = 7 (N = 27) while a relaxed third-band stage is m = 5 (N = 29) and serves three times the rate, so third-band, 4th-band and mixed stages win where the transition is wide (8/1 is ↑2 · 4/3 · ↑3, 1/8 is ↓3 · 3/4 · ↓2, 12/1 is ↑2 · ↑6, the ↑8 / ↓8 stage moves to the middle of the long chains). The rules were not changed; the plan notes that a tie-break toward integer-factor chains at a few per cent would be one edit to
best_chainplus a regeneration, if the maintainer prefers it. The named-chain count for the C ABI is 28 (the union of the two tables' chain columns), where v0.2 counted 34 before taking the union.Nothing is skipped: bridge's rate scale (2.2) had landed, so the twelve k > 0 rows are measured like the rest. The named ratios move from the umbrella to
ratio.hbecausechain.hneeds them; the public header count is 6.Verification
-Werror(the threeTAP_SR_*_WERRORoptions on): the whole family battery passes (rational 99 ctest entries; GCC's run of the rational binary 93/93);scripts/tidy.shclean on the new and changed TUs; clang-format clean.test_matrix.cpp(host suites): every row's MACs per output and latency equal the generated table as exact rationals at all four profiles; every within-family row's compile-time divisors are the table's and itsbasic_chainreports the same numbers; the tone battery per profile (about 1700 tones over the 182 rows: seven passband tones and up to five stopband tones per row at exact bins of an analysis window that makes every rate of the chain a bin, so the rectangular-window DFT measures each image and alias candidate leakage-free): every candidate at or below f_pass is at or below −A — worst −71.34 dB ateconomy, −71.84 atsuper_economy, −71.52 atbalanced, −121.63 attransparent— and the passband within the stages' summed ripple candidates (worst 0.0087 dB at a 70 dB tier, 0.00002 dB attransparent); accounting from three positions, the impulse within half a frame of the latency and flush on every row. About 24 s for the four batteries on a host.test_chain.cpp: the named chains' divisors asstatic_asserts; a chain equals its stages run in sequence bit for bit in float, double, Q15 and Q31; seven chains against sequenced-upfirdn vectors (tools/reference/make_reference_vectors.pyregenerated; the single-stage vectors are byte-identical) within the single stages' tolerances; accounting exact from every position of a superblock of all 20 named chains; chunking, reset, flush as zero padding, latency by impulse, MACs, DC, two channels.test_design.cpp: every entry of the four relaxation tables is whatsearch_nyquist_mfinds on the design grid with the margin (the M2 criterion).bare_metal_main.cpp(a first run with the 70 dB search included hit the 1800 s ctest timeout under soft double; the second commit excludes it and trims the chain accounting sweep to one superblock). Hexagon's row excludes the tone batteries and the transparent search by name. Not run here: M55 and Hexagon, which CI covers.tools/coverage/matrix.py md|pins|rows|diff|ruleare deterministic;diffagainst the v0.6 tables lists the 35 changed rows,rulethe 24 rows where the fewest-stages rule bit.Notes for the reviewer
submodules/dsptappoints at chain: flush, the stream drained stage by stage DspTap#52's branch commit (d4a61a9,chain<>::flush, CI green there); once Align the icount marker's format string and re-record the Hexagon baselines #52 merges by rebase I will repoint at the identical tree on DspTapmainand amend this branch.profilegainsrelaxations(a span of the pinned rows) andrelaxed(divisor); the named ratios (up_2…ratio_3_8) are declared inratio.hnow (the umbrella still exports them). No consumer other than this tree's tests exists yet.tests/coverage/matrix_rows.handtests/reference/reference_vectors.hare generated; the generators' modes are documented in their docstrings.🤖 Generated with Claude Code
https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Generated by Claude Code