Skip to content

rational M4: the chains and the coverage matrix, measured - #59

Merged
tap merged 6 commits into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6
Oct 2, 2026
Merged

tap merged 6 commits into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6

Conversation

@tap

@tap tap commented Oct 2, 2026

Copy link
Copy Markdown
Owner

What this changes

Milestone M4 of rational/PLAN.md: chain.h (basic_chain<S, R...> over DspTap's chain<>, the 20 named multi-stage chains), the committed coverage-matrix generator tools/coverage/matrix.py re-run on the pinned lengths, design.h's relaxation tables (profile::relaxed), and tests/test_matrix.cpp pinning and measuring all 182 rows of section 3. The plan goes to v0.7 with section 3 regenerated from measured counts. The DspTap pin moves to tap/DspTap#52 (chain<>::flush).

Why

A chain of stages has one profile, the chain's: f_pass = p · r_min at its lowest rate, so each stage is designed at p over its design divisor, its lower rate over r_min (the plan's 3.1: stages further up a chain are shorter). basic_chain computes that divisor per stage at compile time (the largest 2^a 3^b at or below the quotient, exact within a family; the quotient itself below 1, which is the one stage at bridge's rate under a 48-family r_min, 147/160) and constructs each stage at the profile relaxed by it, with pinned lengths per divisor so construction never searches. The generator costs every stage the same way and as stage.h trims its rows, so the matrix's MACs per output and latency are the shipping chains' numbers, exact.

Re-running the plan's selection rules on the pinned lengths changed 35 of the 182 rows against the harris-estimate tables (risk S1's ledger, in the plan's section 8): a relaxed half-band saturates at m = 7 (N = 27) while a relaxed third-band stage is m = 5 (N = 29) and serves three times the rate, so third-band, 4th-band and mixed stages win where the transition is wide (8/1 is ↑2 · 4/3 · ↑3, 1/8 is ↓3 · 3/4 · ↓2, 12/1 is ↑2 · ↑6, the ↑8 / ↓8 stage moves to the middle of the long chains). The rules were not changed; the plan notes that a tie-break toward integer-factor chains at a few per cent would be one edit to best_chain plus a regeneration, if the maintainer prefers it. The named-chain count for the C ABI is 28 (the union of the two tables' chain columns), where v0.2 counted 34 before taking the union.

Nothing is skipped: bridge's rate scale (2.2) had landed, so the twelve k > 0 rows are measured like the rest. The named ratios move from the umbrella to ratio.h because chain.h needs them; the public header count is 6.

Verification

  • Host, clang 18 and GCC with -Werror (the three TAP_SR_*_WERROR options on): the whole family battery passes (rational 99 ctest entries; GCC's run of the rational binary 93/93); scripts/tidy.sh clean on the new and changed TUs; clang-format clean.
  • test_matrix.cpp (host suites): every row's MACs per output and latency equal the generated table as exact rationals at all four profiles; every within-family row's compile-time divisors are the table's and its basic_chain reports the same numbers; the tone battery per profile (about 1700 tones over the 182 rows: seven passband tones and up to five stopband tones per row at exact bins of an analysis window that makes every rate of the chain a bin, so the rectangular-window DFT measures each image and alias candidate leakage-free): every candidate at or below f_pass is at or below −A — worst −71.34 dB at economy, −71.84 at super_economy, −71.52 at balanced, −121.63 at transparent — and the passband within the stages' summed ripple candidates (worst 0.0087 dB at a 70 dB tier, 0.00002 dB at transparent); accounting from three positions, the impulse within half a frame of the latency and flush on every row. About 24 s for the four batteries on a host.
  • test_chain.cpp: the named chains' divisors as static_asserts; a chain equals its stages run in sequence bit for bit in float, double, Q15 and Q31; seven chains against sequenced-upfirdn vectors (tools/reference/make_reference_vectors.py regenerated; the single-stage vectors are byte-identical) within the single stages' tolerances; accounting exact from every position of a superblock of all 20 named chains; chunking, reset, flush as zero padding, latency by impulse, MACs, DC, two channels.
  • test_design.cpp: every entry of the four relaxation tables is what search_nyquist_m finds on the design grid with the margin (the M2 criterion).
  • Cortex-M33 QEMU leg, run locally: the rational one-shot passes in 158 s with the matrix suite and the relaxation-table searches excluded by bare_metal_main.cpp (a first run with the 70 dB search included hit the 1800 s ctest timeout under soft double; the second commit excludes it and trims the chain accounting sweep to one superblock). Hexagon's row excludes the tone batteries and the transparent search by name. Not run here: M55 and Hexagon, which CI covers.
  • tools/coverage/matrix.py md|pins|rows|diff|rule are deterministic; diff against the v0.6 tables lists the 35 changed rows, rule the 24 rows where the fewest-stages rule bit.

Notes for the reviewer

  • Submodule pin moved. submodules/dsptap points at chain: flush, the stream drained stage by stage DspTap#52's branch commit (d4a61a9, chain<>::flush, CI green there); once Align the icount marker's format string and re-record the Hexagon baselines #52 merges by rebase I will repoint at the identical tree on DspTap main and amend this branch.
  • Contract change, this engine only. profile gains relaxations (a span of the pinned rows) and relaxed(divisor); the named ratios (up_2 … ratio_3_8) are declared in ratio.h now (the umbrella still exports them). No consumer other than this tree's tests exists yet.
  • The matrix's chain shapes are the plan's rules applied to measured numbers; the ledger in the plan (S1) lists every flip and its cause. If you would rather keep integer-factor chains where a mixed stage saves only a few per cent, say so and I regenerate with that tie-break.
  • tests/coverage/matrix_rows.h and tests/reference/reference_vectors.h are generated; the generators' modes are documented in their docstrings.

🤖 Generated with Claude Code

https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA


Generated by Claude Code

claude added 6 commits October 2, 2026 13:35
chain.h: basic_chain<S, R...> over DspTap's chain<> (its flush landed as
tap/DspTap#52; the submodule pin follows), constructing each stage at the
profile relaxed by the stage's design divisor — its lower rate over the
chain's lowest, the largest 2^a 3^b at or below it, or the quotient
itself below 1 (the stage at bridge's rate under a 48-family r_min,
147/160) — computed at compile time; chain<R...> / chain_q15 / chain_q31;
macs_per_output() exact; the 20 named multi-stage chains of the matrix.
design.h: profile::relaxed(divisor) and the relaxation tables, the pinned
m per band at every divisor the matrix uses, found by the same search as
the M2 pins and verified against the shipping designer by test_design.
The named ratios move from the umbrella to ratio.h (chain.h needs them).

tools/coverage/matrix.py replaces the harris estimates with the numpy
search on the 16384-point grid per (profile, band, divisor), costs every
stage as stage.h trims its rows, and re-runs the plan's selection rules
over all 182 pairs; it writes tests/coverage/matrix_rows.h (every row's
stages, divisors and exact MACs / latency per profile) and prints the
plan's tables, the relaxation tables, the rows that changed and the rows
where the fewest-stages rule bit. 35 rows changed chain against the
estimates (the plan's S1 ledger): a relaxed half-band saturates at m = 7
while a relaxed third-band stage is m = 5, so third-band, 4th-band and
mixed stages win where the transition is wide.

test_matrix.cpp builds every row in double (the cross rows through the
bridge adapter in tests/support, 4.2), pins MACs per output and latency
exactly at all four profiles, checks the within-family rows are the named
chains, measures each row against its promise with a tone battery at
exact bins (every image and alias candidate at or below f_pass is at or
below -A; worst -71.3 dB at economy, -121.6 at transparent; passband
within 0.009 dB), and checks accounting, the impulse and flush. Nothing
is skipped: bridge's rate scale (2.2) had landed, so the k > 0 rows are
measured too. test_chain.cpp: a chain equals its stages in sequence bit
for bit in every format, sequenced-upfirdn vectors for seven chains,
accounting from every position of every named chain, chunking, reset,
flush as zero padding, latency, MACs, DC, channels.

Plan v0.7: section 3 regenerated from the measured counts, the design
divisor in 3.1, the fewest-stages list, the C ABI's 28 named chains (the
union v0.2 counted as 34), M4 done; READMEs and CLAUDE files follow. The
matrix suite and the transparent relaxation search are host suites,
excluded on the QEMU legs; the public header count is 6.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
The relaxation tables' 70 dB search joins the transparent one in the
bare-metal exclusions (design math covered on every host; under
qemu-system-arm's soft double it alone takes minutes), and the chains'
accounting sweep covers one superblock of positions with probes a
superblock long instead of two (every phase of every stage is still
visited; a quarter of the copies).

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
test_matrix.cpp is a host measurement (182 chains in double through a
tone battery, bridge composed in) and bare_metal_main.cpp already
excludes it by name, but its 182 instantiations alone overflow the M55
image's CODE region by 53 KB at link time, so the TU is only part of the
hosted and Hexagon builds.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Every row of the coverage matrix is a distinct chain type, and with the
measurement and streaming code templated per row the TU took eight
minutes to compile on a host (and the Hexagon cross job longer). The
rows are now built once into a type-erased handle (process, accounting,
flush, latency, MACs, clone), so the tone battery, the streaming checks
and the structural pins are written and instantiated once and a row's
template code is its factory alone: the same tests, 129 s to compile.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
… on main

The 70 dB relaxation-table search (3 s on a host) ran past the Hexagon
job's 20-minute budget under qemu-hexagon, where the design grid's
double arithmetic runs some hundreds of times slower (the coarse-grid
comparison takes 187 s there); both searches are host suites now, as
they already were on the Arm legs. The submodule pin moves from
tap/DspTap#52's branch commit to its merge on DspTap main (537ab30, the
identical tree).

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
The regex edit did not apply in ad0bccf (only the submodule pin moved);
this is the exclusion of both relaxation-table searches on Hexagon.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
@tap
tap merged commit 2e9b349 into main Oct 2, 2026
27 checks passed
@tap
tap deleted the claude/sample-rate-expansion-strategies-ezqzu6 branch October 2, 2026 16:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants