rational: Q15 decimators quantized per branch; sparse rows and halved tables declined - #63
Merged
Merged
Conversation
… tables declined These are the codegen levers after M6, each measured (rational/PLAN.md section 6). Q15 decimators: the pin moves to tap/DspTap#54 for finalize_divided. Each of a Q15 decimator's M branches is now quantized at its own unity sum, and the 1/M rides the single rounding. Before, h / M was quantized as one row, which cost about 20 log10 M dB of stopband. - The Q15 decimator's table is now bit for bit its band's interpolator table (FixedPoint.Q15DecimatorTableIsTheInterpolatorsTable). - Attained stopband: down-6 / down-8 are -71.5 / -71.7 dB at economy (were -64.7 / -62.9) and -78.5 / -78.2 dB at transparent (were -62.9 / -61.1). down-3 at economy moves from -72.4 to -70.2 dB, still the 70 dB tier. - The Q15 RMS floor improves by 5 to 12 dB, and full-scale DC is still exact. Q31, float and double are unchanged. - basic_stage gains k_table_gain and finalize_output(), so tests and callers can see the convention. 16 Q15 table pins are re-pinned. Sparse rows for 2/3, 3/4 and 3/8 were built, measured and declined. Float, Q15 and Q31 outputs stayed bit-identical, but the per-output bookkeeping cost more than the third of the MACs it saved: Q15 was +22, +60 and +24 % on M33, M55 and Hexagon. The 2/3 workloads (ratio23_float_eco, ratio23_q15_eco) join the ratchet from that measurement. The symmetry-halved table was declined unbuilt: it saves memory only, 1.6 KB at most here. The generator's diff parser now skips section 6's ratio-keyed table. Baselines are re-recorded on M33, M55 and Hexagon. Construction moves +0.3 / +4.3 / +5.0 % (four branch quantizations where there was one), and down-3 Q15 streaming +2.3 % on Hexagon (the multiply-back). Every other count moved by under 0.6 %, and async and bridge are exactly unchanged. Co-Authored-By: Claude Opus 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
tap
pushed a commit
that referenced
this pull request
Oct 3, 2026
#63 merged with the pin on 54131fd, the PR branch commit; bb08c89 on DspTap main carries the identical tree and stays reachable after branch cleanup. Co-Authored-By: Claude Opus 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
This PR finishes the codegen levers after M6. Each one was measured, and the outcomes are recorded in
rational/PLAN.mdsection 6 (v0.10). The Helium Q15 dot (#62) was the first.tap::dsp::finalize_divided.A Q15 decimator now quantizes each of its M branches at the branch's own unity sum. Every branch of a Nyquist design sums to 1, so each keeps full Q1.14 precision. The 1/M is applied in the single rounding: a shift for ↓2 and ↓8, an exact-at-DC multiply-back for ↓3 and ↓6.
The Q15 decimator's table is now bit for bit its band's Q15 interpolator table. The new test
FixedPoint.Q15DecimatorTableIsTheInterpolatorsTablepins that.basic_stagegainsk_table_gainandfinalize_output(), so tests and callers can see the convention.Q15 decimator numbers (economy shown first, then transparent):
↓3 at economy gets 2.2 dB worse but stays in the 70 dB tier.
Full-scale DC is still exact. Q31, float and double are unchanged: every Q31 table pin holds.
16 Q15 table pins (hash, MACs, stopband, RMS) are re-pinned.
The design: input i goes on sub-line i mod M, so each row's residue classes of lag become contiguous and the class of structural zeros drops out. MACs per output fell to the nonzero count (2/3 at economy: 32.5 → 22.5), no matrix chain changed, and float, Q15 and Q31 outputs stayed bit-identical.
Measured on the new 2/3 workloads at economy:
The per-output bookkeeping costs more than it saves wherever the dot is cheap. The datapath change is not in this PR. The two 2/3 workloads (
ratio23_float_eco,ratio23_q15_eco) join the ratchet from that measurement.tools/coverage/matrix.py diffnow skips section 6's table, which is keyed by ratio. M5 had broken that parser.stage.hcontract,rational/README.md,rational/CLAUDE.md, and the plan's section 6 tables and limits.Why
You asked me to take up the deferred levers, and to drop B and D on their numbers.
Verification
-Werror. With gcc-Werror, 113/113 rational tests pass. clang-tidy and clang-format are clean.--exact, 14/14 matching on each target.matrix.py diffreports 0 of 182 rows changed.Notes for the reviewer
mainbefore this PR merges.d646905, not DspTapmain's5a30d05(identical tree). This PR's pin replaces it.transparent. The plan says so.🤖 Generated with Claude Code
https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Generated by Claude Code