Skip to content

rational: Q15 decimators quantized per branch; sparse rows and halved tables declined - #63

Merged
tap merged 1 commit into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6
Oct 3, 2026
Merged

tap merged 1 commit into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6

Conversation

@tap

@tap tap commented Oct 2, 2026

Copy link
Copy Markdown
Owner

What this changes

This PR finishes the codegen levers after M6. Each one was measured, and the outcomes are recorded in rational/PLAN.md section 6 (v0.10). The Helium Q15 dot (#62) was the first.

  • Q15 decimators quantized per branch (shipped). The pin moves to sample_traits: finalize_divided, the 1/D folded into the single rounding DspTap#54 for tap::dsp::finalize_divided.
    • A Q15 decimator now quantizes each of its M branches at the branch's own unity sum. Every branch of a Nyquist design sums to 1, so each keeps full Q1.14 precision. The 1/M is applied in the single rounding: a shift for ↓2 and ↓8, an exact-at-DC multiply-back for ↓3 and ↓6.

    • The Q15 decimator's table is now bit for bit its band's Q15 interpolator table. The new test FixedPoint.Q15DecimatorTableIsTheInterpolatorsTable pins that.

    • basic_stage gains k_table_gain and finalize_output(), so tests and callers can see the convention.

    • Q15 decimator numbers (economy shown first, then transparent):

      Stage Stopband before Stopband after RMS vs double before RMS vs double after
      ↓2 −71.6 / −69.2 dB −72.3 / −77.3 dB −92.8 / −89.3 dBFS −97.5 / −95.7 dBFS
      ↓3 −72.4 / −63.8 dB −70.2 / −76.1 dB −90.8 / −86.4 dBFS −96.8 / −95.7 dBFS
      ↓6 −64.7 / −62.9 dB −71.5 / −78.5 dB −86.5 / −85.6 dBFS −98.6 / −98.3 dBFS
      ↓8 −62.9 / −61.1 dB −71.7 / −78.2 dB −86.8 / −85.6 dBFS −97.9 / −99.5 dBFS

      ↓3 at economy gets 2.2 dB worse but stays in the 70 dB tier.

    • Full-scale DC is still exact. Q31, float and double are unchanged: every Q31 table pin holds.

    • 16 Q15 table pins (hash, MACs, stopband, RMS) are re-pinned.

  • Sparse rows for the mixed ratios going down (built, measured, declined).
    • The design: input i goes on sub-line i mod M, so each row's residue classes of lag become contiguous and the class of structural zeros drops out. MACs per output fell to the nonzero count (2/3 at economy: 32.5 → 22.5), no matrix chain changed, and float, Q15 and Q31 outputs stayed bit-identical.

    • Measured on the new 2/3 workloads at economy:

      Format M33 M55 Hexagon
      float −22 % +13 % −25 %
      Q15 +22 % +60 % +24 %
    • The per-output bookkeeping costs more than it saves wherever the dot is cheap. The datapath change is not in this PR. The two 2/3 workloads (ratio23_float_eco, ratio23_q15_eco) join the ratchet from that measurement.

  • Symmetry-halved table (declined, not built). It saves memory only: rational's largest table is 399 coefficients (1.6 KB float), and on the M55 the reversed Q15 dot uses a gather load, which is slower than a contiguous one.
  • Generator fix: tools/coverage/matrix.py diff now skips section 6's table, which is keyed by ratio. M5 had broken that parser.
  • Docs: stage.h contract, rational/README.md, rational/CLAUDE.md, and the plan's section 6 tables and limits.

Why

You asked me to take up the deferred levers, and to drop B and D on their numbers.

Verification

  • Host: 286/286 tests pass with clang 18 -Werror. With gcc -Werror, 113/113 rational tests pass. clang-tidy and clang-format are clean.
  • Embedded: M33 and M55 under QEMU pass 17/17 each, every engine.
  • Ratchet: re-recorded on M33, M55 and Hexagon, then re-run with --exact, 14/14 matching on each target.
    • Construction costs +0.3 % (M33), +4.3 % (M55) and +5.0 % (Hexagon): four branch quantizations where there was one.
    • ↓3 Q15 streaming costs +2.3 % on Hexagon, from the multiply-back.
    • Every other count moved by under 0.6 %.
    • async and bridge are exactly unchanged on all three targets.
  • Generator: matrix.py diff reports 0 of 182 rows changed.

Notes for the reviewer

🤖 Generated with Claude Code

https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA


Generated by Claude Code

… tables declined

These are the codegen levers after M6, each measured (rational/PLAN.md
section 6).

Q15 decimators: the pin moves to tap/DspTap#54 for finalize_divided.
Each of a Q15 decimator's M branches is now quantized at its own unity
sum, and the 1/M rides the single rounding. Before, h / M was quantized
as one row, which cost about 20 log10 M dB of stopband.

- The Q15 decimator's table is now bit for bit its band's interpolator
  table (FixedPoint.Q15DecimatorTableIsTheInterpolatorsTable).
- Attained stopband: down-6 / down-8 are -71.5 / -71.7 dB at economy
  (were -64.7 / -62.9) and -78.5 / -78.2 dB at transparent (were -62.9 /
  -61.1). down-3 at economy moves from -72.4 to -70.2 dB, still the
  70 dB tier.
- The Q15 RMS floor improves by 5 to 12 dB, and full-scale DC is still
  exact. Q31, float and double are unchanged.
- basic_stage gains k_table_gain and finalize_output(), so tests and
  callers can see the convention. 16 Q15 table pins are re-pinned.

Sparse rows for 2/3, 3/4 and 3/8 were built, measured and declined.
Float, Q15 and Q31 outputs stayed bit-identical, but the per-output
bookkeeping cost more than the third of the MACs it saved: Q15 was +22,
+60 and +24 % on M33, M55 and Hexagon. The 2/3 workloads
(ratio23_float_eco, ratio23_q15_eco) join the ratchet from that
measurement.

The symmetry-halved table was declined unbuilt: it saves memory only,
1.6 KB at most here. The generator's diff parser now skips section 6's
ratio-keyed table.

Baselines are re-recorded on M33, M55 and Hexagon. Construction moves
+0.3 / +4.3 / +5.0 % (four branch quantizations where there was one),
and down-3 Q15 streaming +2.3 % on Hexagon (the multiply-back). Every
other count moved by under 0.6 %, and async and bridge are exactly
unchanged.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
@tap
tap merged commit 4eb1180 into main Oct 3, 2026
28 checks passed
@tap
tap deleted the claude/sample-rate-expansion-strategies-ezqzu6 branch October 3, 2026 01:18
tap pushed a commit that referenced this pull request Oct 3, 2026
#63 merged with the pin on 54131fd, the PR branch commit; bb08c89 on
DspTap main carries the identical tree and stays reachable after branch
cleanup.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants