Skip to content

Bump DspTap to the Helium Q15 dot; re-record the M55 baselines - #62

Merged
tap merged 1 commit into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6
Oct 2, 2026
Merged

tap merged 1 commit into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6

Conversation

@tap

@tap tap commented Oct 2, 2026

Copy link
Copy Markdown
Owner

What this changes

  • submodules/dsptap moves to fir_kernels: Helium Q15 dot, eight lanes per VMLALDAVA DspTap#53, the Helium Q15 dot kernel (d646905).

  • The M55 baselines are re-recorded in async/, bridge/ and rational/bench/baselines.json, and the README icount tables are regenerated with update_icount_docs.py. Only the M55 Q15 rows move:

    Workload Before After Change
    bridge up_q15_eco 44,556,679 19,095,905 −57 %
    bridge down_q15_eco 57,208,782 22,404,812 −61 %
    bridge up_q15_se 39,050,794 17,493,186 −55 %
    bridge down_q15_se 46,983,269 17,749,381 −62 %
    async pipeline12_q15 398,227,192 231,525,615 −42 %
    async pipeline_q15 137,792,386 109,235,651 −21 %
    rational down2_q15_eco 25,070,113 19,806,582 −21 %
    rational up2_q15_eco 47,493,112 34,150,590 −28 %
    rational up3_q15_eco 74,874,023 51,156,121 −32 %
    rational down3_q15_eco 28,017,447 21,227,446 −24 %
    rational down2_down2_q15_eco 54,506,280 28,178,778 −48 %
  • rational/PLAN.md gets the new M55 counts in its section 6 icount table. It also records that the M6 finding is resolved: Q15 on the M55 was no faster than float, and is now 19–32 % under it.

Why

This is the first of the deferred levers: the MVE Q15 kernel. Under arm-none-eabi-gcc 13 the substrate's Q15 dot compiled to one scalar SMLALBB per tap on Helium. The kernel now does eight lanes per VMLALDAVA. It is substrate code, so it landed in DspTap first.

Verification

  • Bit-exact: the kernel is bit-exact by construction (exact products, associative int64 sum), and DspTap's new EveryTapCountMatchesReference pins it at every tap count from 0 to 40.
  • M55 batteries: every engine's battery passes on the M55 under QEMU (17/17), including rational's and bridge's bit-pinned output hashes.
  • Host: 285/285 tests pass.
  • Baselines: recorded with --update, then re-run with --exact, and every count matched.
    • async 7/7, bridge 10/10, rational 12/12.
    • Every float and Q31 count is unchanged.
    • M33 and Hexagon don't compile the gated kernel, so their counts are unchanged.

Notes for the reviewer

  • Merge order: merge fir_kernels: Helium Q15 dot, eight lanes per VMLALDAVA DspTap#53 first. After that, I'll repoint this pin at the identical tree on DspTap main before this PR merges.
  • What's next: the remaining levers are the sparse rows for mixed-down stages, the Q15 decimators' per-branch quantization, and symmetry-halved tables. Each comes as its own measured PR.

🤖 Generated with Claude Code

https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA


Generated by Claude Code

The pin moves to tap/DspTap#53, whose Q15 dot_row, accumulate_row and
dot_row_reversed reduce eight lanes per VMLALDAVA on Helium cores.
Before it, arm-none-eabi-gcc 13 left them one scalar SMLALBB per tap.
The kernel is bit-exact by construction, so no output bit moves and
every battery passes on every leg.

Only the M55 Q15 counts move, all improvements beyond the ratchet's
tolerance, so they are re-recorded with the README tables regenerated:

- bridge Q15: -55 to -62 % (up_q15_eco 44.56 M -> 19.10 M)
- async Q15 pipelines: -21 % and -42 %
- rational Q15: -21 to -48 % (up3 74.87 M -> 51.16 M)

Every float and Q31 count, and every M33 and Hexagon count, is
unchanged. rational/PLAN.md records that the M6 finding (Q15 no faster
than float on the M55) is resolved: Q15 is now 19 to 32 % under float
there.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
@tap
tap merged commit 3e6592b into main Oct 2, 2026
28 checks passed
@tap
tap deleted the claude/sample-rate-expansion-strategies-ezqzu6 branch October 2, 2026 20:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants