Bump DspTap to the Helium Q15 dot; re-record the M55 baselines - #62
Merged
Merged
Conversation
The pin moves to tap/DspTap#53, whose Q15 dot_row, accumulate_row and dot_row_reversed reduce eight lanes per VMLALDAVA on Helium cores. Before it, arm-none-eabi-gcc 13 left them one scalar SMLALBB per tap. The kernel is bit-exact by construction, so no output bit moves and every battery passes on every leg. Only the M55 Q15 counts move, all improvements beyond the ratchet's tolerance, so they are re-recorded with the README tables regenerated: - bridge Q15: -55 to -62 % (up_q15_eco 44.56 M -> 19.10 M) - async Q15 pipelines: -21 % and -42 % - rational Q15: -21 to -48 % (up3 74.87 M -> 51.16 M) Every float and Q31 count, and every M33 and Hexagon count, is unchanged. rational/PLAN.md records that the M6 finding (Q15 no faster than float on the M55) is resolved: Q15 is now 19 to 32 % under float there. Co-Authored-By: Claude Opus 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
submodules/dsptapmoves to fir_kernels: Helium Q15 dot, eight lanes per VMLALDAVA DspTap#53, the Helium Q15 dot kernel (d646905).The M55 baselines are re-recorded in
async/,bridge/andrational/bench/baselines.json, and the README icount tables are regenerated withupdate_icount_docs.py. Only the M55 Q15 rows move:up_q15_ecodown_q15_ecoup_q15_sedown_q15_sepipeline12_q15pipeline_q15down2_q15_ecoup2_q15_ecoup3_q15_ecodown3_q15_ecodown2_down2_q15_ecorational/PLAN.mdgets the new M55 counts in its section 6 icount table. It also records that the M6 finding is resolved: Q15 on the M55 was no faster than float, and is now 19–32 % under it.Why
This is the first of the deferred levers: the MVE Q15 kernel. Under arm-none-eabi-gcc 13 the substrate's Q15 dot compiled to one scalar
SMLALBBper tap on Helium. The kernel now does eight lanes perVMLALDAVA. It is substrate code, so it landed in DspTap first.Verification
EveryTapCountMatchesReferencepins it at every tap count from 0 to 40.--update, then re-run with--exact, and every count matched.Notes for the reviewer
mainbefore this PR merges.🤖 Generated with Claude Code
https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Generated by Claude Code