perf(memtrack): reduce benchmark noise and resolve workload symbols - #555
Conversation
Merging this PR will improve performance by 11.7%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | WallTime | memtrack track tar |
9.4 s | 8.4 s | +11.7% |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing cod-3690-investigate-flaky-ls-and-tar-benchmarks (407b28e) with main (c01f5d2)
Footnotes
-
6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
|
a68b18a to
fcf3921
Compare
This comment has been minimized.
This comment has been minimized.
The ls workload finishes quickly, while allocator-uprobe teardown accounts for most of the measured command time and its variation. Remove it from the walltime matrix instead of treating kernel grace-period latency as memtrack throughput.
The previous five-second budget allowed only one warmup round. An unusually fast first measured round could then determine the reported minimum. Increase warmup to thirty seconds while retaining the real archive and the existing sixty-second measurement budget.
fcf3921 to
407b28e
Compare
Changes
Missing symbols
The Ubuntu workload binaries lack full symbols, and the pinned Samply fork explicitly blocks Ubuntu debuginfod. Existing memtrack Rust names were already present. Installing local matching debug files resolves workload names and source locations before presymbolication.
All 20 hosted flamegraphs from the five executions have named tar/dd workload nodes. Tar graphs contain 181–212 workload nodes each, all named; dd graphs contain 23–35, all named. Tar's named hot frames consistently include flush_archive and sys_write_archive_buffer/flush_write. dd's leading named costs consistently include BPF CO-RE candidate matching and BTF initialization.
High-address frames with no mapped object remain unresolved (18.5–21.6% of sampled self-CPU weight in tar). Kernel-side attribution is plausible, but its cause is not established by these graphs. This change does not claim to resolve those frames.
Five-run validation
Exact revision: a2c9b21. All five tar and dd benchmark jobs passed. Four complete workflows passed; one failed only because macOS clippy could not resolve index.crates.io.
Each tar variant completed four warmup rounds and seven measured rounds. The first measured workload span was within 0.65% of the following-round median in every execution; event counts stayed at 24,902 normally and 26,403 with physical tracking. These five runs show low tar variation, not a guarantee of future stability. dd remains noisier.
CI executions: 36602936493, 36602947014, 36602957193, 36602966329, 36602976080.