Skip to content

perf(memtrack): reduce benchmark noise and resolve workload symbols - #555

Merged
not-matthias merged 2 commits into
mainfrom
cod-3690-investigate-flaky-ls-and-tar-benchmarks
Oct 1, 2026
Merged

not-matthias merged 2 commits into
mainfrom
cod-3690-investigate-flaky-ls-and-tar-benchmarks

Conversation

@not-matthias

@not-matthias not-matthias commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Changes

  • Remove the short ls benchmark, whose measured time and variation are dominated by allocator-uprobe teardown.
  • Increase tar warmup from 5s to 30s while retaining real archive output and the 60s measurement budget.
  • Install exact-version, architecture-matched tar/coreutils debug packages from Launchpad over HTTPS before profiling. Do not upgrade workload binaries or depend on mutable ddebs indexes.

Missing symbols

The Ubuntu workload binaries lack full symbols, and the pinned Samply fork explicitly blocks Ubuntu debuginfod. Existing memtrack Rust names were already present. Installing local matching debug files resolves workload names and source locations before presymbolication.

All 20 hosted flamegraphs from the five executions have named tar/dd workload nodes. Tar graphs contain 181–212 workload nodes each, all named; dd graphs contain 23–35, all named. Tar's named hot frames consistently include flush_archive and sys_write_archive_buffer/flush_write. dd's leading named costs consistently include BPF CO-RE candidate matching and BTF initialization.

High-address frames with no mapped object remain unresolved (18.5–21.6% of sampled self-CPU weight in tar). Kernel-side attribution is plausible, but its cause is not established by these graphs. This change does not claim to resolve those frames.

Five-run validation

Exact revision: a2c9b21. All five tar and dd benchmark jobs passed. Four complete workflows passed; one failed only because macOS clippy could not resolve index.crates.io.

Benchmark Range (seconds) Sample CV
tar 9.299–9.403 0.45%
tar, physical 9.427–9.497 0.30%
dd 1.139–1.251 3.67%
dd, physical 1.228–1.326 2.92%

Each tar variant completed four warmup rounds and seven measured rounds. The first measured workload span was within 0.65% of the following-round median in every execution; event counts stayed at 24,902 normally and 26,403 with physical tracking. These five runs show low tar variation, not a guarantee of future stability. dd remains noisier.

CI executions: 36602936493, 36602947014, 36602957193, 36602966329, 36602976080.

  • Relevant prek hooks passed.
  • Ubuntu 22.04 container smoke verified unchanged tar/dd executable checksums and build-ID-matched debug files containing symbols and DWARF.
  • An earlier debug-package approach hit inconsistent Ubuntu ARM64 mirror indexes and implicitly upgraded workload packages. Those runs are excluded from the statistics above; exact-version Launchpad downloads avoid both problems.

@codspeed

codspeed Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Merging this PR will improve performance by 11.7%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚡ 1 improved benchmark
✅ 30 untouched benchmarks
⏩ 6 skipped benchmarks1

Performance Changes

Mode Benchmark BASE HEAD Efficiency
⚡ WallTime memtrack track tar 9.4 s 8.4 s +11.7%

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing cod-3690-investigate-flaky-ls-and-tar-benchmarks (407b28e) with main (c01f5d2)

Open in CodSpeed

Footnotes

  1. 6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@not-matthias
not-matthias marked this pull request as ready for review September 30, 2026 07:57
@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Low risk] Adjusts benchmark configuration and removes a benchmark workload.

The PR appears safe to merge; no outstanding or new actionable finding remains.

Summary

The PR removes the teardown-dominated ls benchmark and increases tar warmup from 5 to 30 seconds while retaining the 60-second measurement budget.

  • CI continues to run the dd and tar workload configurations.
  • No new actionable issue was identified.

Reviews (3) · Last reviewed commit: "perf(memtrack): warm up tar for several ..."

Comment thread .github/workflows/ci.yml Outdated
Comment thread .github/workflows/ci.yml Outdated
@not-matthias
not-matthias removed the request for review from GuillaumeLagrange September 30, 2026 08:02
@not-matthias
not-matthias force-pushed the cod-3690-investigate-flaky-ls-and-tar-benchmarks branch from a68b18a to fcf3921 Compare September 30, 2026 08:58
@greptile-apps

This comment has been minimized.

The ls workload finishes quickly, while allocator-uprobe teardown accounts for most of the measured command time and its variation. Remove it from the walltime matrix instead of treating kernel grace-period latency as memtrack throughput.
The previous five-second budget allowed only one warmup round. An unusually fast first measured round could then determine the reported minimum. Increase warmup to thirty seconds while retaining the real archive and the existing sixty-second measurement budget.
@not-matthias
not-matthias force-pushed the cod-3690-investigate-flaky-ls-and-tar-benchmarks branch from fcf3921 to 407b28e Compare October 1, 2026 09:55
@not-matthias
not-matthias merged commit 407b28e into main Oct 1, 2026
56 checks passed
@not-matthias
not-matthias deleted the cod-3690-investigate-flaky-ls-and-tar-benchmarks branch October 1, 2026 10:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants