Skip to content

Preserve Python and Node symbols in memory flamegraphs - #556

Draft
not-matthias wants to merge 4 commits into
mainfrom
cod-3654-support-memory-flamegraphs-for-pythonnode
Draft

not-matthias wants to merge 4 commits into
mainfrom
cod-3654-support-memory-flamegraphs-for-pythonnode

Conversation

@not-matthias

Copy link
Copy Markdown
Member

Summary

  • Capture per-process perf maps and Python JIT unwind data alongside native memory artifacts.
  • Enable V8 perf maps for memory runs while preserving existing NODE_OPTIONS.
  • Document native-allocation coverage and runtime limitations.

Verification

  • cargo test --release --bin codspeed writes_keyed_artifacts_and_metadata_for_a_streamed_mapping
  • cargo fmt --all --check
  • Local Python 3.12 perf-trampoline and Node 22 memory runs produced perf maps and resolved language frames when parsed offline.

@codspeed

codspeed Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

✅ 31 untouched benchmarks
⏩ 6 skipped benchmarks1


Comparing cod-3654-support-memory-flamegraphs-for-pythonnode (9108e6c) with main (214c040)

Open in CodSpeed

Footnotes

  1. 6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

Python and Node write runtime symbols to /tmp/perf-<pid>.map, and Python JIT dumps carry the unwind data needed to walk through interpreter trampolines. Collect both for the benchmark processes before saving the memtrack metadata, reusing the walltime artifact pipeline, so offline allocation stacks keep their runtime frames.
Memory stacks can only name Python and Node frames from the runtime perf maps. Set PYTHONPERFSUPPORT and the Node perf options per benchmark command, as simulation mode already does. Memory mode also passes --interpreted-frames-native-stack, since interpreted JS frames otherwise all resolve to the shared V8 interpreter trampoline.
Track a native allocation made from a Python function (perf trampoline) and a Node function (V8 perf-basic-prof). Assert the allocation carries a captured stack and that the runtime perf map names the allocating function, which offline attribution needs.
@not-matthias
not-matthias force-pushed the cod-3654-support-memory-flamegraphs-for-pythonnode branch from 5c6e4f0 to 62fd7de Compare October 2, 2026 15:19
A process can map the same file at several addresses at once. V8 remaps its
embedded builtins out of the node binary into its code range, so node runs
from both its original text and that copy. Each placement has its own load
bias, but only the last mapping per (path, pid) was kept, so frames in every
earlier placement lost their symbols and unwind data. In a Node memory
profile, all native node frames showed up as unresolved addresses.

Record every distinct load bias and the unwind data of every executable
mapping for each process, and emit them all in the artifact metadata. The
walltime perf path shares this bookkeeping and gets the same fix.

Add a sample that maps its own text a second time and allocates through
both copies. The test feeds that process's real mappings into the artifact
pipeline and checks that both copies resolve.

Refs COD-1377
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant