Skip to content

[df bulk] Prepare for reading TTreeReaderValue in bulks - #22494

Draft
vepadulano wants to merge 5 commits into
root-project:masterfrom
vepadulano:df-bulk-read-bulk-ttreereadervalue
Draft

vepadulano wants to merge 5 commits into
root-project:masterfrom
vepadulano:df-bulk-read-bulk-ttreereadervalue

Conversation

@vepadulano

@vepadulano vepadulano commented Jun 5, 2026 •

Copy link
Copy Markdown
Member

Needs #22313 #22461 #22472 #22479

This is a partial implementation, serving as a starting point. It only considers the case of reading from TTreeReaderValue. It also highlights already a few pre-existing features which may not be supported anymore:

  • Reading polymorphic TBranch with different concrete types stored across different events
  • Reading polymorphic TBranch of concrete types read as their base type

Part 5 of N of the dataframe bulk processing effort.

@vepadulano vepadulano self-assigned this Jun 5, 2026
@vepadulano
vepadulano marked this pull request as draft June 5, 2026 15:55
@github-actions

github-actions Bot commented Jun 5, 2026 •

Copy link
Copy Markdown

Test Results

    24 files      24 suites   4d 1h 44m 57s ⏱️
 3 883 tests  3 880 ✅ 0 💤 3 ❌
83 112 runs  83 108 ✅ 0 💤 4 ❌

For more details on these failures, see this check.

Results for commit eadabf6.

♻️ This comment has been updated with latest results.

@vepadulano
vepadulano force-pushed the df-bulk-read-bulk-ttreereadervalue branch from cda277f to 9e6aa71 Compare October 5, 2026 16:01
vepadulano and others added 5 commits October 6, 2026 09:18
Sets the stage for future bulk reading where the actual values that need to be passed to downstream nodes may need to get stored in a separate location.
Nodes of the computation graph which may operate a selection on which entries are valid or not are modified to accommodate for multiple entries being evaluated at the same time. For now, the size of the bulk is set to one.

Note that this commit does not touch in any way the reading of data.
The RLoopManager decides the current bulk size that all nodes of the computation graph must adhere to. Currently this is set to one, in the future it may vary.
define and variation nodes update their current available values via an Update method, this is now prepared to work in bulks, currently assuming always size one.
Enables reading TTree data via TTreeReaderValue into a pre-allocated memory region that will eventually contain the bulk of entries, assuming for now a bulk size of one. The changes are backwards compatible, all tests work with the new bulk loading logic, with one exception and a minor remark.

The remark is that the test_ctors is modified since now the construction and destruction of class instances happens at the boundary of a bulk. Before, TTree could optimize by IO-constructing the value once and then re-assign it across entries.

The exception is the test called PolymorphicTBranchObject in dataframe_regressions.cxx, which reads a TTree branch of on-disk type TObject where at each entry the concrete type of the value is different (namely a TList in the first entry and a TH1D in the second entry). In particular, this is a TBranchObject with split level zero. The approach taken is to disable the bulk loading from RDataFrame in this case and only allow reading one entry at a time for this type of branch.

Co-authored-by: Philippe Canal <[email protected]>
@vepadulano
vepadulano force-pushed the df-bulk-read-bulk-ttreereadervalue branch from 9e6aa71 to eadabf6 Compare October 6, 2026 11:59

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant