Repository navigation
[df bulk] Prepare for reading TTreeReaderValue in bulks - #22494
Draft
vepadulano wants to merge 5 commits into
Draft
vepadulano wants to merge 5 commits into
vepadulano wants to merge 5 commits into
Conversation
vepadulano
requested review from
bellenot,
dpiparo,
martamaja10 and
pcanal
as code owners
June 5, 2026 15:54
vepadulano
marked this pull request as draft
June 5, 2026 15:55
Test Results 24 files 24 suites 4d 1h 44m 57s ⏱️ For more details on these failures, see this check. Results for commit eadabf6. ♻️ This comment has been updated with latest results. |
vepadulano
force-pushed
the
df-bulk-read-bulk-ttreereadervalue
branch
from
October 5, 2026 16:01
cda277f to
9e6aa71
Compare
Sets the stage for future bulk reading where the actual values that need to be passed to downstream nodes may need to get stored in a separate location.
Nodes of the computation graph which may operate a selection on which entries are valid or not are modified to accommodate for multiple entries being evaluated at the same time. For now, the size of the bulk is set to one. Note that this commit does not touch in any way the reading of data.
The RLoopManager decides the current bulk size that all nodes of the computation graph must adhere to. Currently this is set to one, in the future it may vary.
define and variation nodes update their current available values via an Update method, this is now prepared to work in bulks, currently assuming always size one.
Enables reading TTree data via TTreeReaderValue into a pre-allocated memory region that will eventually contain the bulk of entries, assuming for now a bulk size of one. The changes are backwards compatible, all tests work with the new bulk loading logic, with one exception and a minor remark. The remark is that the test_ctors is modified since now the construction and destruction of class instances happens at the boundary of a bulk. Before, TTree could optimize by IO-constructing the value once and then re-assign it across entries. The exception is the test called PolymorphicTBranchObject in dataframe_regressions.cxx, which reads a TTree branch of on-disk type TObject where at each entry the concrete type of the value is different (namely a TList in the first entry and a TH1D in the second entry). In particular, this is a TBranchObject with split level zero. The approach taken is to disable the bulk loading from RDataFrame in this case and only allow reading one entry at a time for this type of branch. Co-authored-by: Philippe Canal <[email protected]>
vepadulano
force-pushed
the
df-bulk-read-bulk-ttreereadervalue
branch
from
October 6, 2026 11:59
9e6aa71 to
eadabf6
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Needs #22313 #22461 #22472 #22479
This is a partial implementation, serving as a starting point. It only considers the case of reading from TTreeReaderValue. It also highlights already a few pre-existing features which may not be supported anymore:
Part 5 of N of the dataframe bulk processing effort.