Repository navigation
Conversation
SampleBuilder fitted input (or output) processors only when none were supplied. Supplying a pre-fitted processor for one field left every other field on that side without a processor, so its raw value (e.g. a Python list) reached the model with no error. This is the common fit-on-train, reuse-for-test pattern when a user reuses only some processors. - SampleBuilder.fit fits a processor for every schema field that has none; supplied processors are used as they are, never refitted. - SampleBuilder copies the supplied processor dicts, so fitting never adds keys to the caller's dict (e.g. the training set's processors). - tests/core/test_partial_processors.py: missing input and output processors are fitted, supplied ones are not refitted, and create_sample_dataset gives tensors for every field. - docs/api/processors.rst: "Reusing Fitted Processors". - examples/reuse_train_processors.py: fit on train, reuse for test. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Collaborator
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
SampleBuilder.fitfitted processors for a side (input or output) only when no processors were supplied for that side. Supplying a pre-fitted processor for one field therefore left every other field on that side without a processor. Its raw value, e.g. a Python list like[40.0], reached the model with no error.That's easy to hit with the "fit on train, reuse for test" pattern: reuse the code vocabulary, forget one field, and that field is silently left raw.
Change
SampleBuilder.fitfits a processor for every schema field that has none. Supplied processors are used as they are and never refitted.SampleBuildercopies the supplied processor dicts. Without the copy, filling in missing fields would also add them to the caller's dict, for example the training dataset'sinput_processors.This is PR 1 of the plan in the design note on fitting processors on the training split; it's independent of the
split=work.Tests, docs, example
New
tests/core/test_partial_processors.py:create_sample_datasetwith partial processors.Three of these fail on master.
Existing processor-transfer, ignore-processor, sample-builder and schema tests pass.
docs/api/processors.rstgets a new "Reusing Fitted Processors" section, which also documents the fit-on-train pattern.SampleBuildergets a>>>example.New
examples/reuse_train_processors.py: fit on train, reuse for test. A test-only code maps to<unk>.Full core suite:
Ran 1401 tests … OK (skipped=76).tools/check_pr_rules.pypasses.🤖 Generated with Claude Code