Skip to content

Fit processors for schema fields that were not supplied - #1266

Closed
solarsys wants to merge 1 commit into
sunlabuiuc:masterfrom
solarsys:fix/fit-missing-processors
Closed

solarsys wants to merge 1 commit into
sunlabuiuc:masterfrom
solarsys:fix/fit-missing-processors

Conversation

@solarsys

@solarsys solarsys commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Problem

SampleBuilder.fit fitted processors for a side (input or output) only when no processors were supplied for that side. Supplying a pre-fitted processor for one field therefore left every other field on that side without a processor. Its raw value, e.g. a Python list like [40.0], reached the model with no error.

That's easy to hit with the "fit on train, reuse for test" pattern: reuse the code vocabulary, forget one field, and that field is silently left raw.

create_sample_dataset(samples, {"codes": "sequence", "age": "tensor"}, {"label": "binary"},
                      input_processors={"codes": train_codes})
# master: dataset[0]["age"] == [40.0]   (a list, no processor)

Change

  • SampleBuilder.fit fits a processor for every schema field that has none. Supplied processors are used as they are and never refitted.
  • SampleBuilder copies the supplied processor dicts. Without the copy, filling in missing fields would also add them to the caller's dict, for example the training dataset's input_processors.

This is PR 1 of the plan in the design note on fitting processors on the training split; it's independent of the split= work.

Tests, docs, example

  • New tests/core/test_partial_processors.py:

    • a missing input processor is fitted and gives tensors;
    • a missing output processor is fitted;
    • a supplied processor isn't refitted (its vocabulary doesn't grow);
    • create_sample_dataset with partial processors.

    Three of these fail on master.

  • Existing processor-transfer, ignore-processor, sample-builder and schema tests pass.

  • docs/api/processors.rst gets a new "Reusing Fitted Processors" section, which also documents the fit-on-train pattern.

  • SampleBuilder gets a >>> example.

  • New examples/reuse_train_processors.py: fit on train, reuse for test. A test-only code maps to <unk>.

  • Full core suite: Ran 1401 tests … OK (skipped=76). tools/check_pr_rules.py passes.

🤖 Generated with Claude Code

SampleBuilder fitted input (or output) processors only when none were
supplied. Supplying a pre-fitted processor for one field left every other
field on that side without a processor, so its raw value (e.g. a Python
list) reached the model with no error. This is the common fit-on-train,
reuse-for-test pattern when a user reuses only some processors.

- SampleBuilder.fit fits a processor for every schema field that has
  none; supplied processors are used as they are, never refitted.
- SampleBuilder copies the supplied processor dicts, so fitting never
  adds keys to the caller's dict (e.g. the training set's processors).
- tests/core/test_partial_processors.py: missing input and output
  processors are fitted, supplied ones are not refitted, and
  create_sample_dataset gives tensors for every field.
- docs/api/processors.rst: "Reusing Fitted Processors".
- examples/reuse_train_processors.py: fit on train, reuse for test.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@solarsys

solarsys commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator Author

Closing as already merged: this PR's change (fitting processors for schema fields that were not supplied, plus tests/core/test_partial_processors.py) landed on master as part of #1273, which was stacked on it and squash-merged in e94ca1d. Nothing left to merge here.

@solarsys solarsys closed this Oct 7, 2026
@solarsys
solarsys deleted the fix/fit-missing-processors branch October 9, 2026 05:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant