feat: add support for datasets in training - #387
stephantul wants to merge 6 commits into
Conversation
|
Codecov Report❌ Patch coverage is
🚀 New features to boost your workflow:
|
Pringled
left a comment
There was a problem hiding this comment.
Very very nice! I ran some datasets and everything looks good. I only have one architectural suggestion to keep everything related to pyarrow in dataset.py, but up to you if it makes sense
| pa.types.is_list(label_type) or pa.types.is_large_list(label_type) or pa.types.is_fixed_size_list(label_type) | ||
| ) | ||
| value_type = label_type.value_type if multilabel else label_type | ||
| if not (pa.types.is_string(value_type) or pa.types.is_large_string(value_type) or pa.types.is_integer(value_type)): |
There was a problem hiding this comment.
It's a bit hard for me to follow the flow of everything but I think we are checking this several times, is it possible to check this once in dataset.py and then just assume it's correct downstream?
This PR adds datasets to all our training paths.