Skip to content

Remove pyarrow as required dependency, relying on Arrow PyCapsule Interface #1227

Description

@kylebarron

Is your feature request related to a problem or challenge? Please describe what you are trying to do.

PyArrow is a massive dependency. Unpacked, it tends to be >100MB in size, and, until the latest versions (I think?) also required numpy as its own non-optional dependency.

It's also, in effect the only current dependency

dependencies = ["pyarrow>=11.0.0", "typing-extensions;python_version<'3.13'"]

It would be great if we could remove it, and that would greatly lessen the minimal environment size for datafusion python.

Many other Python Arrow libraries implement the PyCapsule Interface, so the user can use nanoarrow, arro3, Polars, DuckDB, etc, or pyarrow. Whatever is best for them.

Describe the solution you'd like

The Arrow PyCapsule Interface is a lightweight, decentralized protocol for sharing Arrow data between Python libraries. We already implement the PyCapsule Interface, so it's just a matter of removing places where we hard-code use of pyarrow.

Describe alternatives you've considered

Keep pyarrow dependency.

Additional context

Activity

  1. ion-elgreco commented on Sep 17, 2025

    @ion-elgreco
    Contributor

    I agree with removing pyarrow as a dependency. We have to consider that there are users out there that want to create small containers but do data operations in them, if you pull in pyarrow which is so heavyweight that complicates it a lot.

    In delta-rs I've already done this, delta-io/delta-rs#3459 some time ago, which was very easy due to Arro3. (Thanks @kylebarron :D)

    As I am thinking about it, @kylebarron a pointer might make more sense indeed since it's more to be interopable between other libs, if you want to analyze the data it can be done directly in datafusion.

  2. kylebarron commented on Sep 17, 2025

    @kylebarron
    MemberAuthor

    We already have implementations of the PyCapsule interface, and we're improving it in #1222 (thanks to @kosiew ). It's just methods like to_batches that are just syntactic sugar for calling pa.table(datafusion_df). What should we do about them?

  3. timsaucer commented on Sep 17, 2025

    @timsaucer
    Member

    The more I think about it, I think this is a good idea. We just need someone who wants to dive in and take on the work. I have too many other things in the air right now to tackle myself.

  4. ion-elgreco commented on Sep 17, 2025

    @ion-elgreco
    Contributor

    Perhaps @H0TB0X420 would like to move this forward?

  5. HairstonE commented on Sep 19, 2025

    @HairstonE
    Contributor

    Sure! I can work on moving this forward. Just to confirm, we would like to remove places where pyarrow is explicitly used and implement the PyCapsule Interface?

  6. kylebarron commented on Sep 19, 2025

    @kylebarron
    MemberAuthor

    First, let's establish a list of all the places where pyarrow is used. Where is it used as an input parameter and where is it used as a return type?

    Wherever pyarrow is used as an input parameter is pretty trivially replaced by the PyCapsule Interface, because the interface can accept pyarrow objects without us needing to depend on pyarrow as a dependency.

    Wherever it's used as a return type is a little more tricky, because we need to return something.

  7. ion-elgreco commented on Sep 19, 2025

    @ion-elgreco
    Contributor

    First, let's establish a list of all the places where pyarrow is used. Where is it used as an input parameter and where is it used as a return type?

    Wherever pyarrow is used as an input parameter is pretty trivially replaced by the PyCapsule Interface, because the interface can accept pyarrow objects without us needing to depend on pyarrow as a dependency.

    Wherever it's used as a return type is a little more tricky, because we need to return something.

    Where we return a pyarrow type we either need to remove it and implement on the Datafusion type directly, for example for the schema we just need to have it be a Datafusion.schema that's how I mentioned it in the initial issue when I had time :p

  8. HairstonE commented on Sep 23, 2025

    @HairstonE
    Contributor

    I'll pick this up this week. Apologies for the delay, I've been traveling.

  9. HairstonE commented on Sep 24, 2025

    @HairstonE
    Contributor

    I used these commands to find where pyarrow is used for input parameters and return types.

    grep -r "import pyarrow" -> pyarrow_imports.txt: All files importing pyarrow
    grep -r "pa\." -> pyarrow_usage.txt: All usage of pa. objects in the codebase

    Input Parameters

    SessionContext methods:

    • from_arrow_table(data: pa.Table)
    • create_dataframe(partitions: list[list[pa.RecordBatch]])
    • register_csv(..., schema: pa.Schema)
    • register_parquet(..., schema: pa.Schema)
    • register_json(..., schema: pa.Schema)
    • register_dataset(dataset: pa.dataset.Dataset)

    DataFrame methods:

    • cast(mapping: dict[str, pa.DataType])

    User-defined functions:

    • Various methods accepting pa.DataType, pa.Array, pa.Scalar

    Return Types

    Core API methods:

    • DataFrame.collect() -> List[pa.RecordBatch]
    • DataFrame.collect_partitioned() -> List[List[pa.RecordBatch]]
    • DataFrame.schema() -> pa.Schema
    • DataFrame.to_arrow_table() -> pa.Table
    • RecordBatch.to_pyarrow() -> pa.RecordBatch
    • DataFrame.into_view() -> pa.Table

    Convenience methods:

    • DataFrame.to_pandas()
    • DataFrame.to_polars()

    What's the recommended approach for tackling this?

  10. kylebarron commented on Sep 25, 2025

    @kylebarron
    MemberAuthor

    I think it'll be easiest to focus on input parameters first. You can either use the arrow-rs pyarrow feature flag or use my crate pyo3-arrow. I'm partial to pyo3-arrow because of these issues. So, for example, if we ever have a place where we need a column input then we need to use pyo3-arrow (arrow doesn't have a way to import a column).

    Input Parameters

    SessionContext methods:

    • from_arrow_table(data: pa.Table)

    This can be updated to import a table via the PyCapsule Interface, without any breaking change. On the Python side the type hint can be updated to

    from typing import Tuple, Protocol
    
    class ArrowStreamExportable(Protocol):
        def __arrow_c_stream__(
            self,
            requested_schema: object | None = None
        ) -> object:
            ...
    • create_dataframe(partitions: list[list[pa.RecordBatch]])

    Similarly, each RecordBatch can be imported via pycapsules.

    • register_csv(..., schema: pa.Schema)
    • register_parquet(..., schema: pa.Schema)
    • register_json(..., schema: pa.Schema)

    These can import a schema via the pycapsule interface, without any breaking change. On the python side the type hint can be updated to this protocol

    from typing import Tuple, Protocol
    
    class ArrowSchemaExportable(Protocol):
        def __arrow_c_schema__(self) -> object: ...
    • register_dataset(dataset: pa.dataset.Dataset)

    This is specific to a pyarrow API, and so this can stay as-is with pyarrow as an optional dependency (I wish this was named register_pyarrow_dataset to avoid confusion)

    DataFrame methods:

    • cast(mapping: dict[str, pa.DataType])

    Why isn't this a schema? Or is intentional that you only want to cast a couple specified columns, leaving the others alone, not "projecting" to the specific schema?

  11. HairstonE commented on Sep 26, 2025

    @HairstonE
    Contributor

    I'm hitting a compatibility issue implementing PyCapsule support with pyo3-arrow.

    The issue is pyo3-arrow requires the buffer protocol, which isn't available in pyo3's stable ABI (abi3) for Python <3.11. datafusion-python currently uses abi3-py39.

    The error: buffer module is gated behind #![cfg(any(not(Py_LIMITED_API), Py_3_11))]

    How would you recommend tackling this?

  12. kylebarron commented on Sep 26, 2025

    @kylebarron
    MemberAuthor

    You can turn off the default features of pyo3-arrow

  13. HairstonE commented on Sep 28, 2025

    @HairstonE
    Contributor

    I tested removing the PyArrow dependency but ran into blocking issues:

    Keep arrow pyarrow: Compiles successfully but doesn't remove PyArrow dependency, so it doesn't solve the problem of making PyArrow optional.

    pyo3-arrow: Hit multiple blockers including requiring updates to 15 or more files importing from arrow::pyarrow, pyo3-arrow conflicts with abi3 for Python <3.11, version incompatibility (arrow pyarrow needs pyo3 0.25, pyo3-arrow needs 0.26), and with default-features = false, pyo3-arrow doesn't export PyArrowType at all (only has buffer_protocol feature).

    I'm a bit stumped and could use some direction on how to proceed.

  14. kylebarron commented on Sep 30, 2025

    @kylebarron
    MemberAuthor

    Keep arrow pyarrow: Compiles successfully but doesn't remove PyArrow dependency, so it doesn't solve the problem of making PyArrow optional.

    The pyarrow feature of Rust arrow doesn't itself require PyArrow at runtime for data import IIRC.

    pyo3-arrow: Hit multiple blockers including requiring updates to 15 or more files importing from arrow::pyarrow

    If maintainers are interested in switching to pyo3-arrow, we could have a first PR making this switch.

    pyo3-arrow conflicts with abi3 for Python <3.11

    You can turn off the buffer protocol support by passing --no-default-features

    version incompatibility (arrow pyarrow needs pyo3 0.25, pyo3-arrow needs 0.26)

    There's a table with version compatibility. You need to match the version of arrow that DataFusion is using.

    and with default-features = false, pyo3-arrow doesn't export PyArrowType at all

    I don't know what you mean by PyArrowType. Do you mean for the data type? There's nothing called PyArrowType in pyo3-arrow or arrow. You can use PyDataType. That doesn't depend on the buffer protocol feature.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions