Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions parquet-hadoop/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -403,6 +403,17 @@ ParquetInputFormat to materialize records. It should be a the descendant class o
**Description:** The number of threads will be spawned to compress in parallel. More workers improve speed, but also increase memory usage. When it is 0, it works as single-threaded mode.
**Default value:** `0`

## Class: ParquetReadOptions

**Property:** `parquet.read.allocation.size`

**Description:** Size in bytes used to split planned column-chunk ranges into read buffers.
For vectored IO, this limits the length of each range requested from the filesystem.
Filesystems may merge or align ranges for checksums and allocate larger buffers.
This setting does not limit other allocations, such as footer reads or decoding, or total memory use.

**Default value:** `8388608` (8 MiB)

## Class: HadoopReadOptions

**Property:** `parquet.crypto.factory.class`
Expand Down Expand Up @@ -547,6 +558,26 @@ if not found then the library will use the classic non-vectored reads: it is saf

---

**JVM system property:** `parquet.hadoop.vectored.io.threads`

**Description:** Maximum number of shared vectored-read workers. Set a positive integer

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This executor is a JVM-global singleton read once at class initialization, so the limit is process-wide and cannot be changed per reader or per query.

Would it make sense to also let the caller supply the executor, with this shared pool as the default? That way an engine could size and share it in the way it already manages its own pools.

with `-Dparquet.hadoop.vectored.io.threads=N` before the first vectored read; this is
read once per JVM, not from the Hadoop configuration. Invalid values fail initialization.
All readers share this limit. Waiting for a worker, submitting the read, and waiting
for its results share the same 300-second deadline. If admission times out, the read
fails before submitting filesystem work and the reader retains its stream.

A worker remains occupied until submission returns, the caller transfers or abandons
the buffers, and any failure cleanup closes the stream. A backend that ignores
interruption therefore retains capacity; repeated timeouts cannot create workers
beyond the configured limit. If all workers are blocked, further vectored reads time
out waiting for capacity. Idle workers expire after 60 seconds. This limit does not
bound threads created by the filesystem or total buffer memory.

**Default value:** `64`

---

**Property:** `parquet.column.statistics.enabled`
**Description:** Whether to enable column statistics collection.
If `true`, statistics will be collected for all columns unless explicitly disabled for specific columns.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -211,6 +211,14 @@ public ByteBufferAllocator getAllocator() {
return allocator;
}

/**
* Returns the size in bytes used to split planned column-chunk ranges into read buffers.
* For vectored IO, this bounds each requested range length; filesystems may merge or
* checksum-align ranges and allocate larger buffers. This does not limit other allocations,
* such as footer reads or decoding, or total memory use.
*
* @return the configured read allocation size in bytes
*/
public int getMaxAllocationSize() {
return maxAllocationSize;
}
Expand Down Expand Up @@ -405,6 +413,12 @@ public Builder withAllocator(ByteBufferAllocator allocator) {
return this;
}

/**
* Sets the read allocation size described by {@link ParquetReadOptions#getMaxAllocationSize()}.
*
* @param allocationSizeInBytes the maximum column-chunk read buffer or requested vectored range size in bytes
* @return this builder
*/
public Builder withMaxAllocationInBytes(int allocationSizeInBytes) {
this.maxAllocationSize = allocationSizeInBytes;
return this;
Expand Down
Loading
Loading