ParquetIO (Apache Beam 2.27.0)

java.lang.Object
- org.apache.beam.sdk.io.parquet.ParquetIO

```
@Experimental(value=SOURCE_SINK)
public class ParquetIO
extends java.lang.Object
```
IO to read and write Parquet files.
Reading Parquet files

ParquetIO source returns a PCollection for Parquet files. The elements in the PCollection are Avro GenericRecord.
To configure the ParquetIO.Read, you have to provide the file patterns (from) of the Parquet files and the schema.
For example:
```
 PCollection<GenericRecord> records = pipeline.apply(ParquetIO.read(SCHEMA).from("/foo/bar"));
 ...
 
```
As ParquetIO.Read is based on FileIO, it supports any filesystem (hdfs, ...).
When using schemas created via reflection, it may be useful to generate GenericRecord instances rather than instances of the class associated with the schema. ParquetIO.Read and ParquetIO.ReadFiles provide ParquetIO.Read.withAvroDataModel(GenericData) allowing implementations to set the data model associated with the AvroParquetReader
For more advanced use cases, like reading each file in a PCollection of FileIO.ReadableFile, use the ParquetIO.ReadFiles transform.
For example:
```
 PCollection<FileIO.ReadableFile> files = pipeline
   .apply(FileIO.match().filepattern(options.getInputFilepattern())
   .apply(FileIO.readMatches());

 PCollection<GenericRecord> output = files.apply(ParquetIO.readFiles(SCHEMA));
 
```
Splittable reading can be enabled by allowing the use of Splittable DoFn. It initially split the files into blocks of 64MB and may dynamically split further for higher read efficiency. It can be enabled by using ParquetIO.Read.withSplit().
For example:
```
 PCollection<GenericRecord> records = pipeline.apply(ParquetIO.read(SCHEMA).from("/foo/bar").withSplit());
 ...
 
```
Reading with projection can be enabled with the projection schema as following. Splittable reading is enabled when reading with projection. The projection_schema contains only the column that we would like to read and encoder_schema contains the schema to encode the output with the unwanted columns changed to nullable. Partial reading provide decrease of reading time due to partial processing of the data and partial encoding. The decrease in the reading time depends on the relative position of the columns. Memory allocation is optimised depending on the encoding schema. Note that the improvement is not as significant comparing to the proportion of the data requested, since the processing time saved is only the time to read the unwanted columns, the reader will still go over the data set according to the encoding schema since data for each column in a row is stored interleaved.
```
 * PCollection<GenericRecord> records = pipeline.apply(ParquetIO.read(SCHEMA).from("/foo/bar").withProjection(Projection_schema,Encoder_Schema));
 * ...
 *
 
```
Writing Parquet files

ParquetIO.Sink allows you to write a PCollection of GenericRecord into a Parquet file. It can be used with the general-purpose FileIO transforms with FileIO.write/writeDynamic specifically.
By default, ParquetIO.Sink produces output files that are compressed using the CompressionCodec.SNAPPY. This default can be changed or overridden using ParquetIO.Sink.withCompressionCodec(CompressionCodecName).
For example:
```
 pipeline
   .apply(...) // PCollection<GenericRecord>
   .apply(FileIO
     .<GenericRecord>write()
     .via(ParquetIO.sink(SCHEMA)
       .withCompressionCodec(CompressionCodecName.SNAPPY))
     .to("destination/path")
     .withSuffix(".parquet"));
 
```
This IO API is considered experimental and may break or receive backwards-incompatible changes in future versions of the Apache Beam SDK.
See Also:

Beam ParquetIO documentation

Nested Class Summary

Nested Classes
Modifier and Type	Class and Description
`static class`	`ParquetIO.Read` Implementation of `read(Schema)`.
`static class`	`ParquetIO.ReadFiles` Implementation of `readFiles(Schema)`.
`static class`	`ParquetIO.Sink` Implementation of `sink(org.apache.avro.Schema)`.

Method Summary

All Methods Static Methods Concrete Methods
Modifier and Type	Method and Description
`static ParquetIO.Read`	`read(Schema schema)` Reads `GenericRecord` from a Parquet file (or multiple Parquet files matching the pattern).
`static ParquetIO.ReadFiles`	`readFiles(Schema schema)` Like `read(Schema)`, but reads each file in a `PCollection` of `FileIO.ReadableFile`, which allows more flexible usage.
`static ParquetIO.Sink`	`sink(Schema schema)` Creates a `ParquetIO.Sink` that, for use with `FileIO.write()`.

Methods inherited from class java.lang.Object
clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait

- Method Detail
  - read
```
public static ParquetIO.Read read(Schema schema)
```
    Reads GenericRecord from a Parquet file (or multiple Parquet files matching the pattern).
  - readFiles
```
public static ParquetIO.ReadFiles readFiles(Schema schema)
```
    Like read(Schema), but reads each file in a PCollection of FileIO.ReadableFile, which allows more flexible usage.
  - sink
```
public static ParquetIO.Sink sink(Schema schema)
```
    Creates a ParquetIO.Sink that, for use with FileIO.write().

Class ParquetIO

Reading Parquet files

Writing Parquet files

Nested Class Summary

Method Summary

Methods inherited from class java.lang.Object

Method Detail

read

readFiles

sink