Describe the enhancement requested
Decouple ParquetReadOptions from the legacy Hadoop ParquetInputFormat/FileInputFormat classes, which currently force pulling in the hadoop-mapreduce-client-core dependency (and its transitive JARs).
Summary
To instantiate a org.apache.parquet.hadoop.ParquetReader we need to use org.apache.parquet.ParquetReadOptions, which references a set of keys located in org.apache.parquet.hadoop.ParquetInputFormat and calls a static getFilter method declared also on ParquetInputFormat, which extends org.apache.hadoop.mapreduce.lib.input.FileInputFormat.
ParquetInputFormat is part of the parquet-hadoop module, while FileInputFormat is declared in the hadoop-mapreduce-client-core JAR from the Hadoop project.
Because ParquetReader (via ParquetReadOptions) needs to call ParquetInputFormat.getFilter(...), the JVM is forced to initialize ParquetInputFormat, and initializing a class triggers the loading and initialization of its superclass (FileInputFormat) along with its entire transitive dependency graph (org.apache.hadoop.mapreduce.*). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-mapreduce dependency into the classpath and link set. AvroParquetReader and ProtoParquetReader extend from ParquetReader and have the same issue.
ParquetInputFormat has three responsibilities:
- define a set of property keys as constants
- deserialize the filter predicates from a configuration value
- support the integration of Parquet files into Hadoop MapReduce
This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the org.apache.parquet.conf package:
ParquetInputProperties: the property-key constants only
ParquetInputFilters: the filter deserialization logic
so that the configuration can be used without ever loading FileInputFormat and its transitive dependencies.
ParquetReader ← org.apache.parquet.hadoop (parquet-hadoop)
│ uses
▼
ParquetReadOptions ← org.apache.parquet (parquet-hadoop)
│ (static import keys + getFilter call)
▼
ParquetInputFormat.getFilter(...) ← org.apache.parquet.hadoop (parquet-hadoop)
│ extends
▼
FileInputFormat<Void, T> ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ extends
▼
InputFormat<K, V> ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
│ (transitive)
▼
{ InputSplit, JobContext, TaskAttemptContext,
RecordReader, ... } ← org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)
With the new ParquetInputProperties / ParquetInputFilters, building ParquetReadOptions no longer touches any org.apache.hadoop.mapreduce type, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.
Public API / Behavioral change
No behavior change. This is a refactoring:
- New classes
org.apache.parquet.conf.ParquetInputProperties and
org.apache.parquet.conf.ParquetInputFilters (in parquet-hadoop) hold the constants and the
ParquetConfiguration-based filter resolution respectively.
- The legacy
org.apache.hadoop.conf.Configuration-based entry points remain on
ParquetInputFormat (they are used only via the legacy MapReduce path) and are kept for binary /
source compatibility.
- All pre-existing constants on
ParquetInputFormat are now @Deprecated and delegate to the new
classes; source and binary compatibility are preserved.
Acceptance criteria
ParquetReadOptions (used for plain file reads) references only ParquetInputProperties / ParquetInputFilters, never ParquetInputFormat / org.apache.hadoop.mapreduce.InputFormat.
- Loading
ParquetInputProperties / ParquetInputFilters / ParquetReadOptions does not initialize FileInputFormat.
- All existing tests still pass (
./mvnw test).
Component(s)
Core
Describe the enhancement requested
Decouple
ParquetReadOptionsfrom the legacy HadoopParquetInputFormat/FileInputFormatclasses, which currently force pulling in thehadoop-mapreduce-client-coredependency (and its transitive JARs).Summary
To instantiate a
org.apache.parquet.hadoop.ParquetReaderwe need to useorg.apache.parquet.ParquetReadOptions, which references a set of keys located inorg.apache.parquet.hadoop.ParquetInputFormatand calls a staticgetFiltermethod declared also onParquetInputFormat, whichextends org.apache.hadoop.mapreduce.lib.input.FileInputFormat.ParquetInputFormatis part of theparquet-hadoopmodule, whileFileInputFormatis declared in thehadoop-mapreduce-client-coreJAR from the Hadoop project.Because
ParquetReader(viaParquetReadOptions) needs to callParquetInputFormat.getFilter(...), the JVM is forced to initializeParquetInputFormat, and initializing a class triggers the loading and initialization of its superclass (FileInputFormat) along with its entire transitive dependency graph (org.apache.hadoop.mapreduce.*). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-mapreducedependency into the classpath and link set.AvroParquetReaderandProtoParquetReaderextend fromParquetReaderand have the same issue.ParquetInputFormathas three responsibilities:This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the
org.apache.parquet.confpackage:ParquetInputProperties: the property-key constants onlyParquetInputFilters: the filter deserialization logicso that the configuration can be used without ever loading
FileInputFormatand its transitive dependencies.With the new
ParquetInputProperties/ParquetInputFilters, buildingParquetReadOptionsno longer touches anyorg.apache.hadoop.mapreducetype, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.Public API / Behavioral change
No behavior change. This is a refactoring:
org.apache.parquet.conf.ParquetInputPropertiesandorg.apache.parquet.conf.ParquetInputFilters(inparquet-hadoop) hold the constants and theParquetConfiguration-based filter resolution respectively.org.apache.hadoop.conf.Configuration-based entry points remain onParquetInputFormat(they are used only via the legacy MapReduce path) and are kept for binary /source compatibility.
ParquetInputFormatare now@Deprecatedand delegate to the newclasses; source and binary compatibility are preserved.
Acceptance criteria
ParquetReadOptions(used for plain file reads) references onlyParquetInputProperties/ParquetInputFilters, neverParquetInputFormat/org.apache.hadoop.mapreduce.InputFormat.ParquetInputProperties/ParquetInputFilters/ParquetReadOptionsdoes not initializeFileInputFormat../mvnw test).Component(s)
Core