Skip to content

Decouple parquet-hadoop module from hadoop-mapreduce-client-core #3780

Description

@jerolba

Describe the enhancement requested

Decouple ParquetReadOptions from the legacy Hadoop ParquetInputFormat/FileInputFormat classes, which currently force pulling in the hadoop-mapreduce-client-core dependency (and its transitive JARs).

Summary

To instantiate a org.apache.parquet.hadoop.ParquetReader we need to use org.apache.parquet.ParquetReadOptions, which references a set of keys located in org.apache.parquet.hadoop.ParquetInputFormat and calls a static getFilter method declared also on ParquetInputFormat, which extends org.apache.hadoop.mapreduce.lib.input.FileInputFormat.

ParquetInputFormat is part of the parquet-hadoop module, while FileInputFormat is declared in the hadoop-mapreduce-client-core JAR from the Hadoop project.

Because ParquetReader (via ParquetReadOptions) needs to call ParquetInputFormat.getFilter(...), the JVM is forced to initialize ParquetInputFormat, and initializing a class triggers the loading and initialization of its superclass (FileInputFormat) along with its entire transitive dependency graph (org.apache.hadoop.mapreduce.*). For code that only needs to read a plain Parquet file (and not the MapReduce input-format machinery), this pulls an unwanted, heavy Hadoop-mapreduce dependency into the classpath and link set. AvroParquetReader and ProtoParquetReader extend from ParquetReader and have the same issue.

ParquetInputFormat has three responsibilities:

  • define a set of property keys as constants
  • deserialize the filter predicates from a configuration value
  • support the integration of Parquet files into Hadoop MapReduce

This issue proposes to extract the first two responsibilities into two new Hadoop-agnostic classes, in the org.apache.parquet.conf package:

  • ParquetInputProperties: the property-key constants only
  • ParquetInputFilters: the filter deserialization logic
    so that the configuration can be used without ever loading FileInputFormat and its transitive dependencies.
ParquetReader                                   ← org.apache.parquet.hadoop (parquet-hadoop)
        │  uses
        ▼
ParquetReadOptions                              ← org.apache.parquet (parquet-hadoop)
        │  (static import keys + getFilter call)
        ▼
ParquetInputFormat.getFilter(...)               ← org.apache.parquet.hadoop (parquet-hadoop)
        │  extends
        ▼
FileInputFormat<Void, T>                        ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
        │  extends
        ▼
InputFormat<K, V>                               ← org.apache.hadoop.mapreduce.lib.input (hadoop-mapreduce-client-core)
        │  (transitive)
        ▼
{ InputSplit, JobContext, TaskAttemptContext,
  RecordReader, ... }                           ← org.apache.hadoop.mapreduce* (hadoop-mapreduce-client-core)

With the new ParquetInputProperties / ParquetInputFilters, building ParquetReadOptions no longer touches any org.apache.hadoop.mapreduce type, so consumers not related to Hadoop avoid transitively including the MapReduce dependency.

Public API / Behavioral change

No behavior change. This is a refactoring:

  • New classes org.apache.parquet.conf.ParquetInputProperties and
    org.apache.parquet.conf.ParquetInputFilters (in parquet-hadoop) hold the constants and the
    ParquetConfiguration-based filter resolution respectively.
  • The legacy org.apache.hadoop.conf.Configuration-based entry points remain on
    ParquetInputFormat (they are used only via the legacy MapReduce path) and are kept for binary /
    source compatibility.
  • All pre-existing constants on ParquetInputFormat are now @Deprecated and delegate to the new
    classes; source and binary compatibility are preserved.

Acceptance criteria

  • ParquetReadOptions (used for plain file reads) references only ParquetInputProperties / ParquetInputFilters, never ParquetInputFormat / org.apache.hadoop.mapreduce.InputFormat.
  • Loading ParquetInputProperties / ParquetInputFilters / ParquetReadOptions does not initialize FileInputFormat.
  • All existing tests still pass (./mvnw test).

Component(s)

Core

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions