Configuring Apache Hadoop to execute DQ Jobs

For large-scale processing and concurrency, a single vertically scaled Spark server may not be sufficient. To address large-scale processing, Data Quality & Observability (self-hosted) can push compute jobs to an external Hadoop cluster. This section describes how to configure the DQ Agent to push DQ jobs to Hadoop.

Important 
Dataproc (Managed Service for Apache Spark) Version 3.0 — Advisory
Dataproc 3.0 (Apache Spark 4.1) is not yet available or certified for production use. Customers may use it in non-production environments at their own risk, but we cannot provide full support until it is available.
  • Roadmap: Official support for Amazon EMR 8 (Spark 4.x) is planned for the 2026.08 release.
  • Compatibility Customers who want to test Dataproc 3.0 before it is available can use the 2026.08 DQ Spark cluster package, subject to the support limitations above.
  • Recommendation: We recommend migrating to Cloud Native (Kubernetes) to avoid version compatibility issues in future releases, as we will continue upgrading Spark versions going forward.

The following diagram shows the Data Quality & Observability (self-hosted) architecture with a Hadoop cluster: