Migration Services
Move Hortonworks HDP data and processing workloads to Dataproc and Google Cloud through a structured migration designed to simplify operations and improve scalability.
Why Dataproc?
Reduce the effort required to maintain HDP clusters, manage Ambari, apply updates, monitor services, and resolve infrastructure issues.
Move persistent HDFS data to Cloud Storage and scale Dataproc processing independently according to workload demand.
Run persistent clusters, temporary job-specific clusters, or serverless Spark workloads based on processing requirements.
Connect Hadoop and Spark workloads with BigQuery, BigLake, Looker, Dataflow, Vertex AI, Gemini, and other Google Cloud services.
How We Migrate
Review HDP clusters, component versions, HDFS data, Spark and Hive workloads, Ambari configurations, security, metadata, reports, and dependencies.
Define the Google Cloud architecture for storage, Dataproc processing, networking, security, metadata, orchestration, analytics, and migration waves.
Move HDFS data to Cloud Storage and adapt Spark, Hive, MapReduce, and compatible Hadoop workloads for Dataproc or other suitable services.
Compare source and target data, processing results, schedules, metadata, permissions, downstream applications, and workload performance.
Transition production workloads and optimize cluster lifecycle, autoscaling, storage access, job configuration, monitoring, and cost controls.
Why Codimite?
Codimite combines Google Cloud architecture, data engineering, workload migration, and governance planning to modernize the complete HDP environment.
Start Your MigrationHDP Workload Assessment. We identify component versions, custom dependencies, outdated workloads, and services that require migration, redesign, or retirement.
Purpose-Built Target Architecture. We map storage, processing, SQL analytics, orchestration, governance, and machine learning to suitable Google Cloud services.
Security and Metadata Migration. We redesign Kerberos, Ranger, Atlas, Hive Metastore, identities, permissions, metadata, and governance for Google Cloud.
Controlled Workload Transition. We reduce risk by migrating data domains and workload groups through pilots, parallel testing, and phased production releases.
Platform Comparison
Hortonworks HDP provides a self-managed Hadoop distribution, while Dataproc delivers managed Hadoop and Spark processing within a flexible Google Cloud data ecosystem.
| Area | Hortonworks HDP | Dataproc and Google Cloud |
|---|---|---|
| Platform model | Self-managed Hadoop distribution containing multiple open-source services | Managed Hadoop and Spark processing integrated with purpose-built Google Cloud services |
| Infrastructure | Requires teams to provision, configure, maintain, and monitor cluster infrastructure | Google Cloud simplifies cluster creation, scaling, management, and retirement |
| Cluster management | Commonly managed through Apache Ambari | Managed through Google Cloud Console, APIs, CLI, automation, and workflow templates |
| Storage | Persistent data commonly stored in HDFS on the cluster | Cloud Storage provides durable storage independently from Dataproc compute |
| Processing | Spark, Hive, MapReduce, Tez, and related services run inside HDP | Dataproc supports managed Spark, Hadoop, Hive, and compatible open-source processing |
| Scaling | Scaling requires adding, configuring, and maintaining cluster nodes | Supports autoscaling, temporary clusters, and serverless Spark execution |
| Cluster lifecycle | Long-running clusters often support both data storage and processing | Clusters can be persistent, temporary, or created only for individual workloads |
| SQL analytics | Commonly uses Hive, Hive LLAP, or connected SQL engines | BigQuery provides fully managed, serverless SQL analytics for suitable workloads |
| Metadata | Uses Hive Metastore and may use Apache Atlas for metadata and lineage | Integrates with managed metadata, catalog, lineage, and governance services |
| Security | Commonly uses Kerberos, Apache Ranger, and cluster-level security controls | Uses Google Cloud IAM, private networking, encryption, audit logs, and fine-grained data controls |
| Orchestration | Uses Oozie, Ambari workflows, or external schedulers | Integrates with Cloud Composer, Workflows, Dataproc workflow templates, and scheduling services |
| Data access | Data access is closely connected to the HDP cluster and HDFS | Cloud Storage, BigQuery, and BigLake support scalable and governed access across services |
| Machine learning | Uses Spark ML and other frameworks installed within the cluster | Connects Spark processing and enterprise data with Vertex AI and Gemini |
| Monitoring | Uses Ambari Metrics, logs, and third-party monitoring tools | Integrates with Cloud Monitoring, Cloud Logging, diagnostics, and Spark History Server |
| Upgrades | Teams plan and manage component, platform, and cluster upgrades | Google manages the Dataproc service while teams select supported image versions |
| Operations | Teams manage cluster health, patching, capacity, failures, and service dependencies | Google manages service infrastructure while teams focus on data and processing workloads |
FAQs
Codimite can migrate HDFS data, Spark jobs, Hive workloads, MapReduce applications, metadata, security policies, schedules, monitoring processes, and connected analytics workflows.
No. Dataproc primarily provides managed Hadoop and Spark processing. HDP storage, SQL analytics, metadata, governance, orchestration, and machine learning capabilities may map to Cloud Storage, BigQuery, BigLake, Vertex AI, Cloud Composer, and other Google Cloud services.
Persistent HDFS data commonly moves to Cloud Storage so it remains available independently of the Dataproc cluster lifecycle. Some datasets may also move to BigQuery or BigLake based on analytics and governance requirements.
Many compatible workloads can be adapted for Dataproc. Runtime versions, libraries, file paths, custom functions, security settings, and HDP-specific dependencies must be reviewed and tested.
Ambari configurations support discovery, while cluster operations are redesigned using Google Cloud tools. Ranger policies, Atlas metadata, and Oozie workflows are mapped to Google Cloud security, governance, and orchestration services.
We compare row counts, business totals, Spark and Hive outputs, schedules, permissions, downstream results, processing performance, and operational behavior before production cutover.
Assess your HDP data, processing workloads, metadata, security, and dependencies to build a practical Google Cloud migration roadmap.
Start Your Migration