[Remote] Data Engineer-Data Platforms-Google - 1
Note: The job is a remote job and is open to candidates in USA. IBM Consulting helps companies advance their hybrid cloud and AI journeys through collaboration, technology, and innovative solutions. The Data Engineer will design, build, maintain, and optimize scalable batch and real-time data pipelines and data engineering solutions using Google Cloud data platforms and open-source technologies.
Responsibilities
- Design Data Pipelines: Design and develop batch and real-time data pipelines for Data Warehouse and Datalake using Google Cloud services such as DataProc, DataFlow, PubSub, BigQuery, and Big Table
- Develop Data Engineering Solutions: Build and maintain data engineering solutions using Google Cloud Storage, BigTable, BigQuery DataProc with Spark and Hadoop, Google DataFlow with Apache Beam or Python, and other open-source technologies like Apache Airflow, dbt, Spark/Python, or Spark/Scala
- Manage Data Platforms: Schedule and manage the data platform using Google Cloud Scheduler and Cloud Composer (Airflow), ensuring seamless data pipeline operations
- Optimize Data Layer: Design and optimize the data layer for efficient data migration and data processing using Google Cloud services
- Ensure Scalability: Ensure scalability and efficiency of data pipelines and data engineering solutions to meet business needs
Skills
- Proven experience designing, building, and maintaining data engineering solutions on Google's Cloud ecosystem, including Google DataProc, DataFlow, PubSub, BigQuery, Big Table, Cloud Spanner, CloudSQL, and AlloyDB
- Experience with Apache Beam, Apache Airflow, dbt, Spark/Python, or Spark/Scala, and ability to integrate these technologies with Google Cloud services
- Experience developing and managing batch and real-time data pipelines for Data Warehouse and Datalake using Google Cloud services
- Experience scheduling and managing data platforms using Google Cloud Scheduler and Cloud Composer (Airflow)
- Experience designing and optimizing data layers for efficient data migration and data processing using Google Cloud services
- Experience with Apache Beam, including integrating it with Google Cloud services such as DataFlow, is highly valued
- Ability to optimize Beam pipelines for scalability and efficiency is a plus
- Familiarity with dbt and data modeling concepts, including data warehousing and data lake architecture, is beneficial for designing and optimizing data layers
- Proficiency in Spark and Scala, including integrating them with Google Cloud services such as DataProc, is desirable for building and maintaining data engineering solutions
Benefits
- This role can be performed from anywhere in the United States of America
- Long-term career development support
Company Overview
Company H1B Sponsorship