Description
You will design, build, and maintain scalable data pipelines and workflows on Google Cloud Platform. You will optimize data models and warehouses to support business analytics while ensuring data quality, security, and compliance.
This role is on-site.
Responsibilities
- Design and develop scalable data pipelines using GCP services such as BigQuery, Dataflow, and Dataproc.
- Build and optimize data schemas and warehouses aligned with business requirements.
- Automate data workflows, monitor system health, and troubleshoot pipeline issues proactively.
- Implement data security, encryption, and compliance policies across all cloud environments.
- Document architecture, processes, and best practices for data engineering solutions.
Required Skills
- 5+ years of experience in data engineering or data architecture.
- Proficiency in SQL and programming languages: Python, Java, or Scala.
- Hands-on experience with GCP data services: BigQuery, Cloud Dataflow, Cloud Storage, and Dataproc.
- Strong understanding of ETL/ELT processes, data modeling, and warehousing concepts.
- Experience with Apache Spark and real-time streaming tools like Kafka or Pub/Sub.
- Familiarity with containerization (Docker) and orchestration (GKE, Kubernetes).
- Knowledge of CI/CD integrations and version control (Git).
Preferred Skills
- Experience with Terraform or Deployment Manager for infrastructure automation.
- Exposure to machine learning workflows on GCP (Vertex AI, AI Platform).