Description
You will build and maintain scalable ETL/ELT pipelines to support AI/ML workloads within the GCP ecosystem.
Responsibilities
- Design and maintain batch and real-time data pipelines using GCP services like Dataflow, Dataproc, and Cloud Composer.
- Prepare and curate structured, unstructured, and streaming datasets to optimize them for model training and serving.
- Implement data quality, security, and governance standards across the entire data lifecycle.
- Automate data processes and develop CI/CD pipelines for both data and ML models.
- Monitor and optimize the performance and cost-effectiveness of data pipelines and AI/ML infrastructure.
Required Skills
- 5+ years of experience as a Data Engineer focused on AI/ML applications.
- In-depth hands-on experience with Google Cloud Platform (BigQuery, Dataflow, Dataproc, Cloud Storage, Cloud Composer).
- Strong proficiency in Python.
- Expertise in SQL and various database technologies, including relational, NoSQL, and data warehouses.
- Experience with machine learning frameworks such as TensorFlow or PyTorch.
- Knowledge of machine learning workflows, including feature engineering, training, and deployment.
- Experience with version control (Git) and CI/CD practices.
- Understanding of distributed systems, big data technologies, and real-time processing.
- Bachelor's or Master's degree in Computer Science, Data Engineering, or a related quantitative field.
Preferred Skills
- Google Cloud Professional Data Engineer certification.
- Experience with Scala or Java.