Description
You will build, maintain, and optimize scalable data pipelines and cloud infrastructure.
Responsibilities
- Design and develop ETL pipelines using Dagster, NiFi, and Apache Spark.
- Implement data transformation workflows and modeling using DBT.
- Architect and manage cloud infrastructure on Google Cloud Platform, including BigQuery, Dataflow, and Cloud Storage.
- Containerize applications and manage production workloads using Docker and Kubernetes.
- Monitor pipeline performance and infrastructure availability with Prometheus and Grafana.
Required Skills
- 5+ years of experience in data engineering or a similar role.
- Proficiency in Python, including Pandas and NumPy for data manipulation.
- Experience with Dagster and DBT for workflow orchestration.
- Hands-on experience with Apache NiFi and Apache Spark.
- Strong knowledge of SQL/T-SQL and relational databases including SQL-Server, PostgreSQL, and MySQL.
- Experience with Google Cloud Platform (GCP) services.
- Practical experience with Docker and Kubernetes in production environments.
- Ability to use Prometheus and Grafana for infrastructure monitoring.
Preferred Skills
- Python optimization techniques such as threading and multi-processing.