Description
You will lead the migration of legacy SAS workloads to cloud and containerized platforms while designing scalable data architectures.
This role is on-site.
Responsibilities
- Convert existing SAS-based pipelines to PySpark for execution on distributed ecosystems.
- Design and implement scalable data pipelines that enforce data governance and quality policies.
- Develop high-level and detailed technical designs focusing on scalability, security, and performance.
- Collaborate with stakeholders to document dataflows and translate business requirements into technical solutions.
- Research and integrate open-source technologies and cloud-based components into the data architecture.
Required Skills
- 8+ years of experience in data engineering and large-scale data platform design.
- 3+ years of experience in a lead or architect role.
- 2+ years of experience in SAS coding.
- 2+ years of experience developing data solutions on AWS or GCP.
- Proficiency with Python, PySpark, and Scala.
- Hands-on experience with Docker and Kubernetes.
- Expert understanding of data warehousing, distributed systems, and cloud platforms (AWS, GCP).
- Experience with machine learning libraries.
- Strong skills in automation and technical documentation.
Preferred Skills
- Knowledge of Agile (Scrum) development methodologies.
- Experience with business process modeling tools.