Description
You will lead the migration of legacy SAS workloads to cloud and containerized platform solutions.
Responsibilities
- Convert existing SAS-based pipelines to PySpark for execution on distributed ecosystems.
- Design and implement scalable, high-performance data pipelines and technical architectures.
- Collaborate with stakeholders to document dataflows, capture requirements, and define target-state solutions.
- Manage the transition of workloads to cloud platforms while enforcing data governance and quality policies.
- Research and integrate open-source technologies and components into system designs.
Required Skills
- 8+ years of experience in data engineering and designing large-scale data platforms.
- 3+ years of experience in a lead or architect role.
- 2+ years of hands-on experience with SAS coding.
- 2+ years of experience developing data solutions on AWS or GCP.
- Proficiency with Python, PySpark, or Scala, including basic machine learning libraries.
- Hands-on experience with containerization using Docker and Kubernetes.
- Expert knowledge of data warehousing concepts and distributed systems.
- Bachelor of Computer Science degree.
Preferred Skills
- Relevant industry certifications.
- Experience with Agile (Scrum) development methodologies.