Description
You will lead the migration of legacy SAS workloads to cloud and containerized platforms while designing scalable data architectures.
This role is on-site.
Responsibilities
- Convert existing SAS-based pipelines to PySpark for distributed execution.
- Design and implement scalable data pipelines and technical architectures.
- Collaborate with stakeholders to document dataflows and define target-state solutions.
- Manage workload transitions to cloud platforms while enforcing data governance and quality policies.
- Research and integrate open-source technologies into system designs.
Required Skills
- 8+ years of experience in data engineering and large-scale data platform design.
- 3+ years in a lead or architect role.
- 2+ years of hands-on SAS coding experience.
- 2+ years developing data solutions on AWS or GCP.
- Proficiency in Python, PySpark, or Scala, including basic machine learning libraries.
- Hands-on experience with Docker and Kubernetes containerization.
- Expert knowledge of data warehousing concepts and distributed systems.
- Bachelor of Computer Science degree.
Preferred Skills
- Relevant industry certifications.
- Experience with Agile (Scrum) development methodologies.