You will develop big data applications within the Hadoop ecosystem and drive DevOps and data engineering best practices.
Responsibilities
Develop big data applications leveraging Hadoop, NoSQL, In-memory Data Grids, Kafka, Spark, and Ab Initio.
Build and support Jenkins and Cloudbees CI/CD pipelines.
Use Sonarqube for code quality analysis and build dashboards for stakeholders.
Create, manage, and execute test cases and scripts to identify process bottlenecks.
Document data asset lineage in enterprise metadata repositories and implement data quality rules.
Identify repeatable engineering processes to develop common reusable tools and frameworks.
Required Skills
Programming experience in Scala, Java, or Python.
2+ years of hands-on experience with shell scripting, complex SQL queries, Hive scripts, Hadoop commands, and Git.
Experience with ETL, data warehousing, or data lake environments.
Ability to write abstracted, reusable code components.
Bachelor's degree in a quantitative field such as Engineering, Computer Science, Statistics, or Econometrics, or 4-5 years of IT experience in lieu of a degree.
Strong understanding of functional and non-functional requirements.
Experience working in an agile development process including backlog grooming and code reviews.
Preferred Skills
Familiarity with Ab Initio tools (GDG, Co>Operating System, Metadata Hub, etc.).
Experience with Hortonworks/Cloudera, ZooKeeper, Oozie, and Kafka.
Exposure to AWS data analytics services and Collibras data management tools.