Description
You will design, build, and automate data pipelines to provide reliable data for BI, advanced analytics, and APIs.
Responsibilities
- Develop and operationalize ETL/ELT processes to create enterprise-certified datasets from internal and external sources.
- Design and implement data pipelines using Informatica Intelligent Cloud Services (IICS) and Azure Data Factory (ADF).
- Optimize data storage, ingestion, quality, and orchestration to improve reliability and efficiency.
- Build and implement scalable big data and NoSQL solutions to drive high-value insights.
- Partner with DevOps, product owners, and engineering teams to automate infrastructure and drive CI/CD integration.
- Conduct ad-hoc data retrieval and assess data integrity across multiple sources.
Required Skills
- 5+ years of experience engineering and operationalizing data pipelines with large, complex datasets.
- Advanced SQL skills with the ability to write complex queries against relational databases.
- Hands-on experience with Informatica PowerCenter or IICS.
- Proficiency with Azure services including Databricks, Synapse (Azure SQL DW), Data Lake Storage (ADLS), Event Hub, and Cosmos DB.
- Experience with Spark, Kafka, and Cribl.
- Expertise in data modeling, ETL, and data warehousing.
- Experience working with various data formats including XML, JSON, CSV, and APIs.
- Experience with DB2, SQL, and Oracle data sources.
- Knowledge of Python, Kafka, and Big Panda.
Preferred Skills
- Experience applying machine learning and statistical approaches to predict business outcomes.
- Experience designing and building ML/DL models.