Description
Key Skills: Python, Java, Terraform, Ansible, AWS, Docker, Kubernetes, TensorFlow, PyTorch, Prometheus
Good to Have Skills: Experience with cloud platforms like Azure and GCP, proficiency in Scikit-learn ML framework, knowledge of observability tools like Grafana and ELK stack, familiarity with CloudFormation for infrastructure automation, understanding of Spinnaker for deployment management, and experience with intelligent agents for workflow automation and decision-making processes.
Roles & Responsibilities:
- Design and implement CI/CD pipelines, infrastructure-as-code frameworks, and container orchestration strategies leveraging various DevOps tools.
- Lead the architecture, deployment, and management of cloud infrastructure in AWS while establishing best practices for reliability and scalability.
- Drive the adoption of AI and machine learning capabilities within DevOps workflows including intelligent monitoring and predictive analytics.
- Lead the integration of intelligent agents for workflow automation, decision-making, and process optimization across development environments.
- Develop AI-powered observability solutions to monitor, analyze, and proactively manage application and infrastructure health using advanced techniques.
- Work closely with cross-functional teams including engineering, product, and operations to identify automation opportunities and deliver solutions.
- Stay abreast of emerging AI/ML technologies, frameworks, and industry trends while driving continuous improvement initiatives.
- Provide hands-on technical guidance to a team of software and DevOps engineers fostering innovation and continuous learning.
- Conduct code reviews, architectural assessments, and design discussions to uphold engineering excellence standards across the organization.
Experience Required: 5+ years in AI/ML engineering with proven expertise in agent-based systems and automation, formal training or certification on software engineering concepts