Architect, build, and manage AI-enabled CI/CD pipelines that improve developer productivity, code quality, release reliability, and deployment speed.
Design and deploy production-grade Model Context Protocol clients and servers to securely connect enterprise LLMs with engineering tools, repositories, cloud infrastructure, and observability platforms.
Develop custom MCP servers using Python, TypeScript, Node.js, or JavaScript to expose logs, infrastructure metrics, deployment data, and internal tools to authorized AI agents.
Integrate LLM agents into developer workflows to support automated code review, vulnerability detection, test generation, release validation, and infrastructure recommendations.
Build and maintain robust CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, ArgoCD, Jenkins, or similar tools.
Implement ChatOps 2.0 capabilities that allow engineers to interact with deployment pipelines, cloud environments, logs, and operational workflows using secure conversational interfaces.
Create safe autonomous remediation workflows for log analysis, incident triage, root-cause analysis, and infrastructure issue resolution.
Build guardrails that allow AI agents to generate, inspect, and safely execute Infrastructure as Code using Terraform, OpenTofu, Terragrunt, Pulumi, Crossplane, or similar tools.
Manage containerized workloads using Docker and Kubernetes platforms such as AWS EKS, Azure AKS, or Google GKE.
Integrate AI-driven observability workflows with platforms such as Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, or ELK.
Implement AI safety controls including role-based access control, least-privilege execution, human-in-the-loop approvals, audit logging, rollback mechanisms, and secure tool access.
Partner with software engineering, DevOps, SRE, security, platform, and data/AI teams to identify opportunities for intelligent automation.
Create reusable automation frameworks, runbooks, dashboards, documentation, and enablement materials for engineering teams.
Drive an “automate everything” culture by reducing manual toil and improving operational efficiency across cloud and software delivery processes.
Basic Qualifications
Minimum 7+ years of experience in DevOps, Cloud Engineering, SRE, Platform Engineering, or Infrastructure Automation.
Minimum 4+ years of hands-on experience designing and managing CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD, or similar platforms.
Minimum 3+ years of experience managing scalable cloud environments in AWS, Azure, or GCP, with strong preference for AWS.
Strong hands-on experience with Kubernetes, Docker, and production container orchestration platforms such as EKS, AKS, or GKE.
Advanced proficiency with Infrastructure as Code tools such as Terraform, OpenTofu, Terragrunt, Pulumi, CloudFormation, or Crossplane.
Strong programming and scripting experience using Python, TypeScript, JavaScript, Bash, or Go.
Practical experience working with LLM APIs such as OpenAI, Anthropic, or similar enterprise AI platforms.
Experience with AI orchestration or agentic frameworks such as LangChain, CrewAI, LlamaIndex, or similar tools.
Strong understanding of the Model Context Protocol ecosystem and experience designing or integrating MCP clients and servers.
Experience integrating DevSecOps controls into CI/CD pipelines, including SAST, DAST, dependency scanning, container scanning, secrets scanning, and vulnerability management.
Strong knowledge of secret management and security tooling such as HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, or similar platforms.
Experience with observability, monitoring, logging, and alerting platforms such as Datadog, Prometheus, Grafana, CloudWatch, Splunk, Dynatrace, or ELK.
Familiarity with security and compliance frameworks such as SOC2, ISO27001, or enterprise audit control environments.
Ability to troubleshoot complex pipeline, infrastructure, deployment, and production issues across cloud-native environments.
Preferred / Nice to Have
Experience building AI-assisted infrastructure provisioning workflows.
Experience implementing autonomous or semi-autonomous incident response and remediation capabilities.
Experience with MLOps, model deployment pipelines, model monitoring, MLflow, SageMaker, or equivalent platforms.
Experience implementing human-in-the-loop approval models for AI-generated operational actions.
Experience with policy-as-code tools such as Open Policy Agent, Sentinel, Checkov, or similar solutions.
Experience working in regulated industries such as banking, financial services, healthcare, or insurance.
Experience with GitOps operating models using ArgoCD, Flux, or similar tools.
AWS, Kubernetes, DevOps, Security, or AI/ML certifications are a plus.
Soft Skills & Mindset
Strong “automate everything” mindset with a passion for reducing repetitive manual tasks and operational toil.
Security-first approach with practical skepticism of autonomous AI actions and a focus on validation, boundaries, approvals, and rollback.
Ability to bridge traditional software engineering, DevOps, SRE, security, and data/AI teams.
Strong communication skills with the ability to explain complex AI-enabled DevOps concepts to both technical and leadership audiences.
Collaborative educator who can help upskill engineering teams on AI-assisted delivery, secure automation, and modern DevOps practices.
Ownership mindset with the ability to design solutions, implement them hands-on, and support them in production.