Strong scripting and automation skills using Python, Bash, and Terraform.
Demonstrated experience implementing SRE practices (monitoring, observability, incident management) with tools like Catchpoint, New Relic, and Grafana.
Knowledge of high availability, scalability, and resiliency patterns in Kubernetes.
Experience setting up CI/CD automation and integrating SonarQube and unit testing frameworks.
Understanding of LLM traceability, monitoring, and model governance.