← Back to jobs

Shimento Inc Logo
Site Reliability Engineer

Shimento Inc

 

Santa Clara, CA, USA

Posted On: 8 days ago
Experience: 5+ years
Availability: Onsite
Openings: 1
Category: Site Reliability Engineer
Tenure: Contract - W2
Related Jobs

No related jobs found

Description

You will design, build, and manage large-scale multi-cluster Kubernetes platforms across cloud and on-prem environments. This role is on-site.

Responsibilities

  • Develop and maintain controllers, CRDs, ingress, DNS, and TLS automation for scalable infrastructure provisioning.
  • Build and secure platform microservices, including CI/CD pipelines, SSO, RBAC, secret management, and monitoring workflows.
  • Own end-to-end production release management, covering Helm deployments, multi-architecture container builds, staged rollouts, and rollback strategies.
  • Implement observability, audit logging, analytics, and automation to support high-scale production operations.
  • Integrate AI-driven tooling and automation into operational workflows to improve platform scalability and efficiency.

Required Skills

  • 6+ years of hands-on DevOps/SRE experience supporting production Kubernetes environments.
  • Strong expertise with Kubernetes operators, CRDs, ingress controllers, and cluster networking.
  • Strong programming skills in Python or Go, with working knowledge of TypeScript/React.
  • Solid experience with AWS or similar cloud platforms, OIDC/SAML authentication, and secret management solutions.
  • Strong understanding of relational databases, caching technologies, and asynchronous communication patterns (WebSockets, SSH tunneling, message queues).
  • Proven experience building internal developer platforms and enterprise CI/CD pipelines.
  • Bachelor’s or Master’s degree in Computer Science or related field.

Preferred Skills

  • Experience with Jenkins, GitLab CI, ArgoCD, Flux, Prometheus, Grafana, VictoriaMetrics, Datadog, Splunk, or Kibana.
  • Deep understanding of agentic AI workflows, developer tooling, CLI frameworks, and MCP ecosystems.
  • Experience building AI-assisted operational tooling including anomaly detection, automated runbooks, and LLM-powered operations workflows.

Education

Bachelor’s/Master’s degree

Related Jobs

No related jobs found

← Back to jobs