← Back to jobs

Intime Infotech Inc Logo
Senior Production Engineer

Intime Infotech Inc

 

Sunnyvale, CA, USA

Posted On: 30+ days ago
Experience: 5+ years
Availability: Onsite
Openings: 1
Category: Production Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

 

•            Collaborate with cross-functional teams to define and evolve availability metrics for Client’s cloud platform, including establishing, measuring, and improving SLIs and SLOs

•            Participate in production incident response, diagnosing and resolving service disruptions while contributing to post-incident reviews and root cause analysis

•            Build, operate, and improve observability across Client’s infrastructure using tools such as Prometheus, Grafana, Alertmanager, and OpenTelemetry

•            Identify reliability risks, performance bottlenecks, and early indicators of potential production issues across distributed systems

•            Develop automation and tooling that reduces operational toil, improves recovery times, and enables self-healing infrastructure

•            Partner with compute, networking, storage, and platform teams to strengthen service resilience and disaster recovery capabilities

•            Contribute to improving operational processes, knowledge sharing, and reliability best practices across the engineering organization

•            Continue growing technical depth through mentorship, training, and hands-on work operating large-scale AI infrastructure

  

•            5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations

•            Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems

•            Strong knowledge of Linux/Unix systems, including debugging complex issues across kernel and user space

•            Previous experience in Infrastructure roles building or managing compute, storage or networking platforms

•            Understanding of modern cloud infrastructure fundamentals including Kubernetes, distributed systems, virtualization, and cloud platforms (AWS/GCP)

•            Familiarity with incident management practices and reliability frameworks (SRE, ITIL, or similar)

•            Experience with monitoring and observability tools such as Prometheus and Grafana, or a strong desire to deepen expertise in this area

•            Familiarity with infrastructure-as-code and configuration management tools such as Terraform or Ansible

•            Scripting or programming experience with languages such as Go, Python, C, or C++

•            Strong communication skills and the ability to collaborate across engineering teams

•            Ability to remain calm and effective while troubleshooting complex issues in high-impact production environments

•            A growth mindset and strong interest in reliability engineering, automation, and operational excellence

 

Additional Qualifications

(Nice to Haves)            •            Experience working with Kubernetes or container orchestration platforms at scale

•            Exposure to change management processes, operational readiness reviews, or structured root cause analysis

•            Experience designing self-healing systems, automated remediation, or event-driven operational tooling

•            Interest in scaling AI or HPC infrastructure and solving reliability challenges in GPU-heavy environments

 

•            Passion for mentorship, learning, and developing deeper expertise in Production Engineering

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs