← Back to jobs

InfiCare Technologies Logo
Site Reliability Engineer

InfiCare Technologies

 

Jersey City, NJ, USA

Posted On: 1 day ago
Experience: 10+ years
Availability: Hybrid
Openings: 1
Category: Site Reliability Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will own AWS production stability, batch job reliability, and incident response for critical financial services workloads.

This role is on-site.

Responsibilities

  • Monitor AWS environments and batch schedules to ensure SLA compliance and rapid incident resolution.
  • Triage job failures, execute reruns, and coordinate with engineering teams to restore service during complex outages.
  • Support deployment activities, including pre/post-implementation checks and rollback execution.
  • Create and maintain runbooks, SOPs, troubleshooting guides, and escalation paths.
  • Drive root cause analysis (RCA) and implement preventive actions to reduce repeat incidents.

Required Skills

  • 10+ years of hands-on production experience with AWS services: EC2, S3, Lambda, CloudWatch, IAM, and VPC networking.
  • Deep expertise in batch scheduling and job management, including dependencies, sequencing, retry logic, cutoff management, and failure handling.
  • Proficiency in SQL for operational validation and troubleshooting.
  • Python and automation scripting experience for operational tooling.
  • Bachelor's degree in a relevant field.

Preferred Skills

  • Experience in enterprise-scale production environments within financial services or regulated industries.
  • Strong incident communication and documentation skills.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs