← Back to jobs
Englewood Cliffs, New Jersey, USA
No related jobs found
Must Have Technical/Functional Skills
• 6-7 years of experience in Site Reliability Engineering, Production Support, DevOps, or Infrastructure Operations.
• Strong understanding of Linux administration and troubleshooting.
• Hands-on experience with AWS cloud services (EC2, RDS, IAM, VPC, CloudWatch, S3).
• Experience with monitoring, alerting, and observability tools.
• Knowledge of incident management, problem management, and RCA processes.
• Experience with automation and scripting using Shell and/or Python.
• Working knowledge of PostgreSQL and MySQL databases.
• Experience with Git version control.
• Understanding of CI/CD concepts and tools such as Jenkins.
Roles & Responsibilities
• Design, implement, and maintain highly available and reliable production systems.
• Automate operational tasks and infrastructure management using Shell, Python, Ansible, or Terraform.
• Manage and support AWS services including EC2, RDS, S3, IAM, VPC, CloudWatch, and related cloud services.
• Perform Linux server administration, troubleshooting, patching, and performance tuning.
• Monitor application and infrastructure health using tools such as Grafana, Prometheus, CloudWatch, Datadog, Splunk.
• Participate in incident management, root cause analysis (RCA), and problem management activities.
• Define and maintain SLIs, SLOs, and SLAs to ensure service reliability.
• Support PostgreSQL and MySQL databases for operational and basic administration tasks.
• Collaborate with development, QA, cloud, and support teams to improve system reliability and deployment processes.
• Drive automation, observability, capacity planning, security, and operational best practices.
• Participate in on-call support and production issue resolution
Bachelor's degree
No related jobs found
← Back to jobs