← Back to jobs
Jersey City, NJ, USA
No related jobs found
Skills
Java
SRE
Job Summary: We are seeking a highly experienced Lead Site Reliability Engineer (SRE) with strong Java development experience to join the technology organization supporting highly available, scalable, resilient, and business-critical applications.
The ideal candidate will have 12+ years of overall technology experience, with strong hands-on expertise in Java, application reliability, production engineering, observability, automation, cloud technologies, CI/CD, incident management, and performance engineering.
The Lead SRE will apply a software engineering mindset to production operations, developing automation and reliability solutions rather than relying solely on traditional infrastructure support. The role will work closely with application developers, architects, DevOps engineers, platform teams, and technology stakeholders to improve system availability, stability, scalability, and operational efficiency.
Key Responsibilities
Lead Site Reliability Engineering initiatives for critical enterprise applications and platforms.
Develop and maintain Java-based automation, reliability, monitoring, and operational tools.
Apply software engineering principles to improve application availability, scalability, resiliency, and performance.
Own production stability and participate in incident management, problem management, and root-cause analysis.
Design and implement proactive monitoring, alerting, health checks, and automated remediation.
Analyze production issues, identify systemic problems, and implement permanent corrective actions.
Establish and improve SRE practices, SLOs, SLIs, SLAs, error budgets, and reliability metrics.
Build automation to eliminate repetitive manual operational activities.
Work closely with Java development teams to improve application reliability and production readiness.
Troubleshoot complex Java/JVM, application, API, database, network, and infrastructure-related issues.
Perform application performance analysis, including JVM, memory, CPU, thread, garbage collection, latency, and throughput analysis.
Design and improve CI/CD pipelines and automated deployment processes.
Support highly available applications across cloud and distributed environments.
Implement resiliency patterns including fault tolerance, failover, disaster recovery, and capacity planning.
Develop dashboards and observability solutions for application and infrastructure health.
Participate in production deployments, release management, and post-production validation.
Drive automation and continuous improvement across the application lifecycle.
Provide technical leadership and mentoring to other engineers.
Collaborate with architecture, development, infrastructure, security, and platform engineering teams.
Required Technical Skills
Must Have:
12+ years of IT/software engineering experience
Strong hands-on Java development
Strong Site Reliability Engineering / Production Engineering experience
Java/Spring Boot or enterprise Java application experience
Strong understanding of JVM internals and Java application performance
Production support and troubleshooting of large-scale applications
Linux/Unix
REST APIs / Microservices
CI/CD
Jenkins / GitHub Actions / GitLab CI or similar
Kubernetes / Docker
Cloud experience — AWS / Azure / GCP
Monitoring and observability tools
Splunk / ELK or equivalent logging platforms
Prometheus / Grafana or equivalent monitoring tools
Strong scripting/automation using Python, Shell, or similar
Incident management and Root Cause Analysis (RCA)
Application performance and capacity management
High availability, resiliency, scalability, and disaster recovery concepts
Strong SQL/database troubleshooting skills
Strong communication and stakeholder-management skills
Any Graduate
No related jobs found
← Back to jobs