← Back to jobs

Hallmark Global Technologies Inc Logo
System Engineer
Posted On: 2 days ago
Experience: 5+ years
Availability: Hybrid
Openings: 1
Category: System Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will design, deploy, and maintain high-performance computing (HPC) infrastructure on AWS, ensuring scalability and reliability for complex workloads.

This role is on-site.

Responsibilities

  • Administer Linux (RHEL) systems and manage HPC applications, including Slurm scheduling and GPU stacks (NVIDIA, CUDA, NCCL).
  • Provision and manage AWS infrastructure using Terraform and CloudFormation, leveraging core services like EC2, VPC, IAM, S3, and EBS.
  • Operate and troubleshoot AWS Parallel Cluster (v3+) and AWS Batch for containerized and burstable workloads.
  • Implement CI/CD automation for infrastructure and applications using GitHub Actions or AWS CodeBuild.
  • Set up monitoring and observability using CloudWatch, Prometheus, or Grafana to track cluster health and job metrics.

Required Skills

  • 5+ years of experience in system engineering or DevOps roles.
  • Hands-on expertise with AWS core services and infrastructure as code (Terraform preferred, CloudFormation required).
  • Strong proficiency in Linux administration (RHEL), including SELinux, systemd, and networking.
  • Experience managing HPC workloads with Slurm, including partitions, QoS, and accounting.
  • Proficiency in scripting and automation using Bash, Python, and Shell Scripting.
  • Familiarity with Windows administration and hybrid environment integration.
  • Experience with containerized workloads (ECS/EKS) and GPU driver installation/troubleshooting.

Preferred Skills

  • AWS certifications (Solutions Architect Associate/Professional, SysOps, or Specialty).
  • Experience deploying or tuning HPC workloads for CFD/FEA applications (e.g., Ansys, Abaqus, OpenFOAM).

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs