Senior HPC & AMD Infrastructure Engineer
AI Summary
Owns the health, reliability, and performance of AMD GPU compute clusters, handling hardware operations, Linux systems engineering, distributed infrastructure, and ML workloads.
About this role
United States | Hybrid or Remote | Full-time
The role
You'll own the health, reliability, and performance of Evergrid's AMD-only GPU compute clusters.
You're the primary custodian of our high-density accelerator environments. The work spans hardware operations, Linux systems engineering, distributed infrastructure, and ML workloads. It covers GPU bring-up and kernel-level debugging, as well as maintaining and optimizing the ROCm-based ML stack behind production-scale AI. If you like getting maximum performance out of hardware, debugging GPUs at scale, and shipping world-class AI infrastructure, this role is for you.
What you'll own
System health and reliability (SRE)
Primary on-call response for outages, GPU failures, node crashes, and cluster-wide incidents.
Being the key point person for POC and active customers' GPU clusters.
Fast diagnosis and resolution that minimizes downtime and keeps SLA-level reliability.
Monitoring for GPU health, thermals, PCIe topology, memory errors, and cluster load.
Repairs, RMAs, and physical maintenance, coordinated with data center operators, hardware vendors, and on-site technicians.
Linux and network administration
Installing, patching, and maintaining Linux (Ubuntu, CentOS, RHEL) across large GPU node fleets.
Kernel tuning, consistent OS configuration, and fleet automation at scale.
Secure networking: VPNs, firewalls (iptables/firewalld), SSH hardening, and routing.
Identity and access systems (LDAP, FreeIPA, Active Directory).
Distributed storage (NFS, GPFS, Lustre).
AMD GPU and ML stack engineering (ROCm-first)
Deployment and bring-up of new GPU nodes, including BIOS configuration, NUMA tuning, and topology validation.
AMD GPU drivers, kernel modules, and the ROCm runtime across production fleets.
The AMD ML stack: ROCm, PyTorch (ROCm builds), JAX (ROCm/XLA), RCCL, hipBLAS/hipDNN, MIOpen, and supporting runtimes.
Debugging complex failures across GPUs, compilers, ML frameworks, and distributed training and inference. Examples include RCCL hangs, HIP memory faults, ROCm kernel crashes, framework build and link issues, and vLLM build failures on ROCm.
Infrastructure that supports both research iteration and production reliability, built with the ML and platform teams.
What you bring
5+ years in HPC, GPU cluster operations, Linux systems engineering, or similar roles.
A bachelor's or master's in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
Deep hands-on experience with AMD MI-series GPUs, including driver and kernel-level debugging.
Strong grasp of Linux internals, kernel modules, hardware bring-up, and performance tuning.
Experience securing and operating production infrastructure: VPNs, firewalls, SSH, and identity systems.
Proficiency in Bash and Python for automation, tooling, and operations.
Strong familiarity with ML stacks and runtime behavior in ROCm environments (ROCm/HIP, MIOpen, RCCL, PyTorch, JAX).
Experience debugging high-performance networking and RDMA (InfiniBand or RoCE), including cluster-level communication failures that affect distributed training.
Bonus
Schedulers and orchestration (Slurm, Kubernetes).
Model serving and inference optimization on ROCm (vLLM, SGLang).
Configuration management and IaC (Ansible, SaltStack, Terraform).
Supporting ML research or production AI teams at a startup or high-growth company.
Benefits & Compensation
Compensation: $180,000 - $220,000
Health insurance
Paid time off and paid holidays
Home office stipend
Skills
Explore related jobs
More jobs at Evergrid
Similar Linux jobs
Jobs in New York City
- I
Shift LeaderInsomnia Cookies · New York City NY (Bronx) - T
Operations InternTransperfect · New York City, New York - T
Payroll AnalystTransperfect · New York City, New York - MManager, Lighting SystemsMSG Entertainment Holdings, LLC · New York City, New York
- WCommunications Manager, Product & TechWaymo · Mountain View, CA
- A
Podcast Media ManagerAnthropic · San Francisco, California | New York City
Browse these categories
Market data for this role
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.
