Helix AI Engineer, Training Performance
AI Summary
Helix AI Engineer at Figure optimizes distributed training frameworks, GPU kernels, and data pipelines for large-scale humanoid robot models across 100k+ GPUs.
About this role
Figure's vision is to deploy autonomous humanoids at a global scale. Our Helix team is looking for an experienced AI Training Performance Engineer to take our model training to the next level. This role is focused on improving distributed training frameworks for large scale model training, optimizing GPU kernels, exploring the relative gains of different accelerator types and co-designing our models to maximize utilization of our hardware.
Responsibilities
- Optimize training performance for a 100B+ parameter models across 100k+ GPUs.
- Collaborate with the broader team on accelerator choice, cluster topology, scheduling, and hardware procurement decisions to inform future scaling.
- Write and optimize custom kernels (Triton/CUDA)
- Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis across training jobs
- Optimize data loading and preprocessing pipelines so I/O never gates the accelerators
- Improve checkpointing, fault tolerance, and elastic restart so large jobs recover quickly from node failures without losing significant wall-clock time
- Partner with researchers to co-design model architectures and training recipes that are performant at scale (e.g., activation checkpointing strategies, mixed precision, sequence packing)
- Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators
- Build and extend agentic systems that automatically generate, benchmark, and iterate on custom kernels
- Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware) for fit with our training workloads, and lead proof-of-concept ports/benchmarks
- Explore different model/data parallelisms (FSDP, context parallel, expert parallel, etc.) to determine optimal configuration per model size.
Requirements
- Bachelor's or Master's degree in Computer Science, Computer/Electrical Engineering, or a related field
- 3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects
- Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy)
- Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ability to translate traces into concrete optimizations
- Solid grasp of collective communication (NCCL) and modern networking concepts (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement).
- Strong Python and CUDA/C++ skills; comfortable reading and modifying framework internals
- Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence)
- Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) and using them to drive optimization priorities
Bonus Qualifications
- Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration
- Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.)
- Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management.
The US base salary range for this full-time position is between $200,000 - $400,000 annually.
The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.
Skills
Explore related jobs
More jobs at Figure
Similar CUDA jobs
Jobs in San Jose
- Customer Service Rep(08240) - 5313 PROSPECT RDDomino's · San Jose, CA
- Service Advisor - Honda of Stevens CreekSonic Automotive · San Jose, CA
Accounting Manager- Construction, Comm GCTGG Accounting · San Jose, California- Director, FinanceWestern Digital · San Jose, CA
- Executive Assistant - VP LegalWestern Digital · San Jose, CA
Ingeniero/a DevSecOps – Nivel IntermedioSoluciones Seguras · San Jose, San Jose
Browse these categories
Market data for ai / ml engineer roles
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.