Senior AI Inference Engineer - Model Optimization & Deployment
Foster City, CAOn-siteFull-time
AI Summary
Senior AI Inference Engineer optimizing, compiling, and deploying large multimodal and LLM models for real-time, on-vehicle execution.
About this role
The Perception team is pioneering the development of a multi-modality foundation model to drive the next generation of autonomous system intelligence.
As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient, production-ready large-scale models to our on-vehicle stack. We are looking for experts with hands-on experience in compressing, accelerating, and deploying complex models (LLMs, VLMs, or FMs) for power- and thermal-constrained vehicle SOCs. You will optimize the ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on edge devices.
In this role, you will:
Qualifications:
Bonus Qualifications:
Skills
C++ConcurrencyCUDAEdge DeploymentFlashAttentionKv-cacheLatency BenchmarkingLinear-attentionMemory SafetyMixed PrecisionModel CompressionONNXPagedAttentionPruningPTQPyTorchQATQuantizationSpeculative DecodingTensorRT
Explore related jobs
More jobs at Zoox
- Part-time Student Worker – Software Development Engineer in TestFoster City, CA
- Part-Time Student Worker – AI Validation and Benchmarking EngineerFoster City, CA
- Senior Commercial Operations AnalystFoster City, CA
- Simulation Expert Triage EngineerFoster City, CA
- Systems Test Engineer, System Behavior AnalysisSan Diego, CA
- Manager, RL Algorithms & DecoderFoster City, CA