
Posted 1 month ago
Software Engineer - ML Infrastructure
AI Summary
Builds and optimizes ML infrastructure and data systems so research teams can train and ship state-of-the-art medical imaging models; owns data pipelines, distributed training, reinforcement learning stacks, and production deployment.
About this role
About Us
We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.
Role Overview
We're seeking a Software Engineer to build the ML infrastructure and data systems that let our research team train and ship state-of-the-art models for medical imaging. Sitting in the Engineering team and working closely with research, you'll own the data pipelines that unify live production traffic with offline datasets, the distributed training and reinforcement learning infrastructure our foundation-model and post-training work runs on, and the inference and evaluation systems that carry models from experimentation into production. This role requires someone who can move fluidly between ML systems, data engineering, distributed training, and production deployment, and who measures success by how quickly the research team can iterate.
Key Responsibilities
Build and optimize distributed training infrastructure for foundation models on large-scale medical imaging, including the long-context parallelism and checkpointing that volumetric CT/MR training demands.
Build the reinforcement learning training stack (high-throughput rollout generation, reward-model serving, and experience collection), enabling the research team to run online, multi-reward RL at scale.
Build high-throughput data loading and preprocessing that keeps GPUs saturated on large volumetric and multimodal datasets.
Design and implement robust data pipelines to collect, process, and store large-scale multimodal medical imaging data from both production traffic and offline sources.
Build centralized data storage solutions with standardized formats (e.g., protobufs) that enable efficient retrieval and training across the organization.
Partner with researchers to prototype new ideas and translate them into production-ready code, owning end-to-end delivery from experimentation through deployment and monitoring.
Contribute to production serving and deployment pipelines — model rollout, canary deployments, and monitoring — alongside the backend team.
Qualifications
5+ years building ML infrastructure, data pipelines, or ML systems in production
Strong Python skills and expertise in PyTorch or JAX
Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and the systems concerns of keeping large GPU jobs efficient
Hands-on experience with data pipeline technologies (e.g., Spark, Airflow, BigQuery, Snowflake, Databricks, Chalk) and schema design
Experience with distributed systems, cloud infrastructure (AWS/GCP), and containerization (Docker/Kubernetes)
Track record of building scalable data systems and shipping production ML infrastructure
Ability to move quickly and handle competing priorities in a fast-paced environment
Preferred Qualifications
Experience building reinforcement learning training infrastructure: rollout generation, reward-model serving, or online/off-policy learning systems
Experience with high-performance inference and serving (vLLM, SGLang, TensorRT, or Triton) for both training-time rollouts and production
Experience building internal training or experimentation platforms used by research teams
Experience supporting A/B testing and experimentation workflows, including canary deployments and monitoring statistical significance
Familiarity with vision-language models (VLMs) or multimodal architectures
Experience with medical imaging formats (DICOM) and healthcare data standards
Familiarity with MLOps practices and model deployment pipelines
Experience with privacy-preserving data systems and HIPAA compliance
Skills
Explore related jobs
More jobs at Epsilon Health
Software Engineer - ProductSan Francisco, CA
Software Engineer - BackendSan Francisco, CA
Research Scientist - VLM PretrainingSan Francisco, CA
Research Scientist - Post-training / RLSan Francisco, CA
Research Scientist - Vision Foundation ModelsSan Francisco, CA
Research Engineer - Data Quality & EvalsSan Francisco, CA
Similar Airflow jobs
Jobs in San Francisco
- EFlik Hospitality Group - United Airlines Lounge - Application Links OnlyExternal Flysfo · San Francisco
- EOn Call LeadExternal Flysfo · CA-San Francisco
- ECargo Operations AgentExternal Flysfo · CA-San Francisco
- EAircraft Appearance TechnicianExternal Flysfo · CA-San Francisco
- ESales AssociateExternal Flysfo · CA-San Francisco
- EMechanic BExternal Flysfo · CA-San Francisco
Browse these categories
Market data for software engineer roles
All reports →- Role reportSoftware engineer jobs, Sep 2026: new listings down 9.7%New software engineer listings fell 9.7% to 1,506 this week, the first weekly drop in a month. Prioritize applications while 1,506 fresh roles remain open.
- Salary reportSoftware engineer salary 2026: median midpoint is $190,587The median published range midpoint for US software engineer roles is $190,587, with the middle half from $166,000 to $222,000.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.