Posted 17 days ago
Senior ML Infrastructure Engineer
AI Summary
Builds and operates high-performance GPU training and inference clusters, designing cloud compute foundations that accelerate scientific ML experimentation at scale.
About this role
Join us at EIT:
At the Ellison Institute of Technology (EIT), we’re on a mission to translate scientific discovery into real world impact. We bring together visionary scientists, technologists, policy makers, and entrepreneurs to tackle humanity’s greatest challenges in four transformative areas:
- Health, Medical Science & Generative Biology
- Food Security & Sustainable Agriculture
- Climate Change & Managing CO₂
- Artificial Intelligence & Robotics
This is ambitious work - work that demands curiosity, courage, and a relentless drive to make a difference. At EIT, you’ll join a community built on excellence, innovation, tenacity, trust, and collaboration, where bold ideas become real-world breakthroughs. Together, we push boundaries, embrace complexity, and create solutions to scale ideas for lab to society. Explore more at www.eit.org
Your Role:
Join our SciComp team to build the cloud and compute foundation that enables scientific breakthroughs. Deliver reliable, secure platforms and self-service guardrails that accelerate experimentation and turn ideas into results - faster, at scale, and with confidence.
Your Responsibilities:
- Build, operate, and continuously optimise our high-performance GPU training and inference clusters, focusing on robust, high-availability scheduling, isolation, and automated lifecycle management.
- Drive systems design and implementation for high-throughput data paths, optimising I/O, caching, and data locality across compute and storage (including our current Lustre implementation).
- Proactively benchmark, profile, and resolve performance bottlenecks across the compute, network, and orchestration layers to maximise efficiency for distributed training and inference.
- Establish comprehensive observability, resilience, and automated security controls to ensure compliance and robust operation of sensitive research environments.
- Partner with Research, Data, and Applied teams to forecast capacity and cost for GPU and storage needs, setting quotas and streamlining ML experimentation pipelines.
Requirements
Essential Skills, Qualifications & Experience:
- Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale
- A proactive, autonomous approach to systems design and the proven ability and desire to ideate, co-create and implement optimal solutions
- Exposure to migrating or transforming ML infrastructure from traditional schedulers to modern, containerised systems
- Expertise with high-throughput storage systems for ML/HPC workloads
- Expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling to resolve bottlenecks
- A solid grasp of IaC and CI/CD practices (e.g., Terraform, Argo CD)
Benefits
We offer the following salary and benefits:
- Competitive salary (dependent on experience) + travel allowance + bonus
- Enhanced holiday. Our annual leave allowance is 25 days plus 8 bank holidays and an additional 3 days between Christmas and New Year. You will also have the opportunity to purchase an additional 5 days annual leave in January and July.
- Pension - Employer contribution 7.5%, minimum employee contribution 5%
- Life Assurance.
- Income Protection
- Private Medical Insurance as standard for you, your partner and any dependents. Including hospital Cash Plan
- Employee discounts
- Electric car scheme
- Nursery Salary Sacrifice scheme
- Cycle to Work Scheme
- Family Planning
- Neurodiversity support including advise and assessments
- Coaching & Therapy services
Working together - what it involves:
You must have the right to work permanently in the UK with a willingness to travel as necessary. In certain cases, we can consider sponsorship, and this will be assessed on a case-by-case basis.
You will live in, or within easy commuting distance of, Oxford (or be willing to relocate) and can commit to being onsite at our Oxford office, a minimum of 3 days per working week.
Skills
Explore related jobs
More jobs at Ellison Institute of Technology
- Executive Assistant - People and EngagementOxford, England
- Senior Software EngineerOxford, England
- Team Coordinator - 12-month Fixed Term ContractOxford, England
- F&B ManagerOxford, England
- Scientist, Plant Cellular Engineering (Metabolic and Synthetic Biology) - PBIOxford, England
- Programming & Events AdministratorOxford, England
Similar Argo CD jobs
Jobs in Oxford
- Senior Front-End Engineer, Quantum ToolsIonQ · Oxford, England
- LBaptist Memorial Hospital Oxford- PRN Certified Registered Nurse Anesthetist (CRNA)Lifelinccorp · Oxford, MS
- LBaptist Memorial Hospital Oxford- Full Time Certified Registered Nurse Anesthetist (CRNA)Lifelinccorp · Oxford, MS
- WSenior Modeler - Construction ModelingWalterpmoore · Oxford, United Kingdom
- WSenior Project Manager/Modeler - Construction ModelingWalterpmoore · Oxford, United Kingdom
- WModeler - Construction ModelingWalterpmoore · Oxford, United Kingdom
Browse these categories
Market data for this role
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.