Group Lead ML Infrastructure
AI Summary
Group Lead ML Infrastructure responsible for owning and scaling the ML platform (training framework, data pipelines, compute efficiency, orchestration, and tooling) to accelerate deep learning research and reliably move prototypes into production.
About this role
We are looking for a ML Infra Group Lead to drive the ML infrastructure that powers our deep learning research and trading. This is one of the most important roles in the company. As our DL strategies become an increasingly central driver of the business each year, the team you will lead will scale its impact across all major markets and asset classes — by enabling our researchers to iterate faster and bring their ideas to production reliably.
While the DL Research Group Lead drives the research agenda, you will own the platform side: the training framework, dataset pipelines, compute efficiency, and release tooling that turn research prototypes into production-grade ML models.
Responsibilities
- Drive the development of our internal ML platform — training framework, dataset and data pipelines, orchestration, experiment tracking and reproducibility tooling, GPU/compute infrastructure, and release tooling — that powers our DL strategies across all markets;
- Own end-to-end performance of large-scale model training: GPU cluster utilization, MFU, training time, reliability, and reproducibility of models from master;
- Shape the technical agenda of the team: identify bottlenecks, prioritize improvements, and deliver measurable infrastructure wins (e.g., a 10% training speedup on our cluster translates into millions of euros of compute value);
- Lead engineers across diverse profiles: hire, motivate, develop, and retain talent, including ML performance engineers, Python infrastructure engineers, and tool builders;
- Make the platform fast, transparent, and value-driven for researchers — shorten iteration cycles and remove infra friction from research workflows;
- Directly code and contribute to the ML platform, not just supervise;
- Partner closely with DL researchers on the Research/Infra interface — code review for shared training and dataset code, debugging non-reproducible models, observability of data and model quality, and release readiness;
Requirements
- A motivated, versatile ML/infra engineer with experience in technical leadership of teams that build platforms for ML researchers;
- Strong background in at least one of: large-scale ML training systems, GPU performance engineering (PyTorch, Triton, CUDA), distributed training, data and ML pipeline infrastructure (storage formats, orchestration, experiment tracking, configuration);
- Hands-on experience accelerating ML research in production or large-scale research settings — reducing training time, improving hardware utilization, or shortening feedback cycles for researchers (e.g., faster data access, better experiment tooling, more reliable pipelines);
- Strong sense of ownership: comfortable taking responsibility for code quality, reproducibility, and observability across a large shared ML codebase — not just reactive fixes, but proactive bug hunting, tech-debt reduction, and rigorous code review;
- A leader who can sustain a strong pace while fostering a healthy, supportive team culture, and who can give clear, growth-oriented feedback;
- Ability to work effectively at the DL/Infra boundary — engaging with research priorities and translating them into infrastructure work that compounds over time;
- While prior trading experience can be useful, it's not a prerequisite. Our priority is a first-principles mindset and a willingness to rethink how things are done in quantitative ML infrastructure.
What we offer
- High base salary and social benefits;
- Generous bonus structure. We are very flexible in discussing salary and conditions of employment;
- Cutting-edge hardware and software in production as well as high technical expertise of the company which allows implementation of bold ideas and boosting great results. Ownership over initiatives that directly solve business problems;
- Ability to trade on dozens of international exchanges;
- Flexible workflow (lack of formalism and bureaucracy, no pressure and over-management) and working schedule;
- Tuition reimbursement, conference and training sponsorship.
Skills
Explore related jobs
More jobs at Pinely
- Execution Research GLAmsterdam, North Holland
- Senior Research Engineer (ML infrastructure )Amsterdam, North Holland
- Senior Deep Learning ResearcherAmsterdam, North Holland
- Senior Quantitative ResearcherAmsterdam, North Holland
- Quantitative Research (Intern)Amsterdam, North Holland
- ML Performance EngineerAmsterdam, North Holland
Similar Compute Infrastructure jobs
Jobs in Amsterdam
- NHost - Madame Tussauds AmsterdamNl Merlinentertainments · Amsterdam, Netherlands
- IHeavy Equipment Operator - DozerInterstatewaste · Amsterdam, OH
- ILaborerInterstatewaste · Amsterdam, OH
Accounts Payable InternHotel Co 51 · Amsterdam, Noord-Holland
Cook On Call - Italian Restaurant - Courtyard AmsterdamHotel Co 51 · Amsterdam, Noord-Holland
Technical Support Specialist (German)VanMoof · Amsterdam, Netherlands
Browse these categories
Market data for this role
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.