Member of Technical Staff- Data Intelligence
AI Summary
Collaborates with researchers and engineers to build high-quality, petabyte-scale data pipelines and infrastructure, defining data quality metrics and automating assessments for model training and evaluation.
About this role
In this role, you’ll work closely with model researchers, data infrastructure engineers, and cross-functional partners to make sure our data is high quality and can be produced at petabyte scale in a reliable, efficient way. From understanding how data choices show up in model behavior, to building processing pipelines and running the compute behind them, you’ll help ensure our models are trained on the best data we can get.
What you’ll do
Work with model researchers to define what “good data” means for our models, including quality metrics, validation checks, and acceptance thresholds
Explore open source datasets and create internal ones most suitable to build fundamental World Models
Build algorithms for automated data quality assessment, data domain mixtures, and domain adaptation from synthetic to real data.
Track datasets, metadata, provenance, and versions so experiments are reproducible and it’s clear what data went into which training and evaluation runs
Own CI/CD and development tooling for the data stack (GitHub, Python, PyTorch), and automate repetitive workflows to reduce friction
Track and optimize throughput, storage, and compute utilization across pipelines and related assets
What we’re looking for
Strong ML and deep learning fundamentals with experience building and operating large-scale data and/or compute systems
Comfortable moving between research questions and production engineering: you can dig into data, run analyses, and also ship reliable systems
Demonstrated research experience with data compositions, quality, and dataset releases
Ability to design and execute experiments with convincing unbiased outcomes
Practical experience with distributed processing and orchestration (Spark, Ray, Airflow, or equivalents)
Solid Python skills, and familiarity with the tooling around modern model training workflows (datasets, checkpoints, experiment tracking)
Strong instincts around data quality: how to measure it, how to monitor it, and how to prevent regressions as things scale
Able to work in a fast-moving environment, prioritize what matters, and communicate clearly with both researchers and engineers
Bonus: experience with large video datasets, dataset curation for training, or building internal tooling for evaluation/analysis in ML environments
Skills
Explore related jobs
More jobs at Reka
Similar Airflow jobs
Jobs in Singapore
Purchasing Team Lead | Material Readiness Planner | NTS SingaporeNts · Singapore, Central Singapore- Compliance Officer & DMLROEbury · Singapore
- LDirector, Cost ManagmentLinesight · Singapore, SG
- LAssociate Director, Cost Management (Data Centre Projects)Linesight · Singapore, SG
- AAdvanced Engineer, Application Development-Technical SalesAsiacelanese · Singapore, SG
- VOptical Source InspectorVe · Singapore, SG
Browse these categories
Market data for this role
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.
