Jobless Developer
FieldAI logo

Posted 10 days ago

Open

Senior Data Platform Engineer

IrvineOn-siteFull-time

AI Summary

Design and build the data platform that manages multimodal robotics data from ingestion through ML training, including scalable pipelines, validation, lineage, and versioned dataset generation.

About this role

FieldAI’s Irvine team is where embodied AI meets real robots, real sensors, and real field deployments. Based in the heart of Southern California’s robotics ecosystem, we build risk-aware, reliable, field-ready AI systems that solve the hardest problems in robotics and unlock the full potential of embodied intelligence. If you want your work to ship, get tested on hardware, and improve through real deployments, Irvine is the place. We go beyond typical data-driven approaches or pure transformer-only architectures, combining rigorous engineering with learning systems proven in globally deployed solutions that deliver results today and get better every time our robots run in the field.

About the Role

We are building the data foundation that powers the full machine learning lifecycle for autonomous robotics. Our robots generate large-scale, multimodal datasets across real-world deployments, and turning that raw experience into reliable, discoverable, high-quality ML data is a core part of improving our autonomy systems.

As a Data Platform Engineer, you will help design and build the platform that manages data from ingestion through processing, validation, labeling, dataset generation, training, and evaluation.

This is not a traditional analytics data engineering role. You will work closely with ML engineers, researchers, labeling teams, robotics engineers, and infrastructure engineers to build scalable systems for robotics and ML data.

What You'll Do

  • Design and build scalable data architecture for large-scale multimodal robotics and ML datasets.
  • Build abstractions and services for ingestion, processing, datasets, metadata, lineage, and data quality.
  • Develop reliable pipelines for transforming raw robot data into versioned, ML-ready data products.
  • Define data models, schemas, contracts, and lifecycle states across data-processing workflows.
  • Build systems for tracking provenance and lineage across raw data, derived artifacts, labels, datasets, and downstream ML workloads.
  • Develop automated validation and data-quality frameworks that detect incomplete, corrupted, or unusable data early.
  • Design for incremental processing, reprocessing, backfills, and versioned transformations.
  • Improve observability and failure diagnosis across complex data workflows.
  • Partner with ML and robotics teams to understand domain-specific data requirements and turn recurring patterns into reusable platform capabilities.
  • Work closely with infrastructure/platform teams on storage, compute, orchestration, reliability, and scalability.
  • What We're Looking For

  • Strong experience building production data platforms or large-scale data-processing systems.
  • Strong software engineering skills, preferably Python and/or C++/Java/Go.
  • Experience with distributed data processing and workflow orchestration.
  • Experience with data lakes/lakehouses, object storage, metadata systems, schemas, and data versioning.
  • Strong understanding of data quality, lineage, reproducibility, and reliable pipeline design.
  • Experience with technologies such as S3, Airflow/Dagster, Spark/Ray, Kubernetes, Parquet, or similar systems.
  • Ability to work across ambiguous organizational and technical boundaries.
  • Strong systems-design and engineering judgment.
  • Nice to Have

  • Experience with ML datasets or ML infrastructure.
  • Robotics, autonomous vehicles, sensor, video, image, LiDAR, or other multimodal data.
  • Experience building internal developer/platform products.
  • Experience operating pipelines at TB/PB scale.
  • Skills

    AirflowC++DagsterData LakesData LineageData Quality FrameworksData VersioningDistributed Data ProcessingGOJavaKubernetesLakehousesMetadata SystemsObject StorageParquetPythonRayS3SparkWorkflow Orchestration

    Explore related jobs

    Browse these categories

    Market data for devops / sre / platform engineer roles

    All reports →