
Posted Today
Lead Data Engineer
AI Summary
Lead Data Engineer at Prevalent AI designs and delivers advanced data engineering projects for clients, writing code, reviewing team code, and ensuring solutions perform in production with large data volumes and tight processing windows.
About this role
Role Purpose :
As a Lead Data Engineer at Prevalent AI, you will lead the design and delivery of advanced data engineering projects for our clients. You will write code, review the team's code, and stay accountable for how the solution performs in production.
The work involves large data volumes, complex joins and transformations, and tight processing windows. Clients usually come to us when their pipelines are slow, expensive or difficult to change. Working out why, and then fixing it, is a large part of the role.
This role suits engineers who know Spark internals well enough to read the logs and explain what is going wrong, who ask the right questions about the data before they write code, and who can run a small team to a delivery date.
Key Accountabilities :
- Lead the design and delivery of client data engineering projects, from understanding the requirement through to production.
- Design pipelines that meet their SLA and continue to do so as data volumes grow.
- Diagnose failures and performance problems in large distributed jobs using logs and metrics.
- Set the data layout and storage strategy so jobs read only the data they need.
- Build reusable, configuration-driven components that cut development effort across the team.
- Review the team's code and agree the standards the team works to.
- Work with client teams to pin down the requirement and understand the data before development starts.
- Contribute to the platform work around delivery, including object storage, containerisation and CI/CD.
- Mentor developers, and document design decisions, tuning approaches and operational runbooks.
Skills & Experience :
Must Have :
- 10+ years in data engineering, with at least 5 years hands-on PySpark on large-volume workloads, and still writing code daily.
- Deep understanding of Spark internals: execution model, stages and tasks, shuffle, partitioning, broadcast versus sort-merge joins, adaptive query execution, and memory and spill behaviour.
- Experience tuning long-running jobs down to an SLA and handling data skew, with specific examples of how a job was diagnosed and what changed.
- Hands-on experience with time-series data at scale: event-time semantics, sliding and tumbling windows, and as-of or point-in-time joins across billions of rows.
- Strong Delta Lake or Iceberg experience, covering partitioning, compaction, sort keys, data pruning and versioning.
- Able to read Spark logs, the Spark UI and the history server and work out why a job is slow or failing.
- Experience running Spark on self-managed clusters, not only on managed platforms.
- Strong Python, with experience building reusable, configuration-driven frameworks rather than one-off scripts.
- Able to lead a small team through code reviews, mentoring and unblocking, and hold a delivery date.
Good to Have
- Trade surveillance, capital markets or FX experience, and familiarity with trade, order and market data.
- KDB/q exposure, or experience migrating off a vendor surveillance platform.
- Experience moving Spark workloads to object storage such as MinIO or S3, and running Spark on Kubernetes.
- CI/CD, automated testing and observability for data pipelines.
- Exposure to regulatory or compliance-driven delivery in a banking environment.
Skills
Explore related jobs
More jobs at Prevalent AI
Similar Apache Iceberg jobs
Browse these categories
Market data for data engineer roles
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.