Posted 2 days ago
AI Data Engineer
AI Summary
Designs, builds, and maintains scalable data pipelines and architectures that power AI, ML, and LLM applications, and migrates legacy data platforms to modern, governed, AI-ready infrastructures.
About this role
Position Purpose:
We are seeking a passionate and motivated AI Data Engineer to join our international team as part of a strategic expansion of our AI and Data capabilities.
In this role, you will be responsible for designing, building, and maintaining the data foundations that power foundational applications, advanced analytics, machine learning, Generative AI, and Large Language Model (LLM) applications across the organization. You will work at the intersection of data engineering, AI platform development, and data governance, ensuring that high-quality, trusted, and scalable data is available to support AI-driven innovation.
Your work will play a critical role in enabling next-generation AI solutions by building robust data pipelines, optimizing large-scale data architectures, and creating AI-ready datasets that support model training, retrieval-augmented generation (RAG), semantic search, and intelligent automation initiatives.
As an AI Data Engineer, you will collaborate closely with Data Engineers, Machine Learning Engineers, Data Scientists, Data Architects, Product teams, and business stakeholders to ensure that data assets are reliable, accessible, secure, and aligned with both business and technical requirements.
What You Will Be Doing
Design, develop, and maintain scalable data ingestion, transformation, and enrichment pipelines supporting AI and analytics initiatives.
Assess, modernize, and migrate legacy data platforms, tools, and operational processes to scalable, governed, and AI-ready data infrastructures.
Lead initiatives to replace manual, siloed, or obsolete data workflows with automated, cloud-enabled, and data-driven solutions.
Design and execute migration strategies for data assets stored in legacy applications, databases, file systems, and on-premise environments while ensuring data quality and business continuity.
Refactor and optimize existing ETL/ELT pipelines, data models, and integration processes to align with modern architectural standards and AI platform requirements.
Develop data processing solutions using Python, Spark, SQL, and distributed computing frameworks.
Design and implement data architectures supporting LLM-based applications, Retrieval-Augmented Generation (RAG), semantic search, and knowledge management platforms.
Develop and manage metadata, lineage, and cataloging solutions to improve data discoverability and governance.
Collaborate with Machine Learning Engineers and Data Scientists to provide high-quality feature stores, training datasets, and inference data pipelines.
Support the creation and maintenance of Data-as-a-Service (DaaS) platforms and reusable data products.
Implement automated data quality checks, validation frameworks, and monitoring solutions to ensure data integrity and reliability.
Contribute to AI platform engineering initiatives, including model-serving infrastructure, feature management, and AI data lifecycle management.
Ensure compliance with data privacy regulations, security standards, and organizational governance policies.
Define and implement best practices to ensure migrated systems comply with enterprise standards for data governance, security, lineage, observability, and documentation.
Participate in the design of AI governance processes related to data sourcing, quality management, access control, and responsible AI.
Collaborate with business stakeholders to identify opportunities for process automation, technical debt reduction, and platform modernization.
Stay informed on emerging technologies in data engineering, Generative AI, knowledge graphs, vector search, and AI platform architectures.
Participate in agile ceremonies including sprint planning, solution design reviews, architecture discussions, and continuous improvement initiatives.
What You Need for this Position
Bachelor's or Master's degree in Computer Science, Data Science, or a related field.
Proven at least 5 years of working experience as a Big Data Engineer or similar.
Strong proficiency in programming languages such as Scala, Python, or Java.
Familiarity with APIs, data integration frameworks, and event-driven architectures.
At least 5 years of experience in migrating legacy data platforms, repositories, and operational workflows to modern data architectures while ensuring business continuity and data integrity.
Building data solutions using Python, SQL, Spark, and distributed computing frameworks experience.
Knowledge of data security, encryption, access controls, and compliance requirements.
Understanding of modern data architectures, including data lakes, lakehouses, and data mesh concepts.
Familiarity with AI and ML workflows, including training data preparation, feature engineering, and model-serving data pipelines
Experience with Git, CI/CD pipelines, and infrastructure-as-code practices.
Strong analytical and problem-solving capabilities, with the ability to navigate complex and ambiguous challenges.
Passion for AI, data innovation, and continuous improvement.
Ability to assess existing processes and identify opportunities for automation, modernization, and operational efficiency.
Excellent communication and stakeholder management skills.
Ability to prioritize tasks, manage multiple projects simultaneously, and deliver high-quality results within defined timelines.
Ability to collaborate effectively across Data Engineering, Data Science, AI, Product, Architecture, and Business teams.
Strong organizational skills and attention to detail.
Comfortable working in fast-paced, multicultural, and international environments.
Curiosity and the ability to rapidly learn new technologies, platforms, and engineering practices.
The ability to thrive in a fast-paced environment and take on new responsibilities quickly.
Skills
Explore related jobs
More jobs at DealerSocket (India) Pvt. Ltd.
Similar Apache Spark jobs
- SBig Data Engineer – Java, Apache Spark, AWS EMR, Lambda, EKS & AirflowSynechron, Inc._USA Company · Chennai - Taramani (Ascendas)
- CFull-stack Engineer 4 (Apache Spark and Java)( Enterprise Platforms Technology)capital one · McLean, Virginia
- CSoftware Engineer 2 (Hybrid) - Linux/Bash/Apache Airflow/Apache Spark/Docker/Podman/Git/AWSCaptivation Software · Annapolis Junction, MD - Hybrid
Jobs in Prague
- ISenior Cyber Security EngineerIntervet Inc. USA · CZE - Central Bohemian - Prague (IT Riverview)
- AIntern Analyst - SEE ConsultingARG IQVIA RDS Argentina SRL · Prague, Czech Republic
- AData Collector - part timeARG IQVIA RDS Argentina SRL · Prague, Czech Republic
- AJob Without JobARG IQVIA RDS Argentina SRL · Prague, Czech Republic
- ARecruiter, Medical Sales & Medical AffairsARG IQVIA RDS Argentina SRL · Prague, Czech Republic
- CSenior People Solutions Analystcollibra · Prague, Czech Republic
Browse these categories
Market data for data engineer roles
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.