
Posted 10 days ago
Data Engineer
Remote
AI Summary
Designs, develops, and optimizes scalable data pipelines; builds ETL/data ingestion workflows; writes PySpark, Python, and SQL; integrates MongoDB and relational sources; ensures data quality and performance.
About this role
About the Role
We are looking for a Data Pipeline Engineer with strong expertise inDatabricks, AWS Glue, Python, SQL, and MongoDB to design, develop, and optimize scalable data pipelines. The ideal candidate should have experience building cloud-native ETL solutions, transforming large datasets, and delivering high-quality data for analytics and business applications.
Key Responsibilities
- Develop and maintain scalable ETL/data pipelines usingDatabricks andAWS Glue.
- Build data ingestion and transformation workflows for structured and semi-structured data.
- Write efficientPySpark, Python, and SQL code for large-scale data processing.
- Integrate data from multiple sources, includingMongoDB and relational databases.
- Implement data quality, validation, and reconciliation checks.
- Optimize pipeline performance, reliability, and scalability.
- Collaborate with architects, analysts, and application teams to deliver data solutions.
- Support production deployments, monitoring, and troubleshooting.
Required Skills
- 5+ years of experience in Data Engineering or ETL development.
- Strong hands-on experience withDatabricks andPySpark.
- Experience withAWS Glue and AWS data services.
- Proficiency inPython andSQL.
- Experience working withMongoDB.
- Knowledge of ETL, data modeling, and data warehousing concepts.
- Familiarity with Delta Lake/Lakehouse architecture is preferred.
- Experience with Git and Agile development practices.
Good to Have
- Delta Lake
- Apache Spark optimization
- AWS S3
- CI/CD pipelines
- Healthcare or Benefits domain experience
Skills
AgileAWS GlueDatabricksDelta LakeETLGitMongoDBPySparkPythonSQL