
Posted 3 months ago
Developer
AI Summary
Develops and maintains scalable data pipelines using Python and PySpark; designs ETL processes, optimizes Spark applications, and collaborates with cross-functional teams to implement data solutions.
About this role
Must-Have**
Strong proficiency in Python programming.
Hands-on experience with PySpark and Apache Spark.
Knowledge of Big Data technologies (Hadoop, Hive, Kafka, etc.).
Experience with SQL and relational/non-relational databases.
Familiarity with distributed computing and parallel processing.
Understanding data engineering best practices.
Experience with REST APIs, JSON/XML, and data serialization.
Exposure to cloud computing environments.
5+ years of experience in Python and PySpark development.
Experience with data warehousing and data lakes.
Knowledge of machine learning libraries (e.g., MLlib) is a plus.
Strong problem-solving and debugging skills.
Excellent communication and collaboration abilities.
Requirements
Develop and maintain scalable data pipelines using Python and PySpark.
Design and implement ETL (Extract, Transform, Load) processes.
Optimize and troubleshoot existing PySpark applications for performance.
Collaborate with cross-functional teams to understand data requirements.
Write clean, efficient, and well-documented code.
Conduct code reviews and participate in design discussions.
Ensure data integrity and quality across the data lifecycle.
Integrate with cloud platforms like AWS, Azure, or GCP.
Implement data storage solutions and manage large-scale datasets.
Skills
Explore related jobs
More jobs at GSB Solutions
Sr. SRE | Hybrid in Mexico CityMexico, México
Advisor de seguridad | Hybrid in Mexico CityCDMX, Ciudad de México
Technical Support Engineer L3 (México-CDMX)CDMX, Ciudad de México
Application Security Engineer – Threat Modeling & Secure SDLC | Hybrid in Mexico CityCDMX, Ciudad de México
Commissioning resource | RemoteRemote Job
Oracle Fusion cloud HCM consultantHyderabad, Andhra Pradesh
Similar Apache Spark jobs
- Software Engineer 2 (Hybrid) - Linux/Bash/Apache Airflow/Apache Spark/Docker/Podman/Git/AWSCaptivation Software · Annapolis Junction, MD - Hybrid
- Senior Data Engineer – Apache Spark | Kafka | Flink | Trino | Iceberg | Big Data | Streaming | Data Platform 8–12 YearsCisco Systems, Inc. · Bangalore, India
Software Engineer, Spark PlatformDoorDash USA · San Francisco, CA
Jobs in Hyderabad
- Job Posting Title Senior Software Engineer - Liquidity Book- Java FullstackFactSet Hong Kong Limited · Hyderabad, IND
- Enterprise Recruitment RepresentativeApplyBoard India Private Limited · Hyderabad
- Senior Software Engineer - Liquidity Book- Java FullstackFactSet Hong Kong Limited · Hyderabad, IND
- Senior Data Recoverability Engineer (Data Governance)FactSet Hong Kong Limited · India, Hyderabad
- Associate scientistLaboratorio de Control ARJ, S. A. de C. V. · Hyderabad, India
Senior Data ScientistDazn · India - Hyderabad
Browse these categories
Market data for this role
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.