Jobless Developer
TOPPAN MERRILL LLC logo

Posted 12 months ago

Open

Senior Data Engineer

ChennaiOn-siteFull-time

AI Summary

Builds and optimizes distributed data pipelines on Azure Databricks using PySpark, implements Unity Catalog governance, and integrates with cloud data platforms to deliver scalable, secure data solutions.

About this role

Job Description:

Responsibilities:

  • Develop & Optimize Data Pipelines
    • Build, test, and maintainETL/ELTdata pipelines using AzureDatabricks & Apache Spark (PySpark).
    • Optimizeperformance and cost-efficiencyof Spark jobs.
    • Ensure data quality through validation, monitoring, and alerting mechanisms.
    • Understand cluster types, configuration, and use-case for serverless
  • Implement Unity Catalog for Data Governance
    • Design and enforceaccess control policiesusing Unity Catalog.
    • Managedata lineage, auditing, and metadata governance.
    • Enable secure data sharing across teams and external stakeholders.
  • Integrate with Cloud Data Platforms
    • Work withAzure Data Lake Storage / Azure Blob Storage/ Azure Event Hubto integrate Databricks with cloud-baseddata lakes, data warehouses, and event streams.
    • ImplementDelta Lakefor scalable, ACID-compliant storage.
  • Automate & Orchestrate Workflows
    • DevelopCI/CDpipelines for data workflows usingAzureDatabricks Workflows or Azure Data Factory.
    • Monitor and troubleshoot failures injob execution and cluster performance.
  • Collaborate with Stakeholders
    • Work withData Analysts, Scientists, and Business Teamsto understand requirements.
    • Translate business needs intoscalable data engineering solutions.
  • API expertise
    • Ability to pull data from a wide variety of APIs using different strategies and methods

Required Skills & Experience:

  • Azure Databricks & Apache Spark (PySpark)– Strong experience in buildingdistributed data pipelines.
  • Python– Proficiency in writing optimized and maintainable Python code for data engineering.
  • Unity Catalog– Hands-on experience implementingdata governance, access controls, and lineage tracking.
  • SQL– Strong knowledge of SQL for data transformations and optimizations.
  • Delta Lake– Understanding oftime travel, schema evolution, and performance tuning.
  • Workflow Orchestration– Experience withAzureDatabricks Jobs or Azure Data Factory.
  • CI/CD & Infrastructure as Code (IaC)– Familiarity with Databricks CLI, Databricks DABs, and DevOps principles.
  • Security & Compliance– Knowledge ofIAM, role-based access control (RBAC), and encryption.

Preferred Qualifications:

  • Experience withMLflowfor model tracking & deployment in Databricks.
  • Familiarity withstreaming technologies(Kafka, Delta Live Tables, Azure Event Hub, Azure Event Grid).
  • Hands-on experience withdbt (Data Build Tool)for modular ETL development.
  • Certification inDatabricks, Azure is a plus.
  • Experience with Azure Databricks Lakehouse connectors for SalesForce and SQL Server
  • Experience with Azure Synapse Link for Dynamics, dataverse
  • Familiarity with other data pipeline strategies, like Azure Functions, Fabric, ADF, etc

Soft Skills:

  • Strongproblem-solvingand debugging skills.
  • Ability towork independently and in teams.
  • Excellentcommunication and documentationskills.

Skills

Apache SparkAzure Blob StorageAzure DatabricksAzure Data FactoryAzure Data Lake StorageAzure Event HubAzure FunctionsAzure Synapse LinkCI/CDDatabricks CLIDatabricks DABsDbtDelta LakeIAMKafkaMLflowPySparkPythonRBACSQLUnity Catalog

Explore related jobs

Browse these categories

Market data for data engineer roles

All reports →