About the Role
We are seeking a highly skilled Data Engineer with strong expertise in Azure Databricks, SQL, PySpark, and Data Modeling. The ideal candidate will have experience designing and implementing scalable data pipelines, optimizing data workflows, and building modern data platforms on the Azure ecosystem.
*Key Responsibilities
Design, develop, and maintain ETL/ELT pipelines using Azure Databricks & PySpark.
Build and manage Delta Lakehouse solutions including Bronze, Silver, Gold layers.
Collaborate with data architects, analysts, and business stakeholders to design data models (Star/Snowflake schemas, Fact & Dimension tables).
Optimize Databricks clusters, jobs, and queries for performance and cost-efficiency.
Implement CI/CD pipelines for Databricks notebooks and data workflows.
Manage schema evolution, data governance, and quality checks across pipelines.
Work with Azure Data Lake Storage (ADLS), Azure Synapse Analytics, and SQL Databases for end-to-end data solutions.
Implement data partitioning, caching, and broadcast joins to optimize PySpark jobs.
Ensure best practices in data security, compliance, and access management.
Troubleshoot and optimize slow SQL queries, indexes, and data warehouse performance.
Support business reporting and analytics needs by designing and maintaining scalable data models.
*Required Skills
Azure Databricks: Notebooks, Delta Tables, Auto-scaling, Job Orchestration.
SQL: Joins, Window Functions, Indexing, Query Optimization, CTEs, SCD handling.
PySpark: RDD, Data Frame API, Lazy evaluation, Transformations, Optimizations.
Data Modeling: OLTP vs OLAP, Star & Snowflake Schema, Fact/Dimension Tables, Normalization/Denormalization.
Azure Ecosystem: ADLS, Synapse Analytics, Azure Data Factory (ADF) is a plus.
Strong understanding of ETL best practices, data quality frameworks, and large-scale distributed data processing.
*Nice-to-Have Skills*
Experience with DataBricks CI/CD pipelines (Azure DevOps/GitHub).
Familiarity with Power BI/Tableau reporting dashboards.
Knowledge of Kafka, Event Hub, or real-time streaming.
Exposure to machine learning pipelines within Databricks.
*Qualifications
Bachelor’s or Master’s degree in Computer Science, Information Technology, or related field.
6-7 years of experience in Data Engineering.
Proven track record of building scalable data solutions using Azure Databricks and PySpark.
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.