Level: Mid-level to senior, based on the required 5+ years of hands-on Azure Databricks and Spark experience.
Summary:
This role is focused on building and supporting large-scale data pipelines using Azure Databricks, Spark, Python/PySpark, and SQL. The candidate should have strong hands-on engineering experience with production data processing, data lake architecture, ETL, data quality, and Azure analytics services. This is a technical delivery role that also requires strong documentation and communication skills.
Main Responsibilities:
- Build and implement data pipelines using Azure Databricks and Spark.
- Write optimized Python/PySpark and SQL code for very large, TB-scale data processing.
- Design and maintain data lake architecture and ETL workflows.
- Manage data quality, integrity, and validation.
- Use Azure data services for storage, processing, and analytics.
- Create technical documentation, requirements, and testing documents.
- Collaborate with business and technical teams to understand requirements and deliver solutions.
- Use Git for version control.
Must-Have Skills:
- 5+ years of hands-on experience with Azure Databricks and Spark.
- Strong Python/PySpark and SQL coding skills.
- Experience optimizing performance for TB-scale data processing.
- Strong understanding of data lake architecture and ETL processes.
- Experience with Azure data analytics services.
- Data quality management experience.
- Strong technical documentation skills.
- Ability to clearly communicate technical concepts.
- Git experience.
Most Important Fit Criteria:
The strongest candidates will have production-level Azure Databricks and Spark experience, strong PySpark and SQL skills, and a proven ability to build and optimize large-scale data pipelines. Prioritize candidates who have worked with TB-scale data, data lakes, ETL, and data quality in Azure environments.
Nice-to-Have:
- Experience working across cross-functional teams.
- Strong requirements gathering and testing documentation experience.
- Prior experience supporting production data platforms.