Responsibilities
• Lead or support discovery sessions, requirements workshops, architecture discussions, and solution reviews
• Develop and optimize batch and streaming data ingestion pipelines from enterprise applications, databases, APIs, and file-based sources.
• Implement medallion/lakehouse architectures, dimensional models, and data transformation workflows to support analytics and reporting use cases
• Engineer solutions using technologies such as PySpark, Spark SQL, SQL, Python, Delta Lake, and orchestration tools within Azure and Databricks
• Recommend best practices for data modeling, governance, lineage, monitoring, DevOps, and security
Minimum Requirements
• 6+ years of experience in data engineering, data platform development, or cloud data solutions
• 3+ years of hands-on experience with Azure Databricks, Apache Spark, or similar distributed data processing technologies
• Expertise with Microsoft Azure infrastructure and data resources, including Fabric, Azure Data Factory, Synapse Data Analytics, Power BI, Azure SQL, Azure Cosmos DB, and Azure Database for PostgreSQL
• Expertise with Databricks, specifically the ability to design enterprise-level strategy and architecture including Unity Catalog, data warehousing, data sharing, and Mosaic AI
• DevOps for data, GitHub, automated testing, and working with containers (AKS, Docker, registries, etc.)
• Excellent communication skills, ability to clearly explain concepts to teammates and customers, and quickly learn new concepts and technologies
Preferred Qualifications
• Experience building data agents, including NLQ, Databricks Genie, and Fabric Data Agents
• Experience with data management, including data governance, data security, master data management, and familiarity with different industry security requirements
• Broad experience with data/reporting tools, architectures, cloud vendors, and data/AI concepts other than Microsoft