Lead Databricks Data Engineer to lead the data platform and MLOps workstream for a large U.S. enterprise client. You will own a governed Databricks lakehouse, production-grade pipelines across multiple enterprise sources, trusted feature and prediction data products, and the engineering path that moves machine-learning workloads from experimentation into reliable production operations.
This is a technical leadership role—not a coordination-only position. You will write and review code, establish platform standards, guide engineers, and work directly with client data science and IT leaders to make and defend architecture decisions. The role is remote within the United States, with occasional travel to client locations for key workshops, planning sessions, and delivery milestones.
Key Responsibilities
• Lead the Databricks data platform workstream; set technical direction, sequence delivery, and own outcomes with client data science and IT stakeholders.
• Design and build production medallion architecture (Bronze, Silver, and Gold) using Databricks, Apache Spark, Delta Lake, and Delta Live Tables or Lakeflow Declarative Pipelines.
• Engineer reliable batch and, where needed, streaming pipelines that integrate operational, CRM, HRIS/payroll, web, finance, and other enterprise sources.
• Implement Unity Catalog governance, lineage, RBAC, service principals, secrets management, environment isolation, and least-privilege controls.
• Build reusable feature, training, and prediction tables with versioning and retained history so models are reproducible and retrainable.
• Establish CI/CD and Databricks Asset Bundles for data and ML workloads, including automated testing, environment parameterization, approvals, and rollback.
• Own orchestration through Databricks Workflows/Jobs, including dependencies, retries, recovery, observability, alerting, and pipeline SLAs.
• Build the MLOps layer with MLflow experiment tracking, model registry/versioning, batch inference, model promotion, monitoring, and rollback.
• Implement data-quality and observability controls that detect schema drift, null spikes, freshness failures, and out-of-range values before downstream impact.
• Tune Spark jobs, clusters, SQL warehouses, storage layout, and Delta tables for performance, reliability, and cost efficiency.
• Enable governed access from Python and R notebooks, BI tools such as Power BI, and downstream APIs.
• Mentor engineers, conduct code and design reviews, produce architecture decisions and runbooks, and keep delivery unblocked.
Required Skills & Experience
• 6+ years of data engineering or data platform experience, including technical leadership or end-to-end ownership of a production platform.
• At least 3 years of recent, hands-on Databricks experience in production environments. Candidates whose primary experience is only Snowflake, Microsoft Fabric, BigQuery, or another platform will not meet this requirement.
• Demonstrated production expertise with Apache Spark/PySpark, Delta Lake, Databricks Workflows/Jobs, Unity Catalog, and MLflow.
• Strong Python and SQL skills, including complex transformations, performance tuning, testing, and maintainable production code.
• Experience building scalable ETL/ELT pipelines with incremental processing, schema evolution, change data capture, backfills, and dimensional/data-product modeling.
• Hands-on CI/CD for Databricks using Git and automated deployment patterns; experience with Databricks Asset Bundles, Terraform, or equivalent infrastructure-as-code.
• Experience implementing data quality, lineage, observability, security, and environment promotion across development, test, and production.
• Working knowledge of a major cloud platform; Azure experience is preferred, but AWS or GCP is acceptable when paired with strong Databricks depth.
• Ability to lead client working sessions, document decisions clearly, explain tradeoffs to senior stakeholders, and remain accountable for technical outcomes.
• Excellent written and verbal communication and the ability to operate effectively with distributed U.S. and offshore teams.
• Ability and willingness to travel occasionally within the United States based on client and project needs.
Preferred Qualifications
• Databricks Data Engineer Professional, Databricks Machine Learning Professional, or comparable certification.
• Experience with Azure Databricks, ADLS Gen2, Azure DevOps, Entra ID, Key Vault, Event Hubs, and Power BI.
• Experience with Structured Streaming, Kafka/Event Hubs, Change Data Feed, Auto Loader, and low-latency data or inference patterns.
• Machine-learning engineering experience spanning feature engineering, model deployment, batch scoring, monitoring, and drift detection.
• Familiarity with dbt, Airflow, R-based data science workflows, semantic layers, and enterprise API integration.
• Prior consulting experience delivering data platforms for large U.S. enterprise clients.
• Experience using AI-assisted engineering tools responsibly to accelerate implementation, testing, documentation, and code review.
Recruiter Screening Requirements
Please submit candidates only when all of the following are confirmed:
• Direct, recent Databricks production experience—not training-only, certification-only, or adjacent-platform experience.
• Hands-on depth in PySpark, Delta Lake, Workflows/Jobs, Unity Catalog, and MLflow, with clear examples of what the candidate personally designed and built.
• Currently located in the United States and able to work remotely during U.S. business hours and travel occasionally.
• Interested in direct W-2 employment with NRnP Technology. No C2C candidates or third-party submissions.
How to Apply
Send a résumé and a brief summary of relevant Databricks implementations to [email protected]. Please include the candidate’s current location, work authorization, availability, and willingness to travel occasionally.
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.