Your mission
Data pipelines and integration
• Design, build and operate scalable batch, micro-batch and streaming data pipelines using Python, PySpark and Azure Databricks.
• Integrate internal and external data sources, including REST APIs, GraphQL, WebSocket and gRPC interfaces, databases, files, event streams and third-party data feeds.
• Develop robust ingestion solutions for structured, semi-structured and unstructured data, including JSON, CSV, Parquet, Delta and API-based payloads.
• Build reliable web scraping and data acquisition components where APIs or managed integration mechanisms are unavailable.
• Implement pagination, throttling, retries, exponential backoff, checkpointing, schema evolution and recovery patterns for external integrations.
• Design pipelines that support idempotent processing, reprocessing, controlled backfills and graceful recovery from partial failures.
• Develop and maintain batch and streaming patterns for market, weather, fundamental and time-series data.
Software engineering
• Develop modular, reusable and testable Python and PySpark components rather than relying on monolithic notebooks.
• Apply object-oriented and functional design principles appropriately to data-focused development.
• Structure solutions as maintainable software projects with clear separation between source code, configuration, tests, deployment assets and notebooks.
• Write clean, readable and well-documented code using type hints, meaningful interfaces and appropriate design patterns.
• Build automated unit, integration, contract and data-quality tests and incorporate them into delivery pipelines.
• Conduct code reviews and promote engineering standards covering readability, testability, security, performance and maintainability.
• Package reusable functionality as Python modules or wheels where appropriate.
• Troubleshoot complex issues across source systems, APIs, processing logic, infrastructure and production runtime environments.
Data architecture and modelling
• Design maintainable data models that support analysts, traders, reporting solutions and downstream data products.
• Implement Lakehouse and Medallion architecture patterns across Bronze, Silver and Gold layers.
• Preserve raw data appropriately while applying cleansing, validation, standardisation and business transformations in downstream layers.
• Design solutions for schema evolution, data retention, lineage and reproducible processing.
• Apply sound data architecture principles across operational, analytical, event-based and time-series workloads.
• Optimise data layouts, partitioning, joins, file sizes, caching and Spark execution plans for performance and cost.
Orchestration and DataOps
• Design, schedule and operate workflows using Databricks Workflows and Astronomer.
• Implement dependency management, parameterisation, environment-specific configuration and controlled promotion across development, test and production environments.
• Define operational runbooks and support effective diagnosis, recovery and problem management.
• Monitor pipeline health, freshness, completeness, performance and data-quality indicators.
• Use production-safe release patterns, including controlled rollouts, rollback and validation where appropriate.
GitOps, CI/CD and Infrastructure as Code
• Manage all production code through Git using clear branching, pull-request and review practices.
• Build and maintain automated CI/CD pipelines using GitHub Actions and/or Azure DevOps.
• Deploy Databricks jobs, pipelines and application artefacts using Databricks Asset Bundles or equivalent approved mechanisms.
• Provision and configure relevant cloud and Databricks resources through Terraform.
• Treat application code, infrastructure, data pipeline definitions and operational configuration as version-controlled artefacts.
• Apply automated validation, security scanning and testing before production deployment.
• Contribute to reusable pipeline templates, engineering standards and platform automation.
Reliability, performance and cost efficiency
• Engineer solutions for availability, recoverability, scalability and predictable operational behaviour.
• Optimise Spark workloads through appropriate partitioning, built-in Spark functions, efficient joins, adaptive execution and avoidance of unnecessary shuffles or UDFs.
• Select suitable compute models and cluster configurations based on workload characteristics.
• Apply cost-awareness to pipeline design, compute sizing, scheduling, storage and data-retention decisions.
• Monitor resource consumption and identify opportunities to reduce processing times and cloud costs without compromising reliability or data quality.
• Balance immediate delivery requirements with sustainable architecture and long-term maintainability.
Data quality, governance and security
• Implement automated data validation, schema checks, null checks, referential-integrity controls and business quality rules.
• Detect and manage schema drift and unexpected changes in source data.
• Use Delta Lake and Unity Catalog capabilities to support data lineage, access control, metadata and governance.
• Ensure secrets and credentials are handled securely using approved secret-management mechanisms and managed identities.
• Maintain technical documentation, metadata and operational information for assigned data products.
• Collaborate with data governance, architecture, security and platform teams to ensure alignment with enterprise standards.
Collaboration and delivery
• Work closely with traders, analysts, data scientists, software engineers, product owners and platform teams to translate business requirements into robust technical solutions.
• Communicate design decisions, risks, dependencies and technical trade-offs clearly to technical and non-technical stakeholders.
• Contribute reusable components, templates, documentation and engineering guidelines for the wider data community.
• Work effectively in a distributed, international and cross-functional environment.
Your profile
An experienced data engineer with at elast 5 years of experience, ideally in the energy sector and/or trading.
Be able to operate fundamental power-price forecasting models for short- to mid-term trading and large amounts of data.
Be able to work fully remote in a collaborative environment with interdisciplinary teams.
Good communication skills and professional behaviour.