Broad Function:
The Senior Data Engineer – VARTA SENSE will be responsible for designing, developing, and managing scalable and reliable data pipelines that form the core data foundation of the VARTA SENSE platform.
The role will work closely with the ML Lead, Solution Architect, Backend Engineering team and other technology stakeholders to transform complex banking and transaction data into standardized, reusable and production-ready datasets and features.
The position requires strong hands-on expertise in Python, SQL, Spark/PySpark, data pipelines, orchestration, data quality and distributed data processing, with particular emphasis on reliability, performance, scalability and secure deployment within restricted banking environments.
Roles and Responsibilities (not limited to):
1. Data Pipeline Development & Engineering
- Design and develop scalable batch, incremental and production-grade data pipelines for large volumes of transaction and customer data.
- Build standardized and reusable data models and canonical schemas for transactions, customer attributes, offers, exposure and outcome events.
- Develop configurable ingestion and mapping frameworks to integrate data received from different banks and source systems.
- Manage pipeline orchestration including dependencies, checkpointing, retries, partial failures, backfills and safe replay mechanisms.
- Implement deduplication, late-arriving data handling, incremental loads, CDC, watermarking and merge/upsert processing.
2. Data Quality, Reconciliation & Governance
- Establish strong data-quality controls, automated testing, source-to-target reconciliation and schema validation.
- Ensure data lineage and schema evolution are properly managed across pipelines.
- Proactively identify and prevent issues such as missing data, duplicate records, incorrect aggregations, silent data loss and double counting.
- Maintain reliable and auditable data-processing standards across production environments.
3. ML Feature Engineering Support
- Work closely with the ML Lead to develop and maintain analytical and ML feature pipelines.
- Ensure consistency between training and production/inference datasets.
- Maintain point-in-time correctness and prevent data leakage within feature pipelines.
- Support versioned analytical features and reusable feature definitions.
4. Performance & Scalability
- Optimize large-scale data-processing workloads through effective partitioning, query optimization, join strategies, memory management and I/O optimization.
- Benchmark workloads against expected transaction volumes and available infrastructure.
- Identify bottlenecks and continuously improve processing speed, stability and infrastructure utilization.
- Ensure data pipelines remain scalable as transaction volumes and customer deployments increase.
5. Deployment, Security & Production Reliability
- Develop and support pipelines for deployment within bank-controlled on-premise and private-cloud environments.
- Ensure pipelines can operate in restricted or offline environments without public-internet dependency at runtime.
- Implement appropriate controls for PII handling, encryption, tokenization, masking, access management and secure connectivity.
- Establish monitoring and alerting for pipeline failures, data-quality issues, infrastructure bottlenecks and processing delays.
- Define and maintain backfill, replay, recovery and disaster-recovery procedures for critical pipelines.
6. Cross-functional Collaboration
- Work closely with ML, Architecture, Backend Engineering, DevOps and Infrastructure teams to design and deliver reliable data solutions.
- Collaborate with development partners to reproduce, transition and operationalize data pipelines within FCI.
- Support integration between data platforms and downstream application / ML services.
- Participate in technical discussions, architecture reviews and production troubleshooting.
7. Reusability, Documentation & Knowledge Transfer
- Build reusable adapters and configuration layers so that onboarding a new bank can be managed through configuration and mapping rather than changes to the core product code.
- Maintain clear technical documentation, runbooks and operational procedures.
- Document architecture decisions, troubleshooting steps and recovery procedures.
- Enable alternate engineering resources to independently operate and troubleshoot critical pipelines.
- Participate in code reviews, Git-based development, CI/CD practices and continuous improvement of engineering standards.