Incedo is a US-based consulting, data science and technology services firm with over 3000 people helping clients
from our six offices across US, Mexico and India. We help our clients achieve competitive advantage through
end-to-end digital transformation. Our uniqueness lies in bringing together strong engineering, data science, and
design capabilities coupled with deep domain understanding. We combine services and products to maximize
business impact for our clients in telecom, Banking, Wealth Management, product engineering and life science
& healthcare industries.
Working at Incedo will provide you an opportunity to work with industry leading client organizations, deep
technology and domain experts, and global teams. Incedo University, our learning platform, provides ample
learning opportunities starting with a structured onboarding program and carrying throughout various stages of
your career. A variety of fun activities is also an integral part of our friendly work environment. Our flexible
career paths allow you to grow into a program manager, a technical architect or a domain expert based on your
skills and interests.
Our Mission is to enable our clients to maximize business impact from technology by
We are looking for a skilled Data Engineer with expertise in columnar/OLAP databases, Python-based data
pipelines, and cloud/on-premises infrastructure. The ideal candidate will be comfortable owning the data lifecycle
— from ingestion and transformation to scalable storage and query optimization — in high-throughput, latencysensitive
environments
Key Responsibilities
Design, build, & maintain scalable pipelines for ingest, transform & serve structured and semi-structured data.
Work with columnar/OLAP (Online Analytics Process) & near real-time databases for analytical workloads —
schema design, partitioning, and query optimization.
Configure horizontal and vertical scaling strategies to sustain database performance under growing data
volumes and query loads.
Develop and maintain Python/PySpark ETL workflows integrated with data warehouse architectures.
Deploy and manage data infrastructure on cloud platforms (preferably AWS) and on-premises environments
along with Open-sources Technologies.
Containerize data services using Docker and orchestrate workloads via Kubernetes.
Perform analysis on data quality issues, pipeline failures, and performance bottlenecks.
Document data models, pipeline flows, and operational runbooks.
Collaborate with team to translate requirements into reliable data products.
Required Technical Skills
Columnar & OLAP Databases
Hands-on experience with one or more columnar/OLAP (Online Analytics Process) databases: ClickHouse,
Apache Druid, Amazon Redshift, Snowflake, Apache Cassandra, or MariaDB Column Store.
Experienced with columnar storage internals — compression, materialized views, and partition pruning.
Proficiency in designing schemas optimized for analytical query patterns — star/snowflake schemas, wide
tables, and time-series partitioning.
Data Engineering & Programming
Strong Python programming — shell scripts, data manipulation, API integrations, and automation.
Experience with PySpark for distributed data processing at scale.
Working knowledge of SQL — complex joins, window functions, CTEs, and query profiling.
Familiarity with data warehousing concepts - slowly changing dimensions and ELT patterns.
Scaling & Performance
Demonstrated experience with horizontal scaling: sharding, replication, and distributed query execution.
Proficiency in resource tuning, index optimization, and memory management.
Ability to benchmark and profile database query performance and identify bottlenecks.
Infrastructure & Cloud
Experience with AWS data services: S3, Glue, EMR, RDS, Redshift, or equivalents.
Comfortable working in on-premises Linux environments for database administration and operations.
Proficiency with Docker for containerizing data services, experience with Kubernetes for orchestrating
workloads.
.
We are looking for a skilled Data Engineer with expertise in columnar/OLAP databases, Python-based data
pipelines, and cloud/on-premises infrastructure. The ideal candidate will be comfortable owning the data lifecycle
— from ingestion and transformation to scalable storage and query optimization — in high-throughput, latencysensitive
environments
Key Responsibilities
Design, build, & maintain scalable pipelines for ingest, transform & serve structured and semi-structured data.
Work with columnar/OLAP (Online Analytics Process) & near real-time databases for analytical workloads —
schema design, partitioning, and query optimization.
Configure horizontal and vertical scaling strategies to sustain database performance under growing data
volumes and query loads.
Develop and maintain Python/PySpark ETL workflows integrated with data warehouse architectures.
Deploy and manage data infrastructure on cloud platforms (preferably AWS) and on-premises environments
along with Open-sources Technologies.
Containerize data services using Docker and orchestrate workloads via Kubernetes.
Perform analysis on data quality issues, pipeline failures, and performance bottlenecks.
Document data models, pipeline flows, and operational runbooks.
Collaborate with team to translate requirements into reliable data products.
Required Technical Skills
Columnar & OLAP Databases
Hands-on experience with one or more columnar/OLAP (Online Analytics Process) databases: ClickHouse,
Apache Druid, Amazon Redshift, Snowflake, Apache Cassandra, or MariaDB Column Store.
Experienced with columnar storage internals — compression, materialized views, and partition pruning.
Proficiency in designing schemas optimized for analytical query patterns — star/snowflake schemas, wide
tables, and time-series partitioning.
Data Engineering & Programming
Strong Python programming — shell scripts, data manipulation, API integrations, and automation.
Experience with PySpark for distributed data processing at scale.
Working knowledge of SQL — complex joins, window functions, CTEs, and query profiling.
Familiarity with data warehousing concepts - slowly changing dimensions and ELT patterns.
Scaling & Performance
Demonstrated experience with horizontal scaling: sharding, replication, and distributed query execution.
Proficiency in resource tuning, index optimization, and memory management.
Ability to benchmark and profile database query performance and identify bottlenecks.
Infrastructure & Cloud
Experience with AWS data services: S3, Glue, EMR, RDS, Redshift, or equivalents.
Comfortable working in on-premises Linux environments for database administration and operations.
Proficiency with Docker for containerizing data services, experience with Kubernetes for orchestrating
workloads.
.
Data & DevOps Tooling
Experience with CI/CD pipelines — Jenkins, GitHub, or similar — for deploying data pipeline code. Familiarity with Apache Kafka for real-time data streaming and event-driven architectures.
Experience with Elasticsearch for log ingestion, full-text search, or observability use cases.
Networking (Domain-Specific)
Understanding of OSI stack (L2–L3) and network protocols: TCP/IP, UDP, Modbus, EtherNet/IP, Profinet.
Exposure to telemetry or time-series data from industrial devices or SCADA systems is a plus.
Proficiency in Atlassian tools: JIRA, Confluence and Bitbucket for task management & documentation
Experience working in Scrum teams with evolving requirements and sprint-based delivery.
Strong problem-solving mindset with the ability to independently scope, prioritize, and deliver tasks
Exposure to C/C++ for performance-critical components or interfacing with native libraries
Qualifications
We value diversity at Incedo. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.