Logo-of-Aligned-Automation-Services-hiring-for-jobs-in-India-on-GrabJobs

Data Engineer Pyspark and Mongo DB

Job Description - Data Engineer Pyspark and Mongo DB

Job Title: Data Engineer (MongoDB, PySpark & Python)

Experience

5–8 Years

Location

As per business requirement

Job Summary

We are looking for an experienced Data Engineer with strong expertise in MongoDB, PySpark, and Python to design, develop, and optimize scalable data pipelines. The ideal candidate should have experience working with large datasets, NoSQL databases, distributed data processing, and cloud-based data platforms.

Key Responsibilities

  • Design, build, and maintain scalable ETL/ELT data pipelines using PySpark and Python.
  • Develop data ingestion frameworks to process structured, semi-structured, and unstructured data.
  • Work extensively with MongoDB for data modeling, querying, indexing, aggregation, and performance optimization.
  • Optimize Spark jobs for high-performance processing of large datasets.
  • Build reusable data transformation and validation frameworks.
  • Develop REST API integrations and automate data ingestion using Python.
  • Monitor, troubleshoot, and optimize data pipelines for reliability and performance.
  • Collaborate with business analysts, data scientists, and application teams to deliver data solutions.
  • Implement data quality, governance, and security best practices.
  • Participate in code reviews and follow CI/CD and Agile development practices.

Required Skills

  • Strong experience in Python programming.
  • Hands-on experience with PySpark and Spark SQL.
  • Strong knowledge of MongoDB, including:
    • CRUD Operations
    • Aggregation Framework
    • Indexing
    • Replication
    • Sharding
    • Performance Tuning
  • Good understanding of data structures and algorithms.
  • Experience in developing ETL/ELT pipelines.
  • Strong SQL skills.
  • Experience with Git version control.
  • Knowledge of Linux/Unix commands.
  • Experience working with JSON, XML, and Parquet data formats.

Preferred Skills

  • Experience with Databricks.
  • Experience with cloud platforms such as Azure, AWS, or GCP.
  • Knowledge of Apache Kafka or other streaming technologies.
  • Experience with orchestration tools such as Apache Airflow.
  • Understanding of Delta Lake and Lakehouse architecture.
  • Familiarity with CI/CD pipelines.


Original job Data Engineer Pyspark and Mongo DB posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Data Engineer Jobs in India

GrabJobs is the no1 job portal in India, connecting you to thousands of jobs fast! Find the best jobs in India, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.