Design, develop, and maintain ETL (Extract, Transform, Load) processes to ensure the seamless integration of raw data from various sources into our data lakes or warehouses.
Utilize Python, PySpark, SQL and AirFlow etc., to process, analyze, and store large-scale datasets efficiently.
Write and maintain SQL queries for data retrieval, transformation, and storage in relational databases like Redshift or PostgreSQL.
Support cloud-based data platforms such as AWS, Azure, or GCP, with a focus on orchestrating AI retraining cycles, versioning, and automated pipeline monitoring.
Familiarity in converting unstructured data into vectors using frameworks like LangChain or LlamaIndex and storing them.
Collaborate with cross-functional teams, including data scientists, ML engineers, and domain experts to design and implement scalable solutions.
Troubleshoot and resolve performance issues, data quality problems, and errors in data pipelines.
Document processes, code, and best practices for future reference and team training.
Requirements
Additional Information:
Experience level 3+ years.
Strong understanding of data governance, security, and compliance principles is preferred.
Ability to work independently and as part of a team in a fast-paced environment.
Excellent problem-solving skills with the ability to identify inefficiencies and propose solutions.
Experience with version control systems (e.g., Git) and scripting languages for automation tasks.
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in India.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip