We are looking for an experienced AI/ML MLOps Engineer with strong hands-on expertise in LLM fine-tuning, model deployment, AWS GPU infrastructure, and MLOps. The role involves fine-tuning and deploying self-hosted Large Language Models (LLMs), building training and evaluation pipelines, and implementing reliable production deployment and monitoring practices.The ideal candidate should have practical experience working across the complete ML lifecycle — data preparation, model fine-tuning, evaluation, deployment, monitoring, and continuous improvement.
Key Responsibilities
Fine-tune Large Language Models using Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO).
Develop and maintain training data pipelines, including data transformation, formatting, deduplication, filtering, and quality validation.
Work extensively with the Hugging Face ecosystem, including Transformers, Datasets, and PEFT.
Build and automate model evaluation and benchmarking frameworks to assess model quality and performance.
Deploy and serve LLM models using AWS GPU/EC2 infrastructure and Amazon SageMaker.
Optimize models for production through model quantization, inference optimization, and resource utilization.
Build robust MLOps and ML CI/CD pipelines covering model training, evaluation, packaging, deployment, and monitoring.
Implement A/B testing, Canary, and Shadow-mode deployments for safely introducing new model versions into production.
Develop mechanisms for automated model promotion and rollback based on predefined performance and operational metrics.
Implement production monitoring for model performance, latency, throughput, errors, GPU utilization, and resource consumption.
Containerize ML workloads using Docker and deploy/manage them using Kubernetes/Amazon EKS.
Collaborate with Data Scientists, ML Engineers, DevOps teams, and other stakeholders to build scalable and reliable AI/ML solutions.
Requirements
Strong programming experience in Python.
Hands-on experience with LLM fine-tuning, particularly SFT and DPO.
Strong knowledge of Hugging Face Transformers, Datasets, and PEFT.
Experience working with AWS GPU/EC2 and SageMaker for ML workloads.
Strong understanding of MLOps, ML CI/CD, and model lifecycle management.
Experience with LLM model serving and production deployment.
Experience building training data preparation and processing pipelines.
Knowledge of model evaluation, benchmarking, and performance optimization.
Hands-on experience with model quantization.
Experience implementing A/B, Canary, and Shadow-mode deployments
Benefits
Comprehensive Medical Coverage:
Health insurance of INR 7.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
Robust Protection Plans:
Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
Retirement Benefits:
PF and Gratuity provided as per standard government regulations.
Flexible Work Options:
Enjoy hybrid work arrangements & flexible working hours
Generous Leave Policy:
21 days of annual leave, in addition to 10 company-declared holidays.
Employee Well-being Spaces:
Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in India.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip