Fine-tune and adapt open-source speech models on our proprietary call audio
Build training and evaluation pipelines for multilingual and code-mixed speech
Own model quality against metrics that matter operationally — entity and numeric accuracy, latency to first response — not just aggregate error rates
Design and run data curation at scale: pseudo-labelling, speech enhancement, quality filtering on messy real-world audio
Work with our linguist on text normalisation and pronunciation handling
Evaluate candidate architectures, make the call with evidence, and ship the result to production with the platform team
Requirements
Must have skills:
3–4 years in ML, with at least 18 months on speech or audio specifically
Strong Python and PyTorch; comfortable reading a paper and implementing it
Hands-on experience fine-tuning at least one production speech model
Solid grasp of speech fundamentals — mel-spectrograms, acoustic models and vocoders, encoder-decoder vs transducer architectures, evaluation methodology, sampling rates and what they cost you
Understanding of how modern speech systems are actually built: self-supervised encoders, neural audio codecs, LM-based generation, flow matching
Experience with genuinely messy audio, not only clean benchmark datasets
Nice to have:
NeMo, ESPnet, SpeechBrain, or Coqui
Telephony-band or contact-centre audio
Multilingual or code-switched speech work
LoRA/PEFT, distributed training
Open-source contributions or publications in speech
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in India.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip