Researcher - Reinforcement Learning

Company : Huawei Technologies Canada Co.,

Job Type : Full Time

Edmonton, Canada

Job Description - Researcher - Reinforcement Learning

Job description

Huawei Canada has an immediate 12-month contract opening for a Reinforcement Learning Researcher.

About the team:

Founded in 2012, the Noah’s Ark lab has evolved into a prominent research organization with notable achievements in academia and industry. The lab’s mission focuses on advancing artificial intelligence and related fields to benefit the company and society. Driven by impactful, long-term projects, the aim is to enhance state-of-the-art research while integrating innovations into the company's products and services, including LLMs, RL, NLP, computer vision, AI theory, and Autonomous driving.

About the job:

Enabling Large Language Models (LLMs) to learn from experience, interaction, and environment feedback, moving beyond static fine-tuning toward continual, agentic self-improvement.
LLM post-training paradigms (e.g., RLHF, GRPO, reward-free methods, etc.).
Agentic reinforcement learning for tool-using and browsing-based LLMs trained in interactive environments.
Agentic evaluation and benchmarking, including design of multi-turn, verifiable reasoning tasks.
Your work will involve implementing and evaluating new training and evaluation pipelines for reasoning-enhanced LLMs and tool-using agents, scaling experiments on large GPU clusters, and contributing to scientific insights and publications in this emerging area.

Job requirements

About the ideal candidate:

PhD degree in Computer Science or related fields or master's degree with comparable experience.
Strong foundation in deep learning, including architectures such as Transformers and optimization techniques for large models.
Practical or research experience in reinforcement learning, self-supervised learning, or language model fine-tuning.
Proven research record in AI by having at least one paper as the first author in top tier venues, such as NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ICRA.
Solid proficiency in Python and experience with PyTorch, DeepSpeed, Megatron and other distributed training frameworks.
Familiarity with LLM post-training pipelines (RLHF, GRPO/PPO, SFT, LoRA, MoE, etc.) is an asset.
Experience with multi-agent RL, tool-use / browser/coding agents, is an asset.
Strong communication and writing skills; enthusiasm for open research and collaborative problem-solving.

Huawei aims to support a French-speaking work environment for its employees in Quebec. We have taken steps to avoid requiring a language other than French for this position. However, proficiency in English is essential for this role for the following reasons:

The person will be required to communicate regularly with colleagues located outside Quebec, where English is the primary language used for communication between offices. In addition, the nature of the tasks related to this position, which falls within a highly specialized field of artificial intelligence, also requires knowledge of English.

Original job Researcher - Reinforcement Learning posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.

Share Job

Get your Resume Reviewed for Free

About the Company

Huawei Technologies Canada Co.,

Huawei is a global leader of ICT solutions. Continuously innovating based on customer needs, we are committed to enhancing customer experiences and creating maximum value for telecom carriers, enterprises, and consumers. Our telecom network equipment, IT products and solutions, and smart devices are...

Similar Researcher - Reinforcement Learning Jobs in Canada

Get your Resume Reviewed for Free

Email address

Why are you reporting this job?

I think it’s a discriminatory or offensive

I think it’s fraudulent or a scam

I think it’s trying to sell something unrelated to the job / it’s asking for money

I think it contains incorrect or broken information

Other

All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.

Setup your job alert:

Frequency

By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime. Skip

Researcher - Reinforcement Learning

Job Description - Researcher - Reinforcement Learning

Job description

Job requirements

All done!

You've already applied for this job

About the Company

Similar Researcher - Reinforcement Learning Jobs in Canada

Mobile Apps