P

Staff Machine Learning Engineer, Computer Vision

Job Description - Staff Machine Learning Engineer, Computer Vision

About Pinterest:


Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product.


Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible.


At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI.


Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here.


Within Pinterest, the Pinterest Labs organization focuses on applied ML research and development. Labs works across a broad variety of AI/ML initiatives—including core computer vision, multimodal representation learning, heterogeneous graph neural networks, generative modeling, and recommender systems. This is the group that develops the foundation ML models that fully leverage the tens of billions of Pins and the associated knowledge graph to improve the core product.


We are currently hiring for the Visual Modeling team in Labs, which develops Pinterest's in-house visual encoder. In this role, you'll work with Pinterest's rich visual-text dataset to train large-scale models from scratch that are continuously shipped to production to power visualization features. You'll build multimodal representations that power applications such as recommender systems, Semantic IDs, and a range of downstream ML models. The visual encoder also produces visual tokens that power our in-house VLM and composed image retrieval models. The core visual pod is a small group (~10 engineers) inside Labs, which allows for deep collaboration. For example, engineers working on multimodal representation also contribute to our internal text-to-image generation Canvas project—collaborating on autoencoder design or on reward function development for RL training.


 


What you’ll do:



  • Prototype state-of-the-art visual encoders that power Pinterest's recommender systems and internal visual language models.

  • Experiment with billion-scale datasets and gain hands-on experience with large-scale GPU computing.

  • Build flexible visual reasoning tools such as composed image retrieval, promptable detection/segmentation, and instruction-tuned embedding and generative models.

  • Read research papers, participate in group discussions, and help brainstorm the company's overall visual generative strategy.

  • Help collect relevant visual instruction training data that can be shared across multimodal representation, composed image retrieval, text-to-image generation and visual language modeling.

  • Publish and share your work through conferences, paper submissions, and blog posts.

  • Mentor junior researchers and research interns within the Pinterest Labs organization.

  •  


What we’re looking for:



  • Research engineers and scientists with experience building and training computer vision models.

  • Experience with multimodal representations and visual language modeling is strongly preferred.

  • A track record of research contributions (e.g., publications, open-source work) and/or shipping ML models to production.

  • Hands-on experience with large-scale model training and modern deep learning frameworks (e.g., PyTorch).

  • Strong collaboration skills and a demonstrated ability to work effectively in a small, fast-moving team.

  • M.S. or PhD in Machine Learning or related academic areas, or equivalent work experience.

  • Publications at top ML conferences

  • Experience using Cursor, Copilot, Codex, or similar AI coding assistants for development, debugging, testing, and refactoring


 


Relocation Statement:



  • This position is not eligible for relocation assistance. Visit our PinFlex page to learn more about our working model.


 


In-Office Requirement Statement:



  • We let the type of work you do guide the collaboration style. That means we're not always working in an office, but we continue to gather for key moments of collaboration and connection.

  • This role will need to be in the office for in-person collaboration 1-2 times/quarter and therefore can be situated anywhere in the country.


 


#LI-REMOTE
#LI-AK7

Original job Staff Machine Learning Engineer, Computer Vision posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Staff Machine Learning Engineer Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.