Hiring Research Scientists to help define and advance the core technical direction of an early-stage AI company focused on multimodal representation learning.
This role is for a researcher who has trained models natively across multiple modalities, audio, images, video, text, and other structured inputs, and understands how richer representations can improve downstream reasoning and action. You’ll work directly with the founder, help set research direction from the earliest stage, and have significant autonomy over what gets explored and built.
The strongest candidate has worked directly on models that learn across several modalities rather than simply combining independently trained models at inference time. You should understand the challenges of cross-modal representation learning, model training, conditioning, data quality, evaluation, and scaling—and be interested in pushing those systems further.
Strong multimodal researchers from adjacent areas, particularly video generation and multimodal foundation models, are also highly relevant.
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.