SonicJobs Logo
Left arrow iconBack to search

Research Scientist, Multi-Modal Human Understanding

Meta
Posted 2 days ago, valid for a month
Location

Burlingame, CA, US

Salary

$122,000 - $181,000 per year

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Meta is seeking a Research Scientist to advance multi-modal AI technologies for human understanding and synthesis. In this role, you will develop Vision-Language Models (VLMs) and video foundation models that enable machines to perceive, interpret, and generate rich representations of human behavior, expression, and interaction. Your research will span multi-modal reasoning, video understanding, and generative synthesis, enabling more natural and intuitive human-computer interaction at scale.

Responsibilities

  • Design and implement novel multi-modal architectures that fuse vision, language, and temporal signals for holistic human understanding
  • Develop and train Vision-Language Models (VLMs) for tasks including visual question answering, image-text reasoning, and grounded human-centric understanding
  • Build video foundation models capable of temporal reasoning, action synthesis and long-form video synthesis with applications to human behavior synthesis
  • Research generative synthesis techniques for human-centric content including video generation, motion synthesis, and multi-modal content creation
  • Conduct rigorous experiments to evaluate model performance across diverse benchmarks, analyze failure modes, and iterate on architectures to improve accuracy and generalization
  • Contribute to the full research lifecycle from problem formulation and dataset curation through model development and evaluation


Minimum Qualifications

  • Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
  • 2+ years of experience in multi-modal AI research, including hands-on work with Vision-Language Models, video understanding, or human-centric AI systems
  • 2+ years of experience implementing and training large-scale neural networks using frameworks such as PyTorch, with experience on transformer-based architectures
  • Experience designing and executing experiments to evaluate multi-modal model performance, including quantitative analysis across vision, language, and video benchmarks
  • Experience writing production-quality or research-quality code in Python for multi-modal AI applications


Preferred Qualifications

  • Experience developing or fine-tuning Vision-Language Models for human understanding tasks
  • Experience with video foundation models, temporal transformers, or large-scale video pretraining
  • Track record of contributing to published multi-modal AI research at venues such as CVPR, ICCV, or NeurIPS
  • Experience with generative models for human synthesis including diffusion models, GANs, or autoregressive models for video or motion generation


$122,000/year to $181,000/year + bonus + equity + benefits



Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.