Responsibilities
- Design and implement novel multi-modal architectures that fuse vision, language, and temporal signals for holistic human understanding
- Develop and train Vision-Language Models (VLMs) for tasks including visual question answering, image-text reasoning, and grounded human-centric understanding
- Build video foundation models capable of temporal reasoning, action synthesis and long-form video synthesis with applications to human behavior synthesis
- Research generative synthesis techniques for human-centric content including video generation, motion synthesis, and multi-modal content creation
- Conduct rigorous experiments to evaluate model performance across diverse benchmarks, analyze failure modes, and iterate on architectures to improve accuracy and generalization
- Contribute to the full research lifecycle from problem formulation and dataset curation through model development and evaluation
Minimum Qualifications
- Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
- 2+ years of experience in multi-modal AI research, including hands-on work with Vision-Language Models, video understanding, or human-centric AI systems
- 2+ years of experience implementing and training large-scale neural networks using frameworks such as PyTorch, with experience on transformer-based architectures
- Experience designing and executing experiments to evaluate multi-modal model performance, including quantitative analysis across vision, language, and video benchmarks
- Experience writing production-quality or research-quality code in Python for multi-modal AI applications
Preferred Qualifications
- Experience developing or fine-tuning Vision-Language Models for human understanding tasks
- Experience with video foundation models, temporal transformers, or large-scale video pretraining
- Track record of contributing to published multi-modal AI research at venues such as CVPR, ICCV, or NeurIPS
- Experience with generative models for human synthesis including diffusion models, GANs, or autoregressive models for video or motion generation
$122,000/year to $181,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
