Mirage is an AI video company focused on making creation dramatically easier. Our team tackles some of the hardest creative and technical challenges in generative media.
We get to rethink how people collaborate with AI that can actually do creative work with them—defining entirely new interaction paradigms for AI-native video tools. We make that possible by building multimodal agents that turn creative direction into edits, rendering engines that composite multiple layers of video and graphics, and generative video models trained from the ground up that create new media when needed.
All of this comes together in Captions, our creative workspace for generating, editing, and designing video. The technology behind Captions is now available more broadly: Tesseract gives any AI agent native video-editing capabilities, while developers can integrate our models into their own products through APIs.
Explore our work
Captions — Our flagship creative workspace
Tesseract — Our professional video engine for AI agents
Research — The foundation models we build in-house
Updates — The latest from Mirage
TechCrunch, Forbes, Fast Company — Press
Our Investors
We’re very fortunate to have some the best investors and entrepreneurs backing us, including Index Ventures, Kleiner Perkins, Sequoia Capital, Andreessen Horowitz, General Catalyst, Uncommon Projects, Kevin Systrom, Mike Krieger, Lenny Rachitsky, Antoine Martin, Julie Zhuo, Ben Rubin, Jaren Glover, SVAngel, 20VC, Ludlow Ventures, Chapter One, and more.
Please note that all of our roles will require you to be in-person at our NYC HQ (located in Union Square)
About the Role
Mirage is seeking an ML Engineer to push the boundaries of large language models for multimodal creative tasks. You'll develop new approaches for building and extending agentic systems that understand and operate over complex, real-world data, particularly video.
This role focuses on advancing agent capabilities, improving reasoning and control, and enabling new forms of interaction between language models and time-based media.
Responsibilities
Design and build end-to-end agentic systems for creative tasks
Develop novel approaches for training and adapting the large language models that power these agents
Design new objectives, datasets, and fine-tuning strategies to improve agent behavior and reliability
Explore multimodal reasoning and structured generation for creative control
Run systematic experiments to evaluate and improve agent performance in real-world tasks
Design evaluation frameworks for agentic workflows in video analysis and editing
Analyze failure modes across the full agent loop (planning, tool use, execution) and iterate on improvements
What makes you a great fit
BS/MS/PhD in CS, ML, or related field
Strong track record building production ML systems or agentic pipelines
Deep understanding of transformers and modern LLM techniques
Experience with fine-tuning, alignment, or post-training methods, especially for adapting models to generate structured outputs or drive tool use
Comfort owning the full stack, from model-level experiments to deployed agent systems
Strong experimental rigor and good taste for what makes agents actually work in practice
Benefits:
Comprehensive medical, dental, and vision plans
401K with employer match
Commuter Benefits
Catered lunch multiple days per week
Dinner stipend every night if you're working late and want a bite!
Grubhub subscription
Health & Wellness Perks
Multiple team offsites per year with team events every month
Generous PTO policy
Captions provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
Please note benefits apply to full time employees only.