Job Overview
Job description
- Location:
- Redwood City, CA, United States
- Work arrangement:
- On-site
Role Summary
You'll build and train large-scale multimodal agentic models — systems that reason, plan, code, and call tools to do complex, multi-step work over pixels. This is core research shaping how users interact with what Luma's models can do.
It's a multi-stack research role across modeling, data, systems, and evaluation, on novel problems with no existing playbook, treating science and engineering as equally important. It fits someone grounded in foundation models and agentic systems who's trained models at real scale. If you want to work in only one layer of the stack, this deliberately spans several.
Responsibilities
- Architect large-scale multimodal agentic models that use reasoning, planning, coding, and tool calling for complex, multi-step work.
- Design, build, and run robust data pipelines to construct, enrich, and filter massive pixel datasets, and formulate new tasks.
- Train large-scale multimodal models on massive datasets and GPU clusters.
- Define and build novel evaluation frameworks to measure multimodal agents.
First 90 Days
One way the first 90 could unfold.
- Days 1–30 — Immerse & Diagnose: Learn the current models, agentic approaches, and where evaluation and data are weakest.
- Days 30–60 — Ship & Validate: Improve an agentic capability (reasoning, tool use, or coding) and prove it with a new eval.
- Days 60–90 — Scale & Systemize: Scale the approach across datasets and clusters and harden the evaluation framework.
Requirements
- Strong foundation in machine learning, foundation models, and agentic systems.
- Deep understanding of agentic systems and LLM/VLM reasoning, coding models, and tool calling.
- Hands-on PyTorch and large-scale training (distributed, mixed precision, large datasets).
Nice to Have
- Experience with state-of-the-art foundation models in reasoning, coding, or tool calling, or state-of-the-art multimodal agents.
About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.
About the Company
ABOUT LUMA AI
Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence, and the next step beyond language models comes from vision.
Design at Luma paints the vision of what can be, and we are looking for a Visual Designer to help raise the bar for how our product looks and feels. This is a craft-first role for someone who cares deeply about the details that make a product feel exceptional.
- Role:
- Research Scientist / Engineer — Foundation Model (Agent)
- Job Type:
- FullTime
Company profile
lumaai
lumalabs.aiLuma AI (lumaai) builds a multimodal generative AI platform, known as Luma Agents, that creates realistic videos from text, image, video, audio and agentic inputs. The company describes its work as building the future of creative intelligence, with the aim of enabling new forms of human expression through AI. It runs a lean team in Redwood City and offers staff fully covered health plans, flexible time off and a $1,500 yearly learning stipend.