Job Overview
Job description
- Location:
- Palo Alto, United States
- Work arrangement:
- On-site
Role Summary
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.
Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.
You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.
YOU MAY BE A FIT IF
- You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs.
- You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.
- You've worked with Kubernetes, GPU scheduling, or inference infrastructure.
- You think in terms of reliability, SLOs, and honest capacity planning.
- Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.
Responsibilities
- Build and scale the inference platform that serves every request from ollama.com http://ollama.com.
- Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.
- Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.
- Build the reliability, observability, and cost controls for our team and customers
- Role:
- Software Engineer, Cloud
- Job Type:
- FullTime
Company profile
ollama
ollama.comOllama is an open-source software platform for running and managing large language models on local GPU infrastructure and through hosted cloud models. It provides a command-line interface, a native GUI, a local REST API, model-management tools and integrations for using open-weight models with coding assistants. Ollama was developed by Jeffrey Morgan and Michael Chiang and is trusted by more than 9M developers.
- Founded
- 2023