Palo Alto
3 months ago

Job Overview

Job Type
FullTime
Pay
Not disclosed

Job description

Location:
Palo Alto, United States
Work arrangement:
On-site

Role Summary

Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.

Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.

You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.

YOU MAY BE A FIT IF

  • You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs.
  • You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.
  • You've worked with Kubernetes, GPU scheduling, or inference infrastructure.
  • You think in terms of reliability, SLOs, and honest capacity planning.
  • Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.

Responsibilities

  • Build and scale the inference platform that serves every request from ollama.com http://ollama.com.
  • Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.
  • Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.
  • Build the reliability, observability, and cost controls for our team and customers
Role:
Software Engineer, Cloud
Job Type:
FullTime

Company profile

Ollama is an open-source software platform for running and managing large language models on local GPU infrastructure and through hosted cloud models. It provides a command-line interface, a native GUI, a local REST API, model-management tools and integrations for using open-weight models with coding assistants. Ollama was developed by Jeffrey Morgan and Michael Chiang and is trusted by more than 9M developers.

Founded
2023

More jobs at ollama

Similar Software Engineer jobs at other companies