San Francisco
2 months agoJob Overview
Job Type
FullTimePay
Not disclosedJob description
- Location:
- San Francisco, United States
- Work arrangement:
- On-site
Founding Engineer at compresr in San Francisco, United States.
Role Summary
- Type: Full-time permanent contract or an Internship with a potential follow-up offer
- Location: San Francisco or remote with future relocation to San Francisco (sponsored)
- Start: ASAP
Requirements
ABOUT YOU
- A cracked full-stack engineer who enjoys a high-paced startup environment, takes pride in what they build and owns it end to end.
- Solid understanding of cloud infrastructure, deployment, and production systems on AWS.
- Python/basic ML Ops skills. Experience in scaling AI infra products is a plus.
- Proactive, strong communicator with fast response time, team player
TECH REQUIREMENTS
- Strong backend engineering fundamentals
- Experience with concurrency and distributed systems
- Experience deploying and scaling production backend services on AWS
- Ability to work across systems (Python + light frontend)
- Excellent Claude Code (or similar) user
NICE TO HAVE
- Open-source contributions
- Startup experience
- OAuth / API auth flows
STACK
- Backend: Python, FastAPI, PostgreSQL (Supabase), Redis, AWS
- Frontend: Next.js, React, TypeScript
- Tools: GitHub, Docker, Sentry, GitHub Actions
About the Company
- We're building state-of-the-art context compression. Our mission is to become the "Cloudflare for LLMs" — a compression layer embedded into most LLM pipelines by default.
- We're a team of ex-EPFL MSc/PhDs from dlab. We started by publishing papers, then got into YC and started making money helping companies cut their LLM costs.
- We run the business like a research lab: form hypotheses, kill the ones that don't work, double down on the ones that do.
Additional Information
- Intro call (20 min)
- Practical technical interview (60 min)
- Cultural interview (30 min)
- Role:
- Founding Engineer
- Job Type:
- FullTime
Company profile
compresr
compresr.aiCompresr is an LLM context-compression API: users send long context plus a query and get back a shorter context that keeps the answer-bearing tokens and drops the rest, cutting token cost and latency. It is available through a Python SDK, TypeScript SDK or hosted HTTP API. Compresr Inc. is run by a team of ex-EPFL MSc/PhDs who went through YC, and aims to become the Cloudflare for LLMs.
- Founded
- 2026