Job Overview
Job description
- Location:
- San Francisco, United States
- Work arrangement:
- On-site
Role Summary
Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works:
- We ingest everything your agents do in production: traces, tool calls, decisions, outcomes
- Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals
- Teams close the loop, shipping agent improvements validated against real production evidence
You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great. This is not a role where you implement specs handed down.
Responsibilities
WHAT YOU WILL ACCOMPLISH
- Investigation interfaces: Design how engineers understand what their systems did and why. Long traces, tool calls, decisions, failures. How do you make a complex sequence of events legible in minutes?
- Verification: Build the platform for verifying system changes: hosted simulated environments, trajectory replay, and monitors for unintended behavior changes.
- The improvement loop: Build the workflows that turn production data into datasets, evaluations, and regression checks, so the path from "found a problem" to "verified a fix" feels like one motion.
- The platform underneath: Workspaces, roles, permissions, billing, usage, and limits for teams running many workflows across many environments.
Requirements
- Experience building and scaling end-to-end production systems, from data layer to UI
- Strong technical problem-solving skills, especially in fast-changing, ambiguous environments
- A builder and tinkerer's mindset with high agency - you find creative ways to overcome obstacles and ship
- Comfort working directly with customers to understand their needs and solve real-world problems
- Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences
- Role:
- Product engineer, full stack
- Job Type:
- FullTime
Company profile
Judgment Labs
judgmentlabs.aiJudgment Labs builds a continuous-improvement stack for AI agents, helping teams monitor and improve agent behavior at scale. Its learning infrastructure ingests long agent trajectories from production, structures and evaluates them, surfaces what matters and feeds that learning back into the agent. The company turns production experience into data for evals, labeling, rubric generation, context engineering and RL workflows, and helps teams decide which reported failures are worth solving.