Full-Stack Engineer
About
About Judgment Labs
Judgment Labs builds infrastructure for Agent Behavior Monitoring (ABM). While traditional observability focuses on logging exceptions and latency, ABM surfaces behavioral anomalies such as instruction drifts and context retrieval loss in scaled production environments. Hundreds of teams building autonomous agents rely on Judgment to understand how their systems are behaving post-deployment. The team has raised $30M+ across two rounds in the past five months. Investors include Lightspeed, SV Angel, Valor Equity Partners, Nova Global, Chris Manning, Michael Ovitz, Michael Abbott, Cory Levy, and Kevin Hartz. The team is under 20 people and ships at 50+ company velocity, with Olympiad medalists, debate champions, and competitive athletes who bring that same intensity to company building. Everyone is either an ex-founder or a founder-to-be.
About the role
Judgment Labs is hiring a Full-Stack Engineer to own the product experiences and agent infrastructure that make the agent improvement loop legible and actionable for engineering teams. This is not a role where you implement specs handed down: you will talk to customers, define what to build, build it, and iterate until it is great. The role spans from the data layer to the UI, including the agent swarm interfaces, verification platforms, and the SDK layer that lets developers summon Judgment mid-development. About 30% of the role is customer-facing.
What you'll own
- Shape how the Judgment Agent runs large-scale parallel investigations across thousands of production traces, merging failure modes, tool errors, regressions, and drift signals into a single actionable answer
- Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals, trajectory replay against changed agents, and monitors for unintended behavior
- Design how engineers understand long traces, tool calls, decisions, and failures: making a thousand-step reasoning trajectory legible in minutes
- Build the swarm UX so engineers can watch parallel investigations, redirect investigators that go down the wrong path, and consume findings without reading hundreds of reports
- Own the improvement loop: the workflows that turn production trajectories into datasets, judges, and regression checks so the path from found problem to verified fix feels like one motion
- Build and maintain the SDK and terminal-first experience so agent development sessions can summon Judgment as a subagent mid-development
- Own platform infrastructure: workspaces, roles, permissions, billing, usage, and limits for teams running many agents across many environments
Requirements
Must-have
- Experience building and scaling end-to-end production systems, from data layer to UI
- Hands-on experience building with LLMs or agents, or demonstrable drive to get there fast
- Comfort working directly with customers to understand their needs and solve real-world problems (approximately 30% of the role is customer-facing)
- Strong technical problem-solving ability in fast-changing, ambiguous environments
- Excellent communication skills: clear, direct, and persuasive across technical and non-technical audiences
- Based in San Francisco or willing to relocate; 5 days per week in-person
Nice-to-have
- Prior evals, observability, or agent behavior-monitoring product experience
- Strong engineering pedigree from a solid product company (e.g. Roblox, Snapchat, Tesla); top-tier infra background (e.g. Databricks) is a bonus
- Prior forward-deployed or solutions engineering experience
- Prior 0-to-1 startup experience as an early hire
- Founder background or strong founder-to-be signal
- Comfort across the full stack with backend and infrastructure weight
- Junior level (3 years experience): $200K
- Mid to senior level (3-7 years experience): up to $300K
- Exception: exceptional AI-savvy FDE leads can reach $400K on a case-by-case basis
Benefits & perks
- Full benefits package
- Equinox membership
- Private chef
- Competitive compensation
- Direct customer interface and influence on product roadmap
Interview process
- 1Application Review
- 2recruiter screen
- 3evals fde round
- 4booked evals round
- 5coding iq round
- 6booked iq round
- 7fde onsite
- 8Offer
Drop your CV for this role.
One PDF and your email. We read it, score your fit for this role at Judgment Labs, and route the introduction through us.