Full time Remote 9 days ago
Job Description We build the evaluation layer that validates AI agents before they reach customers — an automated system that scores across large volumes of agent traces. You'll own core parts of that platform: the pipeline that runs traces through model-based judges at scale, and the scoring logic that turns raw output into results teams can act on. What you'll do Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety — across LLM-as-a-J…