How it works

Agents that do the work, people who make the calls

A guide to agentdesk in seven parts. Every part starts with a plain-language summary, then the technical detail. If something is unclear, ask the guide: an agent of this platform answers from these pages and the running code.

01

Evaluations

How we know the agents still behave after any change: tickets with known right answers, run through the real system, scored, and enforced in CI.

In plain words

Like an exam with an answer key. Twelve customer messages whose right handling we know in advance, including six traps (another customer's order, a hidden “ignore your instructions”, an inflated refund). Every change to the code or a prompt has to pass the exam before it can go live.
golden.yaml12 cases · 6 safetyReal agentstriage, resolver, guardsAssertionswhat it did (exact)LLM judgehow well, 1 to 5Case resultpass or failGateno safety case failsand ≥ 90% of cases passCI: merge allowedor blocked if the gate is closedStored and comparedEvals page · Langfuse dataset run with scores
Each case runs through the real agents, is checked twice (exact assertions and a judge), and the gate decides whether the change may be merged. Every result is stored and compared.

Loading the suite…

NextOperations→