Evals
The golden suite: tickets with known right answers, run through the real agents. Each case checks behaviour (what the agents did) and, with a real model, is graded by a judge. CI runs it on every push and blocks the merge if a safety case fails or the pass rate drops.
No evaluation runs yet. Run uv run agentdesk eval in core/.