01
Channels in n8n
n8n owns the edges: how customers reach the platform, the traffic that keeps the demo alive, the morning report and the alarms.
In plain words
All six live in the agentdesk folder of the self-hosted n8n. Open any of them to read the sticky notes on what it does and how it fails.
Open the folder in n8n ↗





They are numbered in the order a reader should open them. Their source is n8n/build.py, which writes the JSON that is imported, so the workflows are reviewed like code. They need two credentials created in n8n (the API bearer token and the Telegram bot) and are switched on once the API has its public URL.
02
Tracing in Langfuse
Langfuse records what the agents actually did, call by call, so any reply can be explained after the fact.
In plain words
From the dashboard to a trace
Open any run on the Runs page and click Open full trace in Langfuse. Each ticket is one trace; the trace holds a generation for every model call and a span for every tool call, guard check, retry and fallback, in order. Traces are grouped by ticket as a session, tagged by channel, and labelled with the environment (agentdesk-dev, ci) so local, CI and production never mix.

eval.passed score. This one is the prompt-injection case: the injected “refund 5000 euros” became a refund of the order total, flagged for the reviewer.

What to look at, by question
- Why did this ticket get this reply? Its trace: the prompt, the data looked up, the answer.
- Is a prompt change better or worse? Datasets › agentdesk-golden › Experiments: two runs side by side.
- What does it cost? Dashboards: cost and tokens per model, per day.
- Where does it fail? Traces filtered by level ERROR or WARNING: retries, fallbacks, guard blocks.
03
When things go wrong
Failures are expected and designed for, in the order a ticket meets them. Try “simulate outage” on the top bar, then watch Runs and Queue.
In plain words
| What fails | Detected by | What the system does | Where you see it |
|---|---|---|---|
| Same message delivered twice | Unique external_id on tickets | Second insert is a no-op; the API answers with the existing ticket and duplicate: true. | API response |
| Invalid request to the API | Pydantic validation, bearer token check | 422 with the field errors, or 401. Nothing is written. | API response, n8n “Rejected (400)” branch |
| Model provider slow or down (timeout, 429, 5xx) | Transient ProviderError in the router | Retried on the same provider with jittered backoff, then the next provider in the chain. | retry and fallback steps in Runs, Langfuse spans |
| Provider keeps failing | Three consecutive failures | Its circuit opens: calls skip it for 30s, then one probe decides whether it closes. | Top bar, Telegram alert (deduplicated) |
| Bad key or bad request to a provider (4xx) | Permanent ProviderError | No retry on that provider; straight to the next one. | error step in the run |
| Every provider failed | AllProvidersFailed | The attempt fails; the job is rescheduled with backoff (5s, 10s, 20s). | Queue page, failed run |
| Model answer does not match the schema | Pydantic validation of the tool call | The error goes back to the model for one repair turn; a second failure fails the attempt. | guard step “schema” in the run |
| Model calls a tool it was not granted | Gateway grant check | Refused; the model gets an error result and continues. The database role could not have run it anyway. | blocked tool step |
| Model invents an order, over-refunds, or leaks another customer | Guards, checked against the database | The run is blocked and the ticket goes to a person; nothing is proposed. | guard step, audit log “guard blocked”, Overview counter |
| Prompt injection in the customer message | Pattern check on the ticket | The proposal carries a flag so the reviewer reads it twice; tools and grants are unchanged. | ⚑ flag on the approval card |
| Worker crashes mid-run | Lease expiry (5 minutes) | The job is reclaimed and runs again; finished stages are skipped, effects were never half-written. | Queue page, “lease expired” as last error |
| Job out of attempts | attempts ≥ max_attempts | Moved to the dead letter queue; the ticket is marked failed. | Queue page with Retry, Telegram alert, audit log |
| Proposal no longer valid when approved | Re-validation in the executor (row locked) | Not executed; the proposal is marked failed with the reason. | Approvals “recently decided”, 409 from the API |
| Approved twice (double click, dashboard and Telegram) | Proposal status checked under a row lock | Second decision is a no-op: “already decided”. | Telegram message, API response |
| Telegram unreachable | Send error | The proposal stays unannounced and is retried on the next pass; nothing is lost. | notifications table |
| Langfuse unreachable | Ingestion error in the background thread | Logged and dropped. Tracing never blocks or fails a run; the database copy of every step remains. | worker log |