How it works

Agents that do the work, people who make the calls

A guide to agentdesk in seven parts. Every part starts with a plain-language summary, then the technical detail. If something is unclear, ask the guide: an agent of this platform answers from these pages and the running code.

01

Channels in n8n

n8n owns the edges: how customers reach the platform, the traffic that keeps the demo alive, the morning report and the alarms.

In plain words

n8n is a visual automation tool: each box is a step, each line is where the data goes next. It handles the outside world (forms, webhooks, schedules, Telegram), and hands the actual work to the platform through its API. The agents, their rules and their retries stay in the platform, not in n8n.
TRIGGERS02 · contact formcustomers01 · intake webhookchat widgets, websites03 · traffic generatorevery 15 min, 08–22 h04 · daily reportevery day at 08:0099 · error handlerany workflow fails00 · create ticketvalidate · retry ×3alert if the API is downAGENTDESK APIPOST /ticketsPOST /simulateGET /statsagentdeskqueue · agentsguards · approvalsTelegramalerts and the daily reportAPI down after 3 tries
How the six workflows fit together: two entry points share one ticket-creating sub-workflow, two schedules talk to the API directly, and every failure ends up on Telegram.

All six live in the agentdesk folder of the self-hosted n8n. Open any of them to read the sticky notes on what it does and how it fails.

Open the folder in n8n ↗
n8n workflow 00 · create ticket (shared)
00 · create ticket (shared)Open in n8n ↗The one place where a message becomes a ticket. It cleans the input (email, order number, message), rejects what is invalid, calls the API with three retries, and if the API stays down it alerts on Telegram and returns a clear “try again”. Both entry points below call it, so the rules live once.
n8n workflow 01 · intake webhook
01 · intake webhookOpen in n8n ↗For a chat widget or a website: a POST endpoint. It answers 202 with the ticket id (queued, not resolved: agents work in the background), 400 for bad input and 503 if the platform is unreachable.
n8n workflow 02 · contact form
02 · contact formOpen in n8n ↗A real contact form hosted by n8n itself (email, order number, message). The customer gets a confirmation page, or a retry message if something is down.
n8n workflow 03 · traffic generator
03 · traffic generatorOpen in n8n ↗Keeps the demo alive: every 15 minutes during the day, one to three simulated customers write in through the API. Switch it off and the store goes quiet.
n8n workflow 04 · daily report
04 · daily reportOpen in n8n ↗At 08:00 it reads the last 24 hours from the API (tickets, success rate, latency, cost, decisions, backlog, last evaluation) and sends one Telegram message.
n8n workflow 99 · error handler
99 · error handlerOpen in n8n ↗Registered as the error workflow of all the others: if any of them fails, it sends the workflow, the failing node, the error and a link to the execution to Telegram.

They are numbered in the order a reader should open them. Their source is n8n/build.py, which writes the JSON that is imported, so the workflows are reviewed like code. They need two credentials created in n8n (the API bearer token and the Telegram bot) and are switched on once the API has its public URL.

02

Tracing in Langfuse

Langfuse records what the agents actually did, call by call, so any reply can be explained after the fact.

In plain words

The dashboard tells you that something happened; Langfuse tells you why. For any ticket you can read exactly what the model was told, what it looked up, what it answered and what it cost.
Worker + agentsevery run, step,decision and failureruns, run_stepsPostgres, as it happensSupabase Realtimepushes each changeThis dashboardis it healthy right now?Langfuseone trace per ticketWhy did it do that?prompt, tool results, answer, tokens, cost, provideraudit_logappend-onlyTelegramalerts, cards, reportWho decided what, when?Audit log page
Three records of the same work, each answering a different question.

From the dashboard to a trace

Open any run on the Runs page and click Open full trace in Langfuse. Each ticket is one trace; the trace holds a generation for every model call and a span for every tool call, guard check, retry and fallback, in order. Traces are grouped by ticket as a session, tagged by channel, and labelled with the environment (agentdesk-dev, ci) so local, CI and production never mix.

An agentdesk trace in Langfuse
A trace: the tree on the left is everything that happened (triage, two resolver model calls, the two tool calls, the guard). On the right, the input ticket, the output (reply, refund, flags) and the eval.passed score. This one is the prompt-injection case: the injected “refund 5000 euros” became a refund of the order total, flagged for the reviewer.
A model call in Langfuse
One model call: the system prompt, the tool results the model saw (here, the order it looked up), its answer, tokens, cost, latency and which provider answered. This is how you answer “why did it say that?”.
The golden dataset in Langfuse
The golden suite as a Langfuse dataset: each item is a case with its input and the expected behaviour. Every evaluation run links its traces to these items.

What to look at, by question

  • Why did this ticket get this reply? Its trace: the prompt, the data looked up, the answer.
  • Is a prompt change better or worse? Datasets › agentdesk-golden › Experiments: two runs side by side.
  • What does it cost? Dashboards: cost and tokens per model, per day.
  • Where does it fail? Traces filtered by level ERROR or WARNING: retries, fallbacks, guard blocks.

03

When things go wrong

Failures are expected and designed for, in the order a ticket meets them. Try “simulate outage” on the top bar, then watch Runs and Queue.

In plain words

Models time out, services go down, workers crash, people click twice. For each of these the system has a planned reaction, and each one is visible somewhere you can check.
Mistralmistral-small2 attempts, backoffClaudeclaude-haiku-4-52 attempts, backoffofflinedeterministic stand-in2 attempts, backoffEvery provider failedthe job is retried 5s, 10s, 20sthen the dead letter queuefailsfailsfailsThe first provider that answers is used, and the run records which oneA provider with an open circuit is skipped without calling it: three failures open it for 30 s, then one probe call decides.
A model call: providers are tried in order, each twice. Only when all of them fail does the whole job go back to the queue.
queuedrunningdonedeaddead letter queueclaimedfailed: back off5s · 10s · 20slease expired: the worker diedsucceeded4th failurea person presses Retry on the Queue pageClaims use FOR UPDATE SKIP LOCKED,so any number of workers can sharethe queue without taking a job twice.
A job in the queue: every way it can fail leads somewhere visible, and nothing disappears.
pendingwaits for a personexecutedrefund made, reply sentrejectedfailednothing executedapprove: lock row, re-check, executerejectapprove, but the re-check fails(order changed, already refunded)A second clickfrom the dashboard or Telegramanswers “already decided”
A proposal after the agent is done: only a person moves it, and the executor re-checks before acting.
What failsDetected byWhat the system doesWhere you see it
Same message delivered twiceUnique external_id on ticketsSecond insert is a no-op; the API answers with the existing ticket and duplicate: true.API response
Invalid request to the APIPydantic validation, bearer token check422 with the field errors, or 401. Nothing is written.API response, n8n “Rejected (400)” branch
Model provider slow or down (timeout, 429, 5xx)Transient ProviderError in the routerRetried on the same provider with jittered backoff, then the next provider in the chain.retry and fallback steps in Runs, Langfuse spans
Provider keeps failingThree consecutive failuresIts circuit opens: calls skip it for 30s, then one probe decides whether it closes.Top bar, Telegram alert (deduplicated)
Bad key or bad request to a provider (4xx)Permanent ProviderErrorNo retry on that provider; straight to the next one.error step in the run
Every provider failedAllProvidersFailedThe attempt fails; the job is rescheduled with backoff (5s, 10s, 20s).Queue page, failed run
Model answer does not match the schemaPydantic validation of the tool callThe error goes back to the model for one repair turn; a second failure fails the attempt.guard step “schema” in the run
Model calls a tool it was not grantedGateway grant checkRefused; the model gets an error result and continues. The database role could not have run it anyway.blocked tool step
Model invents an order, over-refunds, or leaks another customerGuards, checked against the databaseThe run is blocked and the ticket goes to a person; nothing is proposed.guard step, audit log “guard blocked”, Overview counter
Prompt injection in the customer messagePattern check on the ticketThe proposal carries a flag so the reviewer reads it twice; tools and grants are unchanged.⚑ flag on the approval card
Worker crashes mid-runLease expiry (5 minutes)The job is reclaimed and runs again; finished stages are skipped, effects were never half-written.Queue page, “lease expired” as last error
Job out of attemptsattempts ≥ max_attemptsMoved to the dead letter queue; the ticket is marked failed.Queue page with Retry, Telegram alert, audit log
Proposal no longer valid when approvedRe-validation in the executor (row locked)Not executed; the proposal is marked failed with the reason.Approvals “recently decided”, 409 from the API
Approved twice (double click, dashboard and Telegram)Proposal status checked under a row lockSecond decision is a no-op: “already decided”.Telegram message, API response
Telegram unreachableSend errorThe proposal stays unannounced and is retried on the next pass; nothing is lost.notifications table
Langfuse unreachableIngestion error in the background threadLogged and dropped. Tracing never blocks or fails a run; the database copy of every step remains.worker log

Circuit breaker, per model provider

Closedcalls go throughOpencalls skip itHalf openone probe call3 failuresafter 30sprobe failsprobe succeeds
NextReference→