Tutorial 3 of 3 · previous: a pipeline
The pipeline in tutorial 2 ran the same steps in the same order every time. This one decides. A support message comes in, a classifier says whether it is a bug report or a question, and the message goes down one branch: bugs become a ticket for engineers, questions get a reply drafted from the help docs. When the classifier can’t tell, the run stops and asks a person, and it can wait for them as long as it takes, through a server restart.
The page is built in stops. After each one you have something that runs, so you can stop after any of them:
- A classifier whose answer is JSON the platform has checked.
- Two handlers, one per kind of message.
- Routing: a graph that sends each message to one handler.
- A person for the messages the classifier can’t place.
- A restart in the middle of a run that is waiting on that person.
Every agent on this page is also a JSON file you can download. Each step says which one.
Before you start
This page assumes you did tutorial 2: you have made an
agent in All Agents, built a graph in the composer, and read a run in
Run History. It does not use changelog-writer.
It needs version 0.2.10 or later, for the docs folder in step 2 and
{{run.input}} in step 3:
npm install -g @upship/agents
agent start
The example is the support inbox of Tally, a made-up shared to-do list app.
Its help docs are four short markdown files. Put them in a folder under
~/code, which the install allows tools to read:
mkdir -p ~/code/helpdesk/docs && cd ~/code/helpdesk/docs
for f in accounts billing exporting sharing; do
curl -sO https://khtdr.com/agents/router/docs/$f.md
done
If your projects live somewhere else, put the folder there instead, or run
agent add-data-dir ~/code/helpdesk to allow it.
1. A classifier
Open All Agents, click New Agent, and fill in:
- Name:
triage-classifier - Description:
Classifies a support message as a bug report, a question, or neither, and says how sure it is - Model Tier:
fast - Max Steps:
1 - System Prompt:
You sort messages sent to the support inbox of Tally, a shared to-do list app.
Decide what the message is:
- bug: something in Tally is broken or behaves differently from how it should.
- question: the person wants to know how to do something, or how something works.
- other: neither, such as a sales pitch, a thank-you, or a message too vague to place.
Give a confidence from 0 to 1 for your choice. Use the whole range. A message
that could honestly be read either way is below 0.6, even if you lean one way.
The new part is Output Schema, above Tools. Click + Classifier starter,
then change the enum to this page’s three kinds and put limits on the
confidence:
{
"type": "object",
"properties": {
"intent": { "type": "string", "enum": ["bug", "question", "other"] },
"confidence": { "type": "number", "minimum": 0, "maximum": 1 }
},
"required": ["intent", "confidence"],
"additionalProperties": false
}
Pick no tools, and click Create Agent. (Or skip the form:
triage-classifier.json, and
curl -s localhost:2137/api/agents -H 'content-type: application/json' -d @triage-classifier.json.)
Run it:
agent triage-classifier "How do I turn on two-step sign-in?"
╭─ triage-classifier ─── streaming
│
│ {
│ "intent": "question",
│ "confidence": 0.95
│ }
│
╰─ ✓ 1 steps · 3.8s · 722 tokens · 561b550d
That looks like a model that was asked for JSON and returned some. The
difference is what happened before you saw it. The platform parsed the answer
and checked it against the schema. intent is one of the three words, and
confidence is a number between 0 and 1. Only then did the run finish. Open
the run in Run History: its output is a block marked Structured Output
and validated, not text. The next step reads fields out of that value,
and it can because the value is guaranteed to have them.
When the model gets it wrong. If the answer is not JSON, or
it breaks the schema, the platform sends it back once with the complaint:
/intent: must be equal to one of the allowed values. One retry
usually fixes a drifted field. If the second answer breaks the schema too, the
run ends as an error rather than handing on text that only looks right. The
retry costs a second model call, and the run's token count includes it.
The mistake this saves you from is routing on prose. With no schema, the classifier would answer “This looks like a bug report”, and the next step would have to search that sentence for the word “bug”. That works until the model writes “This is not a bug”.
2. Two handlers
Each kind of message gets an agent of its own. Neither one knows anything about routing. Each is handed a customer’s message and does one job with it.
The bug writer. New Agent:
- Name:
bug-writer - Description:
Turns a support message that reports a bug into a bug ticket for the engineering queue - Model Tier:
default, Max Steps:1, no tools - System Prompt:
You turn a customer's support message about Tally, a shared to-do list app,
into a bug ticket for engineers.
Write markdown with these parts:
# <one-line title naming what is broken>
**Reported:** the customer's own words, quoted.
**What happens:** what they see.
**Expected:** what should happen instead.
**Unknowns:** what an engineer would need to ask them, as a short list.
Use only what the message says. Do not guess at causes, and do not invent
steps the customer did not describe.
The docs answerer. It needs to read the help docs, and nothing else. The
file tools can read anywhere the install allows, which includes all of
~/code, so this step makes two narrower ones: a search and a read that only
see the docs folder. They are tool definitions, the same layer you met in
tutorial 1’s memory exercise: a row that points at a built-in tool and fixes
part of its config.
The folder has to be a full path. Find yours:
echo ~/code/helpdesk/docs
Open Tools, click New Definition, and fill in:
- Name:
search-helpdesk-docs - Implementation:
search_files - Description:
Searches the Tally help docs for a word or phrase and returns the matching lines with their file names. Search '.' to cover every doc. - Config, with your path:
{ "root": "/Users/you/code/helpdesk/docs" }
Then a second one:
- Name:
read-helpdesk-doc - Implementation:
read_file - Description:
Reads one Tally help doc in full, by its file name, such as billing.md. - Config: the same
root.
(search-helpdesk-docs.json and
read-helpdesk-doc.json, posted to
/api/tools. Change the path first.)
root is the folder the model’s paths start from, and a wall. The model asks
for billing.md, not a full path, and a path that climbs out, like
../../.ssh/id_rsa, is refused even though the install would allow it. The
refusal goes back to the model as the tool’s result, so it can try something
else, and the run carries on. A root outside what the install allows fails
the run as it starts, naming the folder, before any model is called.
Now the agent. New Agent:
- Name:
docs-answerer - Description:
Drafts a reply to a customer's question about Tally from the help docs - Model Tier:
default, Max Steps:8 - Tools:
search-helpdesk-docsandread-helpdesk-doc - System Prompt:
You draft replies to customers' questions about Tally, a shared to-do list app.
Answer only from the help docs. search-helpdesk-docs finds which doc mentions
something, and read-helpdesk-doc reads a doc by its file name. Search at most
three times, then read the whole doc the best match came from. The docs are
short, so if three searches find nothing that answers the question, the docs
do not cover it.
Reply with only this markdown, and nothing before it:
# Re: <the question, in a few words>
The reply to the customer, in two to five sentences, in plain words.
**Source:** the doc file or files the answer came from.
If the docs do not answer the question, say so in the reply, and write
**Source:** none.
The prompt doesn’t say where the docs are. The tools already know.
Try it alone:
agent docs-answerer "Can I export attachments too?"
╭─ docs-answerer ─── streaming
│
│ ⚡ search-helpdesk-docs {"pattern":"export.*attachment","path":".","maxResults":50}
│ ↳ done
│ ⚡ search-helpdesk-docs {"pattern":"attachment.*export","path":".","maxResults":50}
│ ↳ done
│ ⚡ search-helpdesk-docs {"pattern":"export","path":".","maxResults":50}
│ ↳ done
│ ⚡ read-helpdesk-doc {"path":"exporting.md"}
│ ↳ done
│ # Re: Exporting attachments
│
│ No, exports do not include attachments. You'll need to download those directly from each task before exporting if you want to save them.
│
│ **Source:** exporting.md
│
╰─ ✓ 3 steps · 7.5s · 5,272 tokens · 0aa5af2f
Four tool calls in three steps: the model sent some of them together, and the
third step was the answer. Each step is one model call, and Max Steps is
how many it gets. An agent that spends them all and is still calling tools has not answered, so the run
ends as an error rather than handing on half a thought. Here is the same
question with Max Steps set to 2:
│ ⚡ read-helpdesk-doc {"path":"exporting.md"}
│ ↳ done
│ Error: Agent "docs-answerer" used all 2 steps without a final answer. Raise maxSteps or narrow what it has to do.
│
╰─ ✗ 2 steps · 5.2s · 3,169 tokens · 26220abe
In a graph, that error stops the run at this node, so nothing half-finished
reaches save. The “at most three times” in the prompt gives a question the
docs don’t cover a place to stop. Asked about dark mode,
which no doc mentions, it searched three times and wrote Source: none.
3. Routing
The router is a graph agent, like release-notes in tutorial 2, with one
difference. There, every edge fired. Here, an edge can have a condition,
and it fires only when the condition holds.
New Agent:
- Name:
triage - Description:
Sorts a support message into a bug ticket or a drafted answer, and asks a person when it can't tell - System Prompt:
Sorts support messages. The graph does the work.(Required by the form, never sent to a model.)
Under Graph Pipeline, click + Add Graph, and add four nodes, in this order:
- Id
classify, Reftriage-classifier. No input. - Id
bug, Refbug-writer. One input param:input, set to{{run.input}}. - Id
answer, Refdocs-answerer. The same param:input, set to{{run.input}}. - Id
save, Refwrite_file. Two params:pathset toticket.md, andcontentleft as{{input}}.
{{run.input}} is the message the run was started with. Tutorial 2 used
{{input}}, which is whatever the node before handed on. Here that would be
the classifier’s JSON, {"intent":"bug","confidence":0.95}, and the bug writer
would write a ticket about a classification. A handler needs the customer’s
words, so it reaches past the node before it to the run’s own input.
Now the edges. Until now the graph ran in list order. Drag from the handle on
the right of classify to the one on the left of answer. The first edge you draw writes the
list order down as edges too, so the canvas now has four:
classify → bug, bug → answer, answer → save, and the one you drew. One of
them is wrong. Click bug → answer and change To to save. Every node is
now wired: classify to each handler, and each handler to save.
Click classify → bug and give it a Condition:
output.intent == 'bug' && output.confidence >= 0.8
and classify → answer:
output.intent == 'question' && output.confidence >= 0.8
output is the classifier’s checked value from step 1. Save. This is the
graph, and what GET /api/agents/triage returns under graph:
{
"nodes": [
{ "id": "classify", "ref": "triage-classifier" },
{ "id": "bug", "ref": "bug-writer", "input": { "input": "{{run.input}}" } },
{ "id": "answer", "ref": "docs-answerer", "input": { "input": "{{run.input}}" } },
{ "id": "save", "ref": "write_file", "input": { "path": "ticket.md", "content": "{{input}}" } }
],
"edges": [
{ "from": "classify", "to": "bug", "condition": "output.intent == 'bug' && output.confidence >= 0.8" },
{ "from": "classify", "to": "answer", "condition": "output.intent == 'question' && output.confidence >= 0.8" },
{ "from": "bug", "to": "save" },
{ "from": "answer", "to": "save" }
]
}
Send it a question:
agent triage "Can I pay yearly?"
│ ⚡ classify (triage-classifier)
│ ↳ 0.9s
│ ⊘ bug skipped
│ ⚡ answer (docs-answerer)
│ ↳ 7.4s
│ ⚡ save (write_file)
│ ↳ 0.0s
│ {"path":"/Users/you/.agents/output/7d017a71-…/ticket.md","mimeType":"text/markdown","bytesWritten":258}
│
╰─ ✓ 3 steps · 8.3s · 4,672 tokens · 7d017a71
bug was skipped. Its only way in was an edge whose condition was false,
so it could never run, and the runner said so and moved on. It cost nothing:
no model was called for it.
Look at save. It has two edges coming in, and only one of them fired. A
runner that waits for every parent before starting a node would wait here
forever for bug. This one waits until each incoming edge has either fired or
been ruled out, then runs on the ones that fired. Ruling out spreads: a skipped
node’s own edges are ruled out in turn. So save got the answer, and nothing
else. When more than one parent fires, their outputs are joined, which is the
concat in the composer’s Merge (fan-in) box. Here only one branch ever
fires, so there is nothing to choose between.
A typo in a condition does not wait for a run to find it. Click
classify → bug and change == to =. The box turns red as you type, because
the composer runs the same parser the server does. Save anyway, and the server
refuses the whole graph, naming the edge:
{"error":"Validation failed","fields":{"graph.edges.0.condition":["Unexpected character \"=\""]}}
A condition is not JavaScript. It is a small fixed language: comparisons, &&,
||, !, field access, and a few read-only string methods like includes.
It can’t call anything or change anything, which is why it is safe to accept
over HTTP. Docs → Graphs in the install’s own UI lists all of it. Put the ==
back, and save.
Now send a message that is hard to place:
agent triage "My lists look weird since yesterday."
│ ⚡ classify (triage-classifier)
│ ↳ 3.0s
│ ⊘ bug skipped
│ ⊘ answer skipped
│ ⊘ save skipped
│ Error: Graph "triage" produced no output: every terminal branch was skipped by an edge condition.
│
╰─ ✗ 1 steps · 3.0s · 829 tokens · 63fb5faf
Open the run and look at the classifier’s answer: bug, at 0.75. Below the
line, so neither edge fired, both handlers were skipped, and the skipping
spread to save. A graph that ends with nothing is a bug in its conditions,
not a result, so the run fails and says why. It still keeps what did run: the
classifier’s step, with the answer that explains the skip, and the tokens it
cost. That hole is what the next step fills.
Why 0.8 and not 0.6? The classifier's prompt says 0.6, and
it doesn't mean it. In testing, the fast tier never went below
0.65, even for "export??". Clear messages came back at 0.85 to 0.95, and vague
ones at 0.75. A threshold of 0.6 would never have fired. A model's confidence
is its opinion, not a probability, so set the line by reading what yours
actually says in Run History, not by what the prompt asks
for.
4. A person
When the classifier can’t tell, ask someone who can. ask_human is a tool that
stops the run and puts a question to a person. The question is not something
the model writes. It is fixed config on a tool definition, like the docs
folder in step 2, so each question is a row with its own name and version.
Open Tools, click New Definition, and fill in:
- Name:
which-bucket - Implementation:
ask_human - Description:
Asks a person whether a support message is a bug report or a question - Config:
{
"question": "Is this message a bug report or a question?",
"cardinality": "single",
"options": ["bug", "question"]
}
(which-bucket.json, posted to
/api/tools.)
cardinality: single means one of the options, and nothing else. The platform
works out what a valid answer looks like from the options, so an answer of
feature is refused before it reaches the run.
Now edit triage. + Add Node: Id ask, Ref which-bucket, and
one input param, context, set to {{run.input}}. That is the one thing the
person needs beside the question: the message itself.
Draw three edges and give each a condition:
| From | To | Condition |
|---|---|---|
classify | ask | output.intent == 'other' || output.confidence < 0.8 |
ask | bug | output.answer == 'bug' |
ask | answer | output.answer == 'question' |
The first is the exact opposite of the other two edges out of classify, so
every message goes exactly one way. The other two route on the person’s answer.
A person answering ask produces { "answer": "bug" }, a checked value, like
the classifier’s. An edge can’t tell which of the two it is reading, and doesn’t
need to. The person is a classifier that takes longer to answer.
This is also why the handlers read {{run.input}}. When a person routed the
message, the node before bug is ask, and all it hands on is
{ "answer": "bug" }. {{run.input}} gets the handler the message whichever
way it came.
Save. The whole graph is
triage.json. Send the vague message again:
agent triage "My lists look weird since yesterday."
│ ⚡ classify (triage-classifier)
│ ↳ 3.0s
│ ⚡ ask (which-bucket)
│ ⏸ ask waiting token 3f536eff-276e-4231-aa2e-6968c68301f6
│ Waiting on which-bucket (token 3f536eff-276e-4231-aa2e-6968c68301f6).
│
╰─ ⏸ 1 steps · 3.0s · e623d300
⏸ which-bucket (node ask)
Is this message a bug report or a question?
My lists look weird since yesterday.
1) bug
2) question
Choose 1-2:
The terminal asks because you started the run from it, and you are the nearest person. Don’t answer. Press Ctrl-C.
5. A restart
The run is not in your terminal, and it is not in the server’s memory either.
When ask stopped, the run saved what it had so far to its own row, the
classifier’s answer, and marked itself paused. Nothing is holding it open.
To prove it, stop the server:
agent stop
agent start
agent pending triage
Waiting (1)
triage e623d300 · 0m
⏸ Is this message a bug report or a question?
Answer triage e623d300? [y/N]
It is still there. It would be there tomorrow. A paused run has no deadline
unless its tool definition sets one (timeoutMs), and the process that watches
for crashed runs leaves paused ones alone. In the browser, Run History
shows it as paused. Open it, and the question is on the run’s page with an
Answer button.
Answer it here, in the terminal: y, then 1.
Answer triage e623d300? [y/N] y
⏸ which-bucket (node ask)
Is this message a bug report or a question?
My lists look weird since yesterday.
1) bug
2) question
Choose 1-2: 1
✓ resumed
{"path":"/Users/you/.agents/output/e623d300-…/ticket.md","mimeType":"text/markdown","bytesWritten":712}
cat ~/.agents/output/e623d300-*/ticket.md
# Lists displaying incorrectly since yesterday
**Reported:** "My lists look weird since yesterday."
**What happens:** Lists are displaying in an abnormal way as of yesterday.
**Expected:** Lists should display normally as they did before yesterday.
**Unknowns:**
- What specifically looks "weird" about the lists? (layout, formatting, colors, fonts, spacing, missing elements, etc.)
- Which lists are affected? (all lists, specific lists, shared lists, personal lists)
- What platform/device are they using? (web, iOS, Android, desktop app)
...
Now open the run in Run History and expand it. It has two children:
triage-classifier, which ran before the restart, and bug-writer, which ran
after it. There is only one classifier run. On resume, the graph started again
from the top, and every node that already had an output was served from what
was saved instead of being run. classify cost one model call, not two. Then
ask took your answer as its output, the ask → bug edge fired, answer was
skipped, and the rest ran as if nobody had waited.
The run’s steps show the answer too. Between the question and bug-writer is
a step named which-bucket:resolved, holding { "answer": "bug" } and when
it was given, so the record of the run says what the person decided, not only that
someone did.
The run’s Duration includes the wait, because the run really did take that long. Its cost doesn’t: waiting is free.
Why ask is a node in this graph. A run can
pause only at its own top level. If the person's question were asked from
inside one of the handlers, the handler's run would pause under a parent that
has no way to wait for it, and the parent node would fail. That is
nested parking,
which isn't built yet. Put the question where the graph can see it.
What it cost
Three runs from testing, as the Pipeline card on each triage run reported
them:
- A clear bug: about 0.4¢, 8 to 9 seconds. The classifier on
fastand the bug writer ondefault, one model call each. - A clear question: about 1.6¢, 8 to 9 seconds. The docs answerer searches and reads, and each of those is another call with the conversation so far.
- A person decided it was a bug: about 0.4¢, the same as a clear bug. The wait added nothing, and the classifier was not paid for twice.
A skipped branch costs nothing, which is the case for routing over one agent that handles every kind of message. The bug path never pays for the docs search, and nothing pays for a model to decide which path to take after the classifier has.
Now your turn
The classifier has a third answer, other, and so far only the person has
anything to do with it. Give other somewhere to go without asking:
- Add a third option to
which-bucket,neither, and a branch for it: awrite_filenode that saves{{run.input}}asneeds-reply.md, with an edge fromaskonoutput.answer == 'neither'. No new agent. Note that the new node is a second place the graph can end. What happens tosavewhen it is the branch that runs? - Set the threshold from history. Send the classifier twenty messages of your own, clear and vague, and read the confidences in Run History. Where would you put the line so that a person sees the vague ones and never the clear ones?
- Answer one from the browser. Start a vague run, open it in Run History, and answer it there instead of from the terminal.
Where to go from here: interview is ask_human’s multi-turn version, for
when the next question depends on the last answer. Debate and reflection have
several agents check or compete on one answer. All three are on the
features page.
When you are done:
agent stop
Back to the overview
What the platform does and the full feature list.
How it compares
Against other agent frameworks, including what they do that this doesn't.