Agents
An agent is a model given a task, a set of tools and a budget, left to decide for itself which tools to call until it can answer. The agent host runs them against your Fallen-8.
It is the other side of the MCP server. That one exposes this graph’s tools to somebody else’s agent, in somebody else’s client. This one runs agents here, and reaches the graph as a client of that same MCP server, so an agent can call exactly the tools you enabled there and nothing else.
It is a separate deployable (fallen-8-agents), its own process and container image, and it
holds two things deliberately: no model configuration and no provider credential. Every model
call goes to your instance’s own POST /chat with purpose: agent, so you configure a model once,
on the instance, and the host inherits it. Point that instance at Nahil, at a local Ollama, at
OpenAI or at Anthropic and the agents follow, with no second place to change.
What you get for a run
Section titled “What you get for a run”Every agent leaves three things behind, and they are the reason this exists rather than a chat loop in a script:
- A bounded trace. Every model call (with the backend and model that served it), every tool
call (with the arguments and result, byte-capped), the state changes, and the mechanical citation
count. Past
Agents:Trace:MaxStepsthe oldest steps go and a marker row says how many, so a short run is never mistaken for a truncated one. - An event feed. Server-sent events in the same dialect as the change feed, so
a client that can read one can read the other:
id:/event:/data:, keep-alive comments, and declarativeagentsandkindsfilters. Every event carries the live counters, so a subscriber renders cost without polling. - Hard budgets. Model calls, tool calls, wall clock and tokens, each enforced in this process,
because a model that is looping is exactly the model that will not honour an instruction to stop.
A breach ends the run as
budgetExceededand names which budget.
Nothing here is durable. A restart ends every agent and forgets every finished one, which is why every listing and every trace names the host instance that produced it: two traces from different processes are never silently compared.
Running one
Section titled “Running one”# Turn it on. This brings up the sidecar AND opens the instance's /agents routes.F8_AGENTS=true npm run env:up
# Start an agent. 202 with its id: nothing on this control plane waits for a model,# because one call against a remote provider was measured at up to 41 seconds.curl -X POST http://localhost:8080/agents \ -H 'Content-Type: application/json' \ -d '{"task":"How many vertices are in the default namespace?"}'
# Watch it work.curl -N http://localhost:8080/agents/feed
# Read what it did.curl http://localhost:8080/agents/<id>curl http://localhost:8080/agents/<id>/traceOn a secured instance every one of those carries your API key, exactly as the rest of the API
does. With agents off they answer 403 on a keyed instance and 401 on a keyless one: a
client that treats only 403 as “this feature is absent” will misread a bare dotnet run.
A spawn names a role, which is a prompt plus a tool allowlist. The allowlist narrows what the MCP server advertises and can never reach past it, and it is applied to the tool list the model is handed, so a tool outside it is not merely discouraged: the model never sees it.
| Role | Sees | For |
|---|---|---|
assistant |
every tool the MCP server advertises | the default: one agent, one task, one answer |
orchestrator |
the overview read, plus the two swarm tools | breaking a task into parts and composing the results |
worker |
every tool the MCP server advertises | one separable part of a larger task, spawned by an orchestrator |
A worker cannot be spawned over the API. It reports a typed result to the orchestrator that gave
it its part of a task, so one started by hand would have nobody to report to.
Agents:Roles:<role>:Tools replaces a role’s shipped allowlist rather than adding to it:
absent or empty keeps what the role ships with (so an empty list leaves orchestrator on the
overview read), a list of tool names means exactly those, and a lone * means every tool the
server advertises whatever the role ships with, which no list of names can say without enumerating
them. What the server advertises is still the outer bound. A list holding nothing but blanks counts
as empty, and * beside a tool name is refused at startup rather than guessed at. A key naming no
role is refused there too, so a misspelled role name cannot leave the role it was meant to narrow
holding everything. The MCP tiers bound all of them.
Swarm mode
Section titled “Swarm mode”An orchestrator gets two extra tools: spawn_worker(task, name?) and await_workers(ids?). A
worker is a real agent in every sense that matters: listed, traced, budgeted and cancellable
like any other, with its own token budget rather than a slice of its orchestrator’s.
await_workers returns each worker’s typed result (state, answer, citation counts, cost),
never prose about what the worker did. That is deliberate: an orchestrator handed prose would have
to parse another model’s writing, and because it is the only composer, whatever it cannot read
reliably it would paraphrase.
Three caps bound a swarm, all configuration and all enforced when an agent is admitted:
| Cap | Default | Bounds |
|---|---|---|
Agents:Limits:MaxConcurrentAgents |
4 | how many agents run at once on this host |
Agents:Limits:MaxSwarmDepth |
2 | how deep it nests; 2 means “delegate once”, so workers cannot orchestrate in turn |
Agents:Limits:MaxWorkersPerOrchestrator |
4 | how many workers one orchestrator spawns over its whole life |
The last one counts over a life rather than at once on purpose. An orchestrator’s token budget bounds its own calls, not its workers’, so a live-only count would let it spawn its allowance, await them, and spawn again without limit.
A breached cap comes back to the model as a tool error it can act on (delegate less, compose what
it has) and to you as a toolCalled event with success: false and the reason in failure, so
the event names the cap rather than leaving you to find it on the trace. The same is true of a tool
that fails for any other reason: a graph tool reports its failure in its result rather than by
throwing, and that is recorded as the failed call it is.
A worker does not outlive its orchestrator. However an orchestrator ends, whether it answers,
fails, is cancelled or runs out of budget, its live workers are cancelled and each row says which
ending stopped it. Only the orchestrator reads a worker’s result, so a worker whose orchestrator has
gone would spend its budget and hold a concurrency slot for an answer nobody will compose. The same
rule refuses a late delegation: an orchestrator that has already ended cannot take a new worker, and
spawn_worker says so.
Exactly one composer
Section titled “Exactly one composer”Only the orchestrator’s final text is the answer. Workers write for their orchestrator, and the swarm’s mechanics stay where mechanics belong: agent names, spawns and awaits are feed events and trace steps, never narration in the result. The role prompts say so, and they are embedded in the image rather than mounted, because a deployment that could edit or lose a prompt could quietly turn an agent into a confident liar.
Citations, and what they do and do not mean
Section titled “Citations, and what they do and do not mean”The role prompts require every figure, name or id in an answer to carry the tool call it came from,
written [t:<name>]. The host then counts those citations against the calls the run actually made
and records the pair on the trace and the ending event.
A call that failed does not count as a call that was made. A refused spawn, or a graph read that answered 401, is work that did not happen, so citing its name is dangling rather than grounded. That is the case the count exists to expose: an answer whose every tool call failed cannot score as grounded evidence.
It is a count, never a judgement, and the numbers are worth reading precisely:
- A high
validcount is not a correct answer. An agent can cite a real call and still misread it. - A
danglingcount is not proof of a lie. The model may have named a tool it holds but did not call, or one that does not exist. - What the pair is good for is the shape a fabricating run has: every figure asserted, no citations at all, and a trace with no tool calls in it.
Security posture
Section titled “Security posture”- No host port, and one authenticated way in. The container publishes none, so the browser and
your scripts reach the host through the API’s authenticated proxy at
/agents/*. Inside the compose network it is the same story as the integrations runtime: every service there can reach the host, it authenticates nobody, and it holds your API key and the MCP bearer. Security is the one home for what that means and what to do about it. - Bounded captures. A spawn’s
taskmay be at most 8192 bytes, itsnameat most 256 and itssystemPromptAppendixat most 4096, all UTF-8, and each is refused with a 400 naming the bound rather than truncated: a cut instruction would change what the agent answers, and a name rides on every feed event to every subscriber. The route’s own 1 MiB body bound is the transport backstop behind them. - One REST family. The host calls this instance’s
/chatand nothing else; a convention test enforces it, so it cannot quietly grow a graph route of its own. - The MCP tiers are the boundary. An agent can do what that server advertises. Leaving write and admin off is the recommended posture: an agent that can only read cannot be talked into a write.
- Prompt injection is real and is not solved here. A graph’s own content becomes part of what a model reads, so a hostile value in a property can try to steer an agent. The defences that actually hold are the allowlist, the MCP tiers and the budgets, all enforced outside the model; the prompt is not one of them.
- No caller text in telemetry. The host’s meter tags by closed sets only: role, token
direction, outcome, backend, tool name, success. No task and no name you sent reaches a tag or a
span, because the framework is given the agent’s ROLE as its telemetry name. What does travel is
the agent id: Microsoft Agent Framework puts it in its
invoke_agentspan’s name and ingen_ai.agent.id, which is what ties a trace to a run. A collector that derives metrics from span names would therefore see one series per run; the shipped one bounds that name before it does.
Configuration
Section titled “Configuration”Host (Agents:* for the sidecar, Fallen8Target:* for the instance it asks):
| Key | Default | Note |
|---|---|---|
Agents:BindAddress / Port |
127.0.0.1 / 8120 |
the image sets 0.0.0.0; the compose service publishes no host port, and a non-loopback bind is warned about at startup rather than refused |
Fallen8Target:BaseUrl / ApiKey / ApiKeyHeader |
http://localhost:8080 / none / X-Api-Key |
the instance every model call goes to |
Fallen8Target:TimeoutSeconds |
630 |
above the largest chat budget a shipped profile sets, so the instance’s own error wins rather than being cut off here |
Agents:Mcp:Endpoint / BearerToken |
http://localhost:8090 / none |
how an agent reaches the graph |
Agents:Mcp:ConnectTimeoutSeconds |
15 |
the startup handshake; an unreachable server leaves the toolset empty with a reason rather than failing the host |
Agents:Roles:<role>:Tools |
see above | MCP tool names, which replace the role’s shipped list; see Roles |
Agents:Limits:DefaultTokenBudget |
100000 |
what a spawn naming no budget gets |
Agents:Limits:MaxTokenBudget |
400000 |
the ceiling on a caller’s own tokenBudget, which it clamps rather than refuses |
Agents:Limits:MaxStepsPerRun |
24 |
model calls |
Agents:Limits:MaxToolCallsPerRun |
48 |
|
Agents:Limits:MaxRunSeconds |
1800 |
wall clock, measured by the host: a backend’s reported duration does not cover a remote provider’s routing |
Agents:Limits:MaxConcurrentAgents |
4 |
low because the provider quota and the instance’s rate limit are shared by all of them |
Agents:Limits:MaxSwarmDepth / MaxWorkersPerOrchestrator |
2 / 4 |
see Swarm mode |
Agents:Limits:RetainFinishedMinutes / MaxRetainedAgents |
60 / 200 |
how long a finished agent stays readable |
Agents:Trace:MaxSteps / ArgsBytes / ResultBytes |
1000 / 2048 / 8192 |
rows a trace holds (the marker is one of them) and the byte caps on a capture |
Agents:Feed:KeepAliveSeconds / MaxSubscribers / MaxQueuedEvents |
15 / 16 / 512 |
past the queue bound a subscriber is dropped rather than thinned, and its stream ends |
Agents:Observability:Otlp:Endpoint, Agents:Identity:* |
unset | OTLP push and fleet identity; see Observability |
A non-positive value switches a cap OFF for every Agents:Limits:* key, every Agents:Trace:* key
and Agents:Feed:MaxSubscribers, and for the Agents:Limits:* ones the startup line prints
unlimited rather than a bound of zero.
Four keys are floored at 1 instead, and none of them has an “off”:
Fallen8Target:TimeoutSeconds, Agents:Mcp:ConnectTimeoutSeconds, Agents:Feed:KeepAliveSeconds
and Agents:Feed:MaxQueuedEvents. A model call with no deadline, a startup handshake that never
gives up and an unbounded subscriber queue are each worse than the bound they would replace, so a
0 in them is a one second deadline, a one second handshake, a keep-alive every second and a queue
of one. Set the number you mean: to wait longer for a completion, raise
Fallen8Target:TimeoutSeconds above the instance’s own Fallen8:Chat:TimeoutSeconds.
GET /agents/status reports the deadline and the keep-alive in force, so a 0 reads back as 1
there.
The two durations among them have a ceiling as well, at about 24 days, because that is the largest delay the timers behind them can be armed with. Writing a bigger number used to be the other way to ask for “off”, and it made every model call, or every feed connection, fail with an exception naming a parameter rather than the setting.
Instance:
| Key | Default | Note |
|---|---|---|
Fallen8:Agents:Enabled |
false |
the /agents/* routes refuse until this is on |
Fallen8:Agents:Endpoint |
empty | the proxy answers 503 rather than timing out |
Fallen8:Agents:TimeoutSeconds |
30 |
the small routes, and the wait for the feed’s response headers; the feed’s body takes none of it, so a stream stays open as long as its subscriber wants |
Fallen8:Chat:<Backend>:Models:Agent |
phi4-mini:latest on Ollama and Nahil, none on OpenAI and Anthropic |
the model the agent purpose resolves to; see one model per purpose |
What it reports about itself
Section titled “What it reports about itself”GET /agents/status is the first thing to read when an agent fails: whether the host could reach
the chat gateway at startup and when it looked, what backend and model last served a step,
whether the MCP server answered and how many tools it advertises, what each role actually ended up
with, the caps a run is held to, and how many agents are active and retained.
The reachability word is one startup probe and nothing refreshes it, which is why it carries a
timestamp. A host that came up before its instance answered will say unreachable for the life of
the container while every agent runs fine, so read it against that timestamp and against the
last-seen model: a completion can only have been served by a gateway that answered.
See also
Section titled “See also”- MCP server: the tool surface an agent calls, and the tiers that bound it
- Model providers: which backend serves the agent purpose, and how to choose
- Nahil: the default remote backend, and its shared quota
- Observability: where the agent meter and the GenAI spans go
- Configuration: the instance keys and how they are written