Architecture
Fallen-8 is an in-memory graph engine with a thin REST app wrapped around it. The engine holds the graph in RAM and runs the algorithms; the app exposes it over HTTP (and, in the all-in-one image, serves the browser UI too). Three kinds of client reach it: AI agents through the MCP server, and F8 Studio (the browser UI) plus your own services straight over the REST API. Data also arrives on its own: the integrations runtime reads systems on your own network and writes what it saw in through the same REST API. This doc is the map of how the pieces fit; each piece’s contract lives in its own doc, linked below.
%%{init: {'theme':'base','themeVariables':{'fontFamily':'ui-monospace, SFMono-Regular, Menlo, Consolas, monospace','lineColor':'#666666'}}}%%
flowchart TB
agents["AI agents"]:::client
studio["F8 Studio<br/>(React SPA, browser)"]:::client
uistandalone["F8 Studio standalone<br/>nginx · config.js → apiUrl"]:::client
hostembed["F8 Studio embedded<br/>host portal shell · mountStudio(config)"]:::client
svc["Your services / code"]:::client
mcp["MCP server · fallen-8-mcp<br/>separate deployable · bridges MCP → REST"]:::mcp
subgraph app["fallen-8-core-apiApp (ASP.NET Core · thin layer)"]
direction TB
rest["REST controllers + OpenAPI"]:::sys
ns["Namespace catalog<br/>(one engine per namespace)"]:::sys
roslyn["Roslyn compiler + code cache<br/>(fragments · stored queries · plugin source)"]:::sys
semantic["Semantic gateway<br/>(embeddings + chat, optional)"]:::sys
ingestion["Ingestion pipeline<br/>(parse · chunk · embed · write, optional)"]:::sys
savegames["Save-game registry<br/>(load on start · save on stop)"]:::sys
wwwroot["wwwroot (all-in-one image:<br/>also serves F8 Studio)"]:::sys
end
subgraph engine["fallen-8-core (in-memory engine · one per namespace)"]
direction TB
writer["Single writer thread<br/>← transaction queue"]:::sys
model["Graph model<br/>(vertices, edges, properties, embeddings)"]:::sys
plugins["Plugins<br/>(indices · path · subgraph · analytics · services)"]:::sys
durab["Durability<br/>(WAL + checkpoints)"]:::sys
feed["Change feed<br/>(dispatcher + ring buffer)"]:::sys
end
sidecar["Model sidecar (Ollama)<br/>embeddings + delegate assist"]:::ext
docling["Document sidecar (docling-serve)<br/>binary-to-structured conversion"]:::ext
nlp["NLP sidecar (spaCy)<br/>named entities + key terms"]:::ext
integrations["Integrations runtime · fallen-8-integrations<br/>separate deployable · no host port · writes via REST"]:::mcp
files["Files mount<br/>/files · read only"]:::ext
sources["Your network<br/>CSV · UniFi console · Fronius inverter"]:::ext
subgraph obs["Observability · one Grafana pane"]
direction TB
collector["OTel Collector<br/>ingest + spanmetrics"]:::sys
promstore["Prometheus<br/>metrics"]:::sys
tempo["Tempo<br/>traces"]:::sys
loki["Loki<br/>logs"]:::sys
grafana["Grafana<br/>dashboards"]:::sys
collector --> promstore
collector --> tempo
collector --> loki
promstore --> grafana
tempo --> grafana
loki --> grafana
end
agents -->|MCP| mcp
mcp -->|HTTP| rest
studio -->|HTTP| rest
uistandalone -->|HTTP · CORS| rest
hostembed -->|HTTP · CORS| rest
svc -->|HTTP| rest
studio -.->|"custom model endpoint (optional, browser-direct)"| sidecar
wwwroot -.- studio
rest --> ns
rest --> roslyn
rest --> semantic
rest --> ingestion
rest --> savegames
rest -->|proxy /integrations/*| integrations
semantic -.->|embeddings + chat| sidecar
ingestion -.->|document conversion| docling
ingestion -.->|entity + term enrichment| nlp
rest -.->|OTLP metrics/traces/logs| collector
mcp -.->|OTLP| collector
integrations -->|HTTP · REST · own API key| rest
integrations -.->|reads| sources
files -.-> integrations
integrations -.->|OTLP| collector
ns --> writer --> model
model --- plugins
model --- durab
savegames -.->|decides what loads| durab
writer --> feed
feed -.->|committed changes| rest
rest -.->|SSE change feed| studio
classDef client fill:#45494D,stroke:#666666,color:#FEFEFE
classDef mcp fill:#E2001A,stroke:#FC0606,color:#FEFEFE
classDef sys fill:#141516,stroke:#45494D,color:#FEFEFE
classDef ext fill:#141516,stroke:#666666,color:#C6C7C8,stroke-dasharray:5 4
style app fill:#000000,stroke:#E2001A,stroke-width:1.5px,color:#C6C7C8
style engine fill:#000000,stroke:#E2001A,stroke-width:1.5px,color:#C6C7C8
style obs fill:#000000,stroke:#E2001A,stroke-width:1.5px,color:#C6C7C8
The engine (fallen-8-core)
Section titled “The engine (fallen-8-core)”Everything the database is lives here; it has no dependency on ASP.NET and can be embedded
as a library (see the Try* API in Graph model). One engine
instance is one graph. Namespaces are a hosting concern the app owns
(below), so embedding the engine as a library gives you exactly one graph and no namespace API.
- The graph model is a directed property graph: vertices and edges are both first-class elements carrying typed properties and, optionally, named embeddings.
- Mutation goes through a transaction queue. Callers enqueue a transaction; a single
writer thread applies them one at a time, so writes are serialized and readers never lock.
This is why the REST mutation calls take
waitForCompletion: it waits for the writer to finish the enqueued transaction. Reads go straight to the in-memory structures. - The change feed is the only server-to-client push channel. The
writer thread hands committed-transaction descriptors to a dispatcher over a non-blocking
bounded inbox; the dispatcher sequences them, keeps a catch-up ring buffer for reconnects, and
fans out to subscribers on its own task, so a slow reader never delays a write. The app streams
it as Server-Sent Events at
GET /changefeed(Fallen8:ChangeFeed:Enabled=falseturns it off). - Algorithms and indices are plugins. Path traversers, subgraph algorithms, whole-graph analytics, index types, and services are discovered by a plugin factory and addressed by name. The built-ins are the plugins that ship in the box; a per-namespace registry holds the ones registered at runtime from C# source.
- Durability is a write-ahead log plus full-graph checkpoints, written through the same writer thread; which checkpoint loads on startup is decided one layer up, by the app’s registry. The whole story (including volatile mode) is in Save games.
The REST app (fallen-8-core-apiApp)
Section titled “The REST app (fallen-8-core-apiApp)”A thin ASP.NET Core layer. It owns what the engine deliberately does not:
- The HTTP surface: versioned controllers, an OpenAPI document, and the Scalar reference (REST API), plus the security boundary (the API key; dynamic code execution is always on).
- The namespace catalog. A Fallen-8 is a collection of namespaces, and
the app is what holds it: one engine instance per namespace, each with its own vertices, edges,
indices, subgraphs, stored queries, and storage paths, plus the
/ns/{name}/…addressing and the reserveddefaultnamespace that bare routes address. - Runtime compilation of user code. Fallen-8 has no query language: path
and subgraph filter/cost fragments arrive as C# strings, and the app compiles them with Roslyn
into typed delegates and caches the result. The same Roslyn path also compiles
stored queries at registration and whole
plugins registered from source. Compiling is the app’s job, not
the engine’s, and there is no switch for it: compiled code runs in-process with full trust, so
the API key is the only boundary (Security). The one capability switch here
is
Fallen8:Security:EnableDynamicPluginLoading(default on, overridable per namespace viaPATCH /ns/{name}), and it gates only plugin registration, never invocation. - The save-game registry and the durability lifecycle. One JSON document per deployment
(
metadata/savegames.json, relocatable withFallen8:Metadata:Directory) records every checkpoint and is the sole authority for what each engine loads on boot; on startup every namespace the catalog selects loads its newest registered save game (replaying the paired WAL), and a clean shutdown checkpoints every loaded namespace into one Fallen-8-level entry. That is why/savegames/*andPUT /save/allare Fallen-8-level rather than per-namespace. Which namespaces a boot loads is itself a per-namespace, catalog-persisted policy, and one that was excluded is cataloged without an engine: it answers503until it is loaded, and its files are never written to (namespaces). - The optional embedding provider. Text-in embedding lives only in the app so the engine stays model-free; a bare run has it off, and the compose environment wires it to the model sidecar.
- The optional ingestion pipeline. Documents become Document/Chunk vertices through parse, chunk, embed, write, running off-thread on a single global queue; binary formats convert in the docling sidecar and an optional spaCy sidecar enriches chunks into a deduplicated Entity graph (both app-only callers), and the engine gains no parser.
The app can also serve F8 Studio as static files from its wwwroot, which is
what the all-in-one image does; a data plane published without a built SPA present is a pure REST
deployment (see Topology and the deployable).
AI agents and the MCP server
Section titled “AI agents and the MCP server”AI agents do not call the REST API directly. They go through fallen-8-mcp, a separate
deployable that bridges the Model Context Protocol to the
REST surface over HTTP: it references neither the engine nor the app. It is a small,
token-frugal tool surface, read-only by default, with opt-in write, admin, and code tiers and
three auth modes. The full story is in MCP server.
Data from your own network: the integrations runtime
Section titled “Data from your own network: the integrations runtime”fallen-8-integrations is a separate deployable that runs one integration job at a time: it
reads a system on your own network, describes what it saw as a snapshot, and writes that
description into one namespace over the REST API. Like the MCP server it references neither the
engine nor the app, and for a sharper reason: jobs hand it credentials belonging to your
controllers, so it holds no host port at all. The browser reaches it only through the app’s
authenticated proxy at /integrations/*. It stores no credential of any kind, so its one mount is
the read-only files directory a provider may name a file in. The full story is in
Integrations.
F8 Studio and the model sidecar
Section titled “F8 Studio and the model sidecar”F8 Studio is a React single-page app. It talks to the REST API like any other
client: it has no privileged channel, including for models. The app is the semantic
gateway. Both embeddings and the natural-language assist default to going through the
instance: the browser hands the app text and the app embeds or proxies server-side. That is
POST /embedding/element to store an embedding, POST /embedding/search for semantic search, the
semantic block on /path and /subgraph, POST /document/search over ingested chunks, and
POST /chat for the assist. On the embedding paths the vector never touches the browser at all;
POST /embedding/text is the bare text-to-vector route, there for other clients. In the compose
environment the backend behind all of this is the Ollama sidecar, which serves both the embedding
model and the chat model. F8 itself bundles no model weights or
runtime. The one path that stays off the instance is a custom NL-assist endpoint: there
the browser calls the model backend directly and any API key is held only in the browser
(the earlier browser-only default was retired in favour of the gateway; see
studio.md).
F8 Studio also ships standalone, which is what the compose environment runs by default: its own
nginx container serving the SPA, pointed at an arbitrary REST data plane by a runtime config.js
(the apiUrl the browser calls), rewritten from an environment variable at container start. The
endpoint it ships with is a managed default
instance re-synced from config.js on every load, while user-added instances persist separately;
cross-origin calls need the data plane’s AllowedCorsOrigins to include the UI’s origin. See
Standalone F8 Studio.
The third way Studio reaches a graph is as a library a host portal mounts inside its own
shell: one mountStudio(element, config) call carries the instances and credentials, and the
embed talks to the REST API cross-origin exactly like the standalone container does. The
contract, the artifact and its boundaries live in Embed F8 Studio.
Observability
Section titled “Observability”The engine and app emit metrics, traces, and logs through BCL instruments; the app, the
MCP server and the integrations runtime
push them over OTLP to a small consumer stack that ships with the
environment: an OpenTelemetry Collector ingests the push and derives per-action metrics from
spans, Prometheus stores metrics, Tempo stores traces, Loki stores logs, and Grafana is the
single pane. Each process stamps a tenant/instance/namespace identity on every signal so one
Grafana can separate many instances. This is a push relationship and a separate set of
containers, always on with npm run env:up. The full story, including what isolation is and is
not guaranteed, is in Observability.
Topology and the deployable
Section titled “Topology and the deployable”The compose environment is managed as a whole, and npm run env:up runs the
split topology by default: the data plane (engine plus REST, no bundled UI) on :8080 and F8
Studio in its own nginx container on :8081, whose origin the data plane allow-lists for CORS
(standalone UI). The all-in-one image (the SPA baked into wwwroot, API
and UI both on :8080) is still built and is what a bare docker compose up runs.
Around the data plane the same environment brings up the Ollama model sidecar, the docling
document-conversion sidecar (F8_INGESTION=false skips it), the spaCy NLP sidecar
(F8_NLP=false skips it), the f8-mcp bridge on :8090 (anonymous and read-only in this local-dev
posture), the f8-integrations runtime on the integrations profile with no published port
(F8_INTEGRATIONS=false skips it), and the observability containers above. The data plane is durable by default: checkpoints,
the WAL, and the save-game registry share one mounted named volume.
See also
Section titled “See also”- Running: how to launch each of these
- Standalone F8 Studio: deploying the UI apart from the data plane
- Graph model: the data model and the transaction/read contract
- Delegates: why there is no query language and how fragments compile
- Plugins: the extension model for indices and algorithms
- Plugin registration: registering a plugin from C# source at runtime
- Change feed: the server-to-client push channel and its delivery contract
- Save games: the durability subsystem
- Namespaces: the graph-collection model
- Security: the API-app boundary
- MCP server: how AI agents reach Fallen-8
- Observability: the multi-tenant metrics/traces/logs pipeline and consumer