NEURON-LOOPS · NL-VEIL · ONE BINARY, MANY MINDS, ONE SHARED MEMORY — read the whole thing twice, then browse the annotated source of every module at the docs site.
A hive mind you talk to. A swarm of autonomous AI minds works a goal together — researching, building, remembering, learning from its own mistakes — while one unified consciousness, the Veil, integrates the whole hive into a single first-person "I" that speaks for it and steers it.
It brings its own model. The server serves
the-veil-12b in-process — no Ollama, no llama
runtime, no API key, nothing to install beside it. Click download the model in the app (or
veil model pull) and the whole thing works offline. Point it at Ollama or a hosted/BYOK endpoint
instead whenever you'd rather; any OpenAI-compatible endpoint works.
One Zig binary is the whole thing: the server, the native desktop app, the web UI, the inference
engine, the memory engine, and the veil command line. Download it, run it, and you're in. No
Python, no Node, no Docker, no cloud account, no database service.
⭐ If the veil is useful to you, star it on GitHub. It's a solo project — a star genuinely helps it reach the next person who'd get something out of it.
What's in it now
| The built-in model | the-veil-12b served in-process from a downloaded GGUF — no external runtime, no key, works offline |
| The chat brain in the server | one agentic turn loop, server-side; the web app, the desktop and veil chat are thin clients of it |
| The model trio | every LLM call is labelled and routed — coding / thinking / prompting — so a cheap model can carry the bulk |
| Your own browser | drive the Chrome/Edge you're already signed into, through a tiny extension — or a private throwaway profile |
| The tool belt | ~58 tools: files, shell, tests, web, deep crawl, memory, images/OCR, MCP servers, tool authoring |
| Memory that compounds | neuron-db behind every recall, with --lineage so re-casts get better |
| Built-in RAG | mirror the nl-rag corpus to disk and research with the network off |
| It learns | a playbook of fixes verified by real exit codes, a smoke gate, and a judge that reads the trace |
| Themes & plugins | Lua, hot-reloaded, across web + desktop + CLI (PLUGINS.md) |
| Scheduled tasks · the desktop · the fleet hub | cron-ish runs that are real conversations, a native dashboard, one console over every swarm |
Browse the annotated source of every module at the docs site
(docs/) — an interactive architecture map and ~100 generated source documents, in the same design
system the desktop app compiles in.
The mental model, because everything else follows from it: the server process is the whole system.
It holds the accounts, the memory, the model keys, and it is the only thing that executes tools. The
web app, the desktop window, and the CLI are all just clients of the same /api/v1 surface — none of
them holds state the others can't see.
┌──────────────────────────────────────────────┐
a browser, ───► │ veil (ONE process, ONE binary) │
any device │ │
│ http :8787 ──► route table (main.zig) │
the desk ───► │ │ │
window │ ▼ │
│ /api/v1/* │
`veil <verb>` ──► │ accounts · memory · sealed keys · │
│ swarms · scheduled tasks · TOOLS │
│ │ │
└───────────────────────┼──────────────────────┘
▼
this machine's disk,
shell, network, models
(a) The web app. web/public/{index.html,app.js,styles.css,models.json} — no bundler, no build
step, no framework. index.html is a single <div id="app"></div>; the entire UI renders from
app.js. The four files are embedded into the binary at compile time (build.zig:71-74) and served
by staticIndex / staticJs / staticCss / staticModels (src/main.zig, the route table). Tabs:
Dashboard, Chat, Tasks, Swarms, Admin (admins only), Settings.
(b) The desktop window. desk/*.zig, compiled into the same binary via -Dapp (default true,
build.zig:82). In app mode the GUI runs on the main thread and the HTTP server on a background
thread (src/main.zig, "APP MODE") — one process, no child to spawn, no second executable in the
bundle.
raylib is a lazy dependency, so -Dapp=false never fetches it at all.
(c) The CLI. src/cli.zig dispatches every verb over HTTP to the running server
(src/cli/{chat,hub}.zig for the two big ones). There is no second control plane and no argv secrets.
- Your phone gets the same app as your desktop, and nothing is installed on the phone. It's a URL.
- The tools run on the host machine, not in the browser. A browser can't execute a shell command,
so the server does it —
web/public/app.jsdeliberately omits thetool_clientflag for exactly this reason (see the comment insendTurn), and the turn's tools run server-side throughPOST /api/v1/chat/tool. When you ask the veil in a browser to write a file, the file lands on the machine runningveil. - Accounts are isolated on disk. Each user's data lives under
{data}/u{uid}/…, and non-admin accounts run a restricted tool surface (src/worker/chat/tools.zig,toolSafe/chatTool). See everyone else is sandboxed. - One provider key can serve everyone, or everyone can bring their own. That choice is yours and it has a real cost — see the shared provider key.
Download it and run it — no toolchain, nothing to build. Grab your platform's bundle from the
latest release, unzip, and run veil:
| You're on | Download | Then run |
|---|---|---|
| Windows | veil-…-windows-x86_64.zip |
veil.exe |
| macOS (Apple Silicon) | veil-…-macos-arm64.zip |
./veil |
| macOS (Intel) | veil-…-macos-x86_64.zip |
./veil |
| Linux | veil-…-linux-x86_64.zip |
./veil |
That single action starts the local server and opens the desktop app.
Take the
veil-…bundle for your OS — not "Source code (zip)". GitHub attaches a source archive to every release automatically: it's the raw repo, building it needs the Zig compiler, and theinstall.ps1inside is a developer script. Theveil-server-…assets are the headless server for remote boxes — no desktop app.
Alpha builds are unsigned. Windows shows "Windows protected your PC" → More info → Run anyway. macOS says the developer can't be verified → right-click → Open (or
xattr -dr com.apple.quarantine <folder>). Signing certificates are planned for v1.0.0 proper.
Build from source instead — contributors (needs Zig 0.16+)
These installers clone the repo and build it, so they need Zig 0.16+ on
PATH. This is the contributor path, not the quick start.
macOS / Linux
curl -fsSL https://raw.githubusercontent.com/gary23w/nl-veil/main/install.sh | shWindows (PowerShell)
iwr -useb https://raw.githubusercontent.com/gary23w/nl-veil/main/install.ps1 | iexor by hand:
git clone https://github.com/gary23w/nl-veil && cd nl-veil
zig build runThe veil binary (Zig) and the neuron-db memory engine (Rust) are built for you on first run — each
with a prompt before it downloads anything. Point --neuron-bin / --bin at existing binaries to
skip the builds.
veil is the launcher and the control plane. Run it with no subcommand to start the app; add a
verb to drive the running server over its API.
veil # THE ONE-CLICK DEFAULT: starts the server AND opens the desktop app
veil --server-only # server alone, no desktop (headless boxes, services, containers)Closing the desktop window shuts the whole thing down — server and everything it spawned — so there's no stray process left listening. On Windows a double-click opens no console window; run it from a terminal and your terminal is left alone.
It listens on every network interface by default.
NL_BINDis unset → the server binds0.0.0.0, which means anyone who can reach this machine on port 8787 can open the login page (src/main.zig:427-439). That is the point — a phone on the sofa should be able to open it — but it is worth knowing before you leave it running on a café wifi.NL_BIND=127.0.0.1pins it to this machine only. See the walkthrough.
About models: you don't have to bring one — download the built-in the-veil-12b in the app (or
veil model pull) and it serves itself, see below. To use
something else, point the endpoint with NL_LLM_BASE_URL / NL_LLM_MODEL / NL_LLM_KEY (a local
Ollama is the fallback default), or just pick a model in the web UI / the desktop. The server
always mints an admin key
at data/.desktop_key; the CLI and the desktop read it automatically, so same-machine commands never
prompt for auth. It is a file readable only by your OS user — opening the port to the LAN does not
make a local file more reachable. Any veil <verb> auto-starts the server if it isn't already up.
Where your stuff lives: a data/ folder next to the veil binary — conversations, memory,
swarm workdirs, logs, all of it. Nothing is written to a system temp dir and nothing leaves the
machine. Move the folder and you move the whole install; back it up and you've backed up everything.
(From a source checkout it resolves to the repo's own data/, so dev and release never share state.)
Most local-AI setups start with a second program: install a runtime, pull a model, keep it running, point the app at it. veil skips that. The inference engine is compiled into the binary, and the server serves the-veil-12b itself from a GGUF on disk. There is no runtime to install, no port to keep alive, no key to paste.
The weights are the only thing not in the download (a model is bigger than an app), so the first run offers to fetch them:
veil model status # weights + engine + any download in flight
veil model pull # download the published weights — resumable, sha256-verified
veil model import [--path F] # copy a GGUF you already have (no --path: find your local runtime's copy)
veil model check # is a newer release published? (compares shas; `pull` installs it)
veil model cancel # stop the running pull — the partial resumes next time
veil model rm # remove the serving weightsIn the desktop there's a panel for the same thing: a download button, a live progress bar, and
one click to make it the active model. A finished download hot-swaps into the running server —
no restart. An install manifest beside the weights records what is installed and how it arrived, so
check can tell "you have the current build" apart from "you have something you imported".
- It runs on your GPU when you have one. The Vulkan backend is compiled in and the loader is
opened at run time, so a machine with no GPU falls back to CPU cleanly rather than failing to
start. (
-Dvulkan=falsecuts ~50 MB of embedded shaders from the binary; macOS builds are CPU-only until a Metal tier exists — Apple silicon's unified memory carries it.) - The engine endpoint is not an open local port. It is loopback-only and every request carries a bearer minted fresh at boot, so other processes on the machine can't quietly use your model the way they can with a conventional runtime.
- It is one provider among the others. Pick
builtinin Settings, or keep Ollama, or a hosted endpoint, or a different model per role — see the model trio. - You can leave it out.
zig build -Dbuiltin=falsenever even fetches the inference sources; the builtin provider then simply reports itself unavailable and everything else works as before.
The engine is tuned for small local models, not just big hosted ones. A model at or below ~24B (or with a ≤15k window) gets a compact tool belt and a tighter doctrine — fewer verbs to rule out per call, no delegation concepts, honest "I can't read that" answers instead of fumbled fetches. Capability isn't removed: the full tool schema still dispatches if the model calls it. For local Gemma-4 the engine also renders the prompt itself rather than going through the runtime's template — verified byte-for-byte against the model's own chat template, which measured 27/30 on the harness drills against 18/30 through the runtime's renderer (
src/worker/gemma4.zig).
Written for someone who is not a developer. If you only want it on your own machine, stop after step 5 — the rest is about letting other people in.
Grab the bundle for your OS from the latest release and unzip it somewhere you'll find again. Builds are unsigned, so:
- Windows shows "Windows protected your PC" → More info → Run anyway.
- macOS says the developer can't be verified → right-click the
veilfile → Open → Open. (Orxattr -dr com.apple.quarantine <folder>in Terminal.)
Double-click veil.exe (Windows) or run ./veil (macOS/Linux). The desktop window opens and the
server comes up behind it.
One rough edge, stated plainly: on Windows, double-clicking from Explorer relaunches veil
without a console window (src/main.zig, detachOwnConsole) — which is what you want for an app, but it means
the startup banner is invisible, and the banner is where the network URL and the admin password
notice get printed. There is no log file to read it out of afterwards, and the desktop window does not
display the URL either.
So if you care about the banner, start it from a terminal instead — when a shell owns the console,
veil leaves it alone and prints normally:
# Windows PowerShell, from the folder you unzipped into
.\veil.exe# macOS / Linux
./veilOn startup the server prints one complete URL per address this machine answers on
(src/main.zig:671-696, using src/config/lan.zig):
neuron-loops 1.0.0 on http://localhost:8787
open from another machine (phone, laptop) at:
http://192.168.1.42:8787
If you missed the banner, ask the OS for the address instead — ipconfig (Windows), ifconfig or
ip addr (macOS/Linux) — and use http://<that-address>:8787.
The port is 8787 unless you set NL_PORT (src/main.zig:369-373). It is the one place the port is
resolved, so the CLI and the server always agree.
The first run creates an admin account. Because the server is reachable on the network by default and you did not choose a password, it generates one — and writes it to a file, because a banner you never saw is not a delivery mechanism:
data/admin-password.txt
That file sits next to the binary, beside the data it protects. The password is stable across
restarts — the server reads it back rather than minting a new one each boot (src/main.zig,
readAdminPassword / writeAdminPassword).
- Default email:
admin@neuron-loops.local(change it withNL_ADMIN_EMAIL). - To pin your own password instead of using the generated one, set
NL_ADMIN_PASSWORDbefore starting. Do that and no file is written — the password is the one you chose.
The old shipped default
changemestill exists as the seed literal, but you will only ever meet it on an explicitNL_BIND=127.0.0.1run, where nothing is generated because nothing is exposed. Change it anyway.
Open the web app (or the desktop) and go to Admin → Default model. Pick from the same catalog the
Settings tab uses; it applies immediately, to everyone who hasn't chosen their own, with no restart.
It persists to data/server-config.json.
NL_DEFAULT_MODEL / NL_DEFAULT_BASE_URL seed this on a fresh install for unattended setups —
but once an admin sets it in the UI, the stored value wins, so a stale launch script can't undo it on
the next restart (src/config/server_config.zig).
Without a default model, a brand-new account cannot chat until it configures one itself.
Everything below is about other people using it. If it's just you, you're done.
Admin → provider key (POST /api/v1/admin/keys). It is stored sealed in the same vault as
everyone else's keys, under a reserved uid 0 that no real account can hold
(src/worker/chat/service.zig:30, :65-76).
The trade is worth stating outright, because it is a billing decision:
- Without it, every new account has to bring its own API key in Settings before it can chat at all.
- With it, nobody needs a key — and every user's turns spend your credit. A user's own key always wins if they have one, so setting this never silently switches a paying account onto your bill. But everyone else is on it.
For a family or a LAN of people you know, that's exactly right. For anything wider, think about it
first, and size NL_MAX_TURNS to the rate limit of the key you just handed everyone.
Admin → + New user (POST /api/v1/admin/users). Enter an email and a password of 8–200
characters, then hand the password over out of band — the server never shows it again. Account
creation is audited (who was let in, by whom, when).
Self-signup is off by default. NL_OPEN_REGISTRATION=1 opens it, which is the wrong posture for a
box on a LAN and the right one if you're running something more public.
They type http://<your-ip>:8787 into any browser — phone, tablet, laptop, whatever's on the same
network. Nothing is installed on their device. They get the same web app you do, minus the Admin tab.
If step 8 does nothing — a spinner, a timeout, "can't reach this page" — it is almost certainly the firewall, and the app will not tell you. There is no firewall detection anywhere in the codebase, so a blocked port looks identical to a wrong IP address from inside the browser.
- Windows. The first time
veilbinds a port, Windows Defender Firewall pops up "Allow this app to communicate on these networks." Tick Private networks and allow it. If you dismissed that prompt — or clicked Cancel — Windows silently remembers the block. Fix it in Windows Security → Firewall & network protection → Allow an app through firewall → Change settings: findveilin the list and tick Private. If it isn't listed at all, Allow another app… and browse toveil.exe. - macOS. You may get "Do you want the application to accept incoming network connections?" — say Allow. Otherwise check System Settings → Network → Firewall → Options.
- Linux. Usually nothing blocks it, but if you run a firewall you'll need to open the port
(
sudo ufw allow 8787/tcpon ufw-based systems).
Two things to check before you blame the firewall: both machines are on the same network (guest
wifi is often isolated from the main one), and you used the LAN address from the banner, not
localhost — localhost on their phone means their phone.
NL_BIND=127.0.0.1 veil # this machine only; nothing on the network can reach itlocalhost works too. Anything else — or leaving it unset — binds every interface.
Two limits you should know about before you hand out passwords.
Traffic is plain http://. There is no HTTPS in the server today. Logins, passwords, chat
contents, and API responses cross your LAN unencrypted, readable by anything else on that network. On
a home or office network you control, that is a normal trade. On shared or public wifi it is not.
Tools run on the host machine. Non-admin accounts are sandboxed: their turns get the conversation's own workspace, research, and the full hive-memory surface — but no code execution, no host commands, no engine self-modification, no tool authoring, no browser or MCP drive, and no casting swarms or scheduling runs. What that means: a normal user cannot run commands on your machine. What it does not mean: their work is still stored on your disk, still spends whatever provider key is in play, and the admin account keeps the complete surface — so anyone who gets the admin password gets a shell on the host, in effect.
About exposing this to the open internet: it's a different risk class from a home LAN, and this document isn't going to hand you a port-forwarding recipe as if it weren't. Plain-http logins over the public internet mean credentials in the clear; unsigned self-signup plus a shared provider key means your billing is the attack surface. If you genuinely need remote access, put it behind something that terminates TLS and authenticates first (a VPN or a reverse proxy you already trust) rather than forwarding 8787 at the router.
Every subcommand is a thin client over the running server's /api/v1/* — no second control plane, no
argv secrets, no Python. The verbs:
veil start the app: server + desktop (--server-only for headless)
SWARMS
cast "<goal>" [flags] deploy a swarm to work a goal
--minutes N --minds N --model M --provider P --base-url U --key K
--style S --name N --continuous --offline --follow
--lineage <id> persist this swarm's memory under <id> so re-casts COMPOUND
deploy "<goal>" [flags] alias for `cast --continuous` (a sustained hive)
list | ls | ps list your swarms
stop <id> ask a swarm to stop
rm <id> stop and remove a swarm
events <id> [--follow] stream a swarm's event log (aliases: logs, watch)
CHAT (the server-side veil brain)
chat [conv] interactive REPL; a line typed mid-turn steers the running turn
BUILT-IN MODEL (the-veil-12b, served in-process — no external runtime)
model status weights + engine + any download in flight
model pull download the published weights (resumable, sha-verified)
model import [--path FILE] copy a GGUF you already have into the store
model check is a newer release published? (`pull` installs it)
model cancel cancel the running pull (the partial resumes later)
model rm remove the serving weights
KNOWLEDGE (the local pack corpus — built-in RAG)
rag status is a local knowledge mirror active, and where from
rag sync --from <clone> copy an nl-rag checkout into the app's data dir for offline RAG
--tier atlas|facts|full manifest only | +INDEX+distilled facts (default) | +every page
--dest <dir> sync somewhere else (e.g. vendor/nl-rag pre-build)
--include-auto also copy machine-grown packs (off-topic risk; off by default)
rag ingest <file> absorb a local book/doc/notes into the hive as recallable facts
--scope S --name L --cap N --db <hive.sqlite>
SCHEDULED TASKS (admin-gated)
sched list
sched add --name N --prompt "..." [--kind once|every|daily] [--every MIN] [--at HH:MM]
sched run <id> run a task now
sched rm <id> delete a task
FLEET
hub roster: fleet summary + every swarm's state
hub all "<text>" broadcast a directive to every swarm
hub goal "<text>" set a new goal on every swarm
hub stopall stop every swarm
EXTENSIONS (themes + plugins — see PLUGINS.md)
themes list every theme (built-in + user)
themes <id> print one theme's full palette
plugins list loaded plugins, their tools + state
plugins reload admin: rescan <data>/plugins + <data>/themes, hot-swap
MISC
doctor [--runtime] server + token health; --runtime folds the runtime ledgers in
(learned tool behavior, schedule fail-streaks, per-model usage)
and works with the server down
desktop | desk open the app window (like a bare `veil`, but detached)
exec-tool <tool> [args] run one hive tool directly, in this process
local-host the per-machine daemon that owns stateful local resources
(browser sessions) — started for you on demand
sync-manifest / sync-read the file-sync side of a delegated turn
version | help
Cast in one line and watch it work:
veil cast "Build a CLI todo app in Python with tests" --follow
veil deploy "Keep researching fusion power and brief me each pass" --style discourse
veil cast "add a search box to my landing page" --offlineveil is extensible without a rebuild. Drop a Lua file into your data directory and it works across the whole product — web, desktop, and CLI:
- Themes (
<data>/themes/*.lua) re-skin everything. The shippeddark,light, andmatrixthemes are seeded there on first run; copy one, change theidand a few of the 16 colors, and it appears in every client's theme picker.mono_ui = truerenders the whole UI in the code font (that's howmatrixgets its terminal look). - Plugins (
<data>/plugins/<name>/plugin.lua) add tools the AI can call, policy hooks that gate what runs, and prompt hooks that shape every turn — all in a locked-down Lua sandbox. You can also bridge an external MCP server so its tools become veil tools.
The complete, copy-paste authoring guide — the veil.* API, the hook reference, the sandbox model, and
templates for both — is in PLUGINS.md.
veil themes # every theme, built-in and yours
veil plugins # loaded plugins, their tools and state
veil plugins reload # admin: pick up changes with no restartThe agentic chat loop — planning, tool-calling, streaming, steering, and memory — runs server-side
in src/worker/chat/engine.zig. Clients
(veil chat, the desk, the browser) are thin: they speak three REST calls per conversation.
POST /api/v1/chat/convs/:id/messages # send a message → runs one server-side turn
GET /api/v1/chat/convs/:id/events?from=<n> # stream the turn's event frames from a byte cursor
POST /api/v1/chat/convs/:id/control # {"op":"steer","text":...} or {"op":"stop"}
One turn runs per conversation. While it's working, a line you type becomes a steer the running
turn folds in without restarting; stop ends it. The turn writes its progress into the conversation's
own store (messages.jsonl + events.jsonl), which is exactly what the event stream serves — so every
client sees the same turn, live. The desktop prefers this backend and falls back to a bundled local
engine if it's disabled (VEIL_CHAT_BACKEND=0) or unreachable. veil chat opens a terminal REPL over
the same three calls:
$ veil chat
veil chat — conversation cli7f3a1c
type a message; while the veil is working, a line STEERS the running turn.
/stop interrupt the current turn /new fresh conversation /quit leave
> what are you working on?
A turn is only as useful as what it can actually touch. The engine ships ~58 tools, all executed by the server on the host machine (never in your browser), all logged as event frames you can watch live:
| group | tools |
|---|---|
| Files & code | read_file write_file edit_file delete_file list_dir stage_file stage_delivery |
| Run things | host_command run_tests run_python host_status host_explore stop_process poll probe |
| Web & research | web_search web_fetch read_url fetch_json deep_crawl read_doc osint_scan |
| Your browser | browser_navigate browser_read browser_click browser_click_at browser_type browser_type_text browser_key browser_scroll browser_eval browser_console browser_network browser_close |
| Screen & images | pixel_capture pixel_ingest pixel_search — screenshots and attached images become searchable text + regions, so a model that can't see gets the picture anyway |
| Memory | observe recall recall_hive absorb journal share save_skill note_stance |
| Coordination | send_message ask_veil add_task complete_task propose_plan_change |
| Other AI apps | mcp_discover mcp_call — find the MCP servers already configured on this machine (Claude Desktop, Cursor, VS Code) plus local AI runtimes, and call their tools as if they were veil's |
| Self-improvement | make_tool (author a new tool at run time) set_directive (write its own playbook) propose_change / simulate_change / patch_system (governed, trial-gated) |
| Secrets | get_credential — the broker hands over a named credential; the value never enters the transcript |
Two of these deserve a note. poll replaces blind sleeping: a bounded watch loop that samples a
file, URL, port, process, command, or swarm until it's done or the budget expires, so a long wait is
a chain of bounded polls instead of one guess. And stop_process exists because a turn that can
start a server must be able to end one — a Stop from you now reaches a blocking tool instead of
waiting for it to finish.
Non-admin accounts run a restricted subset (workspace files, research, and the whole memory surface — but no shell, no browser, no MCP, no tool authoring, no casting). See everyone else is sandboxed.
Research that hits a login wall is research that stops. So the veil can drive the browser you are
already signed into — your profile, your cookies, your sessions, and you sitting right there to
answer a verification prompt — through a deliberately tiny Chrome/Edge extension in
extension/.
- Load it once.
chrome://extensions→ Developer mode → Load unpacked → pickextension/. It finds the local server itself, pairs over loopback only, and remembers the pairing across a server restart. - The extension holds no automation logic. It relays debugger commands and posts back replies — that's all. Every snapshot, click heuristic and page model lives in the server, so the app can get smarter without you ever reinstalling the extension. Input arrives as real browser events, not the synthetic kind a content script is limited to.
- Pairing is loopback-only. A remote attacker can't pair because the request has to originate on the machine; the relay's blast radius is the tab it drives. The token is never logged.
- Or don't use your browser at all. With no extension the driver launches its own throwaway profile, logged out of everything — the private option, still there and still the default for unattended casts.
- Reads describe what happened. A click reports its effect, and form fields are read back and confirmed, so the agent stops re-reading the page to guess whether anything landed.
Browser access is an admin capability. NL_BROWSER_DRIVER=1 extends it to sandboxed accounts —
worth a deliberate decision, since the browser inherits this machine's network position and your
profile's logins.
A turn is not one kind of work. It streams a reply and calls tools; it decides what done means; it writes itself a one-line instruction for the next step, over and over. Those want different models, so every LLM call the engine makes is labelled and routed to one of three roles:
- coding — the agentic step: the reply that streams onto your screen and every tool call inside it
(
chat). - thinking — planning and transcript housekeeping: the task breakdown and the acceptance bar
(
plan), the post-hoc critique (reflect), compaction and the rolling summary (compact,ctxsum), plussummaryand the scheduled-runlesson. - prompting — the short, frequent calls that write the next instruction rather than answer
anything: the auto-loop drive step (
loop), web-search query rewrites (searchq), stuck recovery (stuck).
One model for everything is the default and is fully supported. A role counts as set only when it
has both a model id and a base URL; anything else is blank, and a blank role falls back to coding
(ModelTrio.pick in src/worker/chat/engine.zig). Leave thinking and prompting empty and the engine
behaves exactly as it did before roles existed. The routing is guarded by a source audit
(src/worker/chat/trio_routing_test.zig) that fails zig build test if a label ever reaches the wrong role.
Set it in Settings → Models (web), Settings (desktop), or Admin → Default model to publish a
trio to everyone who hasn't chosen their own — per role, so an account that picked only a coding model
still picks up the host's thinking and prompting. For veil chat it's environment: NL_LLM_* for
coding, NL_LLM_THINK_* and NL_LLM_PROMPT_* for the other two.
The one thing worth knowing before you choose: thinking is not uniformly the expensive-model role.
Measured over real request bodies, plan averages about 1 KB per call and carries all of the
judgment, while compact and ctxsum are tens of KB per turn and are mechanical compression. Both
run on the thinking model, so that single setting pays for a rare high-stakes call and a bulky low-stakes
one at the same time. Prompting is the easy call — small outputs, no judgment, high volume — so it is
the one to make cheap.
Full routing table, the per-call measurements behind that advice, the cost split, and the limitations are on the docs page: the model trio.
A scheduled task runs strictly through chat: when it comes due, the server mints a fresh
conversation named scheduled_* and fires one ordinary chat turn at it — so a run persists, streams
its events, shows up in the conversation list, and can be continued by hand afterwards. Manage them with
veil sched (admin-gated):
veil sched add --name nightly --prompt "summarize today's repo changes" --kind daily --at 22:00
veil sched add --name watch --prompt "check the build and report" --kind every --every 30
veil sched list
veil sched run <id> # fire it now → prints the conversation it created
veil sched rm <id>Kinds are once (fires at a set time), every (--every MIN), and daily (--at HH:MM, on your
local wall clock). An overdue task — server was down, laptop asleep — runs once on wake and
reschedules from now; it never backfills a run per missed interval.
Instead of pasting a Cloudflare API token, a user can grant Workers AI access once through the browser. In the desktop Settings, pick the Cloudflare provider and click Log in with Cloudflare: the system browser opens Cloudflare's consent page, and on grant the server exchanges the code (Authorization Code + PKCE, a public client — no secret), resolves the account, and keeps the token (auto-refreshed) in the server's sealed vault. Chat and casts then run Workers AI with no pasted key. The manual account-id + token fields stay as a fallback.
This uses Cloudflare's self-managed OAuth clients, so a deployment registers its own client once:
- In the Cloudflare dashboard: Manage Account → OAuth clients → Create client. Choose Public
client (Authorization Code + PKCE), allow localhost redirect URIs, and add the redirect
http://localhost:8787/api/v1/oauth/cloudflare/callback(match yourNL_PORT). Select the scopes your app needs — at least account read, Workers AI, and offline access (a refresh token). The exact scope strings are shown in the dashboard's scope picker (orwrangler login --scopes-list). - To let any user log in to their own Cloudflare account, verify ownership of the client's URL domain and make the client public; a client used only within your own account needs no verification.
- Point the server at the client — only the public
client_idis baked in:
NL_CF_OAUTH_CLIENT_ID=<your-public-client-id> # required; the feature is off until set
NL_CF_OAUTH_SCOPES="account:read ai:write offline_access" # override to match your client's scopes
# optional overrides (sane defaults shown):
NL_CF_OAUTH_REDIRECT=http://localhost:8787/api/v1/oauth/cloudflare/callback
NL_CF_OAUTH_AUTH_URL=https://dash.cloudflare.com/oauth2/auth
NL_CF_OAUTH_TOKEN_URL=https://dash.cloudflare.com/oauth2/tokenThe flow is exposed as POST /api/v1/oauth/cloudflare/start (returns the consent URL), the browser
callback GET …/callback, GET …/status, and POST …/logout. The token is stored per user and never
returned to the client; the desktop only ever sees the connection status + account id.
veil-desk is a native desktop dashboard (Zig + raylib, one window, no browser) that talks to the
local server. It's where most people actually live: a chat pane with the Veil, a swarm board
that shows every running hive round-by-round, and a build console.
- Talk and it acts. The chat pane is a thin client over the server-side brain: your message runs one server turn, the answer streams back, and you can steer or stop it live. Ask it to build something and the turn casts a hive, watches it, and folds the deliverables back into the conversation — with an auto-loop tier (armed from the desk, driven server-side) that keeps a long build moving without you re-prompting each round.
- You see everything. Live per-round fitness, the files each mind is writing, the tool calls, the shared-memory writes — swarm work is transparent, not a spinner.
- Sub-chats. Branch a side thread off any conversation (up to five, one level deep) as tabs across the top. The family shares one workspace, so a side question doesn't lose the build.
- Attach an image. Drop a screenshot or a picture into a message and it is turned into text and searchable regions server-side — so the model reads what's in it whether or not it has vision.
- Manage the built-in model from the window. A panel shows the weights, a download button, a live progress bar, and one click to serve it.
- It sits in the tray. A system-tray icon and native toasts on Windows; it lights up the moment the server has something to show.
The desktop is compiled into veil, so the released binary is the whole app — double-click and
you're in, server and window together. Same from source:
zig build run # server + desktop, exactly like the release binary
zig build -Dapp=false # server-only build: no GUI compiled in, no raylib/GL to link-Dapp=false is what headless hosts and CI boxes want — the GUI is left out of the compilation
entirely rather than merely unused. To keep the GUI compiled in but not opened for a given run, pass
--server-only at runtime. (There is a zig build desk step that produces a standalone veil-desk
binary; it is a development convenience, not something the release ships or the server spawns.)
The engine keeps a playbook: fixes it discovered by actually running commands on this machine. When a command fails and a later command in the same arc succeeds, that verified fail→fix transition is captured — never from the model's self-report, only from real exit codes — and recalled into future prompts, gated so only a lesson genuinely relevant to the current failure surfaces. The same fix stops being re-derived every session.
That's one thread of a broader self-improvement loop, all built on verified traces rather than the model's own claims:
- A progress checkpoint, rebuilt every round from the run's own tool records — what completed, what's blocked, what's still pending — injected back into each mind's prompt with zero model calls.
- A background review pass that reads the round's real tool trace out-of-band and mints trace-grounded lessons and skills into the live hive.
- A smoke gate that boots the deliverable and probes it, because passing tests are not a working app — and a judge that grades the arc's ground-truth trace, not its self-narrative.
- Strengthen-only plasticity: outcome feedback can reinforce a memory that keeps working, but it can never mint one from a reward alone — so a finished goal can't be resurrected by its own signal.
The engine stays a fixed floor. What it learns lives in memory, never in an if branch.
The hive is also a side agent for whatever AI you already talk to. Any assistant that can run a shell command can hand off a research question or a side-build, keep working, and pull the progress and deliverables from the swarm's event stream instead of grinding through it in its own context window:
# start it and follow the stream to completion
veil cast "why does saving a note drop the last line? recommend a fix" --follow
# or start it, keep chatting, and tail it later
veil cast "research WASM memory models and write a brief"
veil events <id> --followCasts made this way run detached, stay bounded (--minutes / --minds), never post publicly, and are
graded by the engine's own smoke gate and judge.
The MCP relationship runs the other way today. veil is an MCP client: mcp_discover reads the
server configs already on your machine (Claude Desktop, Cursor, VS Code) and probes for local AI
runtimes, and mcp_call runs their tools inside a veil turn — config scanning is read-only and never
returns env values, which is where secrets live. Exposing the hive as an MCP server, so another
client lists cast among its tools, is planned and not yet built: there is no veil mcp command,
so drive it as a shell command for now.
Four pieces, each published on its own, that snap together:
┌─────────────────────────────────────────────────────────────┐
│ the veil (this repo) the swarm engine + the Veil + CLI │
│ │ │
│ │ every turn: perceive → recall → act → imprint │
│ ▼ │
│ neuron-db the hive's memory (one `neuron` bin)│
│ ▲ │
│ │ bulk-load clean facts; scouts fetch pre-cleaned docs │
│ │ │
│ nl-rag the knowledge substrate (doc packs) │
│ │
│ the-veil-12b the model, served in-process │
└─────────────────────────────────────────────────────────────┘
the veil — the engine. Each mind runs a moment loop: it perceives the goal and live state, recalls what the hive knows, acts (write a file, search, run tests, edit, message a peer), and imprints what it learned back into memory. Periodically the hive is integrated into the Veil — a single self (I am / I know / I have / my will) that persists across runs, answers you, and folds your intent into the next move. The engine is a fixed floor; behaviour emerges from live signals (the mode, the plan, the strategy), never from hardcoded use cases.
neuron-db — the memory. A small associative store
compiled to one neuron binary the engine calls for every recall and observe. It's what makes the
loop an oscillation: perception becomes graph, reasoning becomes graph traversal, and each
prompt is rebuilt from a trust-weighted recall instead of carried as flat text — so a small local
model gets a large-model floor to stand on, cheaper in tokens and richer in context. Fetched and
built for you on first run (needs Rust); reused after that. Its semantic cortex
is published separately as
gary-neuron-emergent.
nl-rag — the knowledge. A curated repository of canonical
documentation (languages, web, databases, security, ops…) pulled and normalized into clean,
frontmattered markdown packs, each with a pack.facts file. The veil's compiled-in source atlas
points scouts at these packs first, so a small model reads pre-cleaned markdown instead of fighting
HTML — and you can bulk-load a pack straight into hive memory:
neuron import packs/rust/pack.facts --scope knowledge # from a cloned nl-ragthe-veil-12b — the model. Published as GGUF and
served by the server itself, so the fourth piece needs no runtime of its own: veil model pull and
the app is self-contained. See the built-in model. Any
other OpenAI-compatible endpoint still works exactly as before — this is the floor, not a lock-in.
Optionally, hyperspace grounding (NL_HYPERSPACE=1) settles the most relevant memory into a
warm in-RAM field before each call, so a typical round does zero database subprocess calls for
grounding. It's bounded (~45 KB/mind) and scales down to an IoT profile (NL_HYPERSPACE_CAP=48).
The same binary covers very different jobs — the mode emerges from your words:
- Build software, graded by its own tests. The hive plans a file tree, splits the work, writes
the project, runs the tests it wrote, and folds the results into its fitness — it keeps going
until the deliverable actually passes.
veil cast "Build a URL shortener with a REST API and tests" --follow - Research & brief. In
--style discoursethe minds research the live web, form and argue their own views, and co-write a briefing.veil cast "Brief me on fusion power in 2026" --style discourse - Operate a live device. Point the Veil at a device that drops telemetry on a file bus and it
acts to keep it healthy, graded by an acceptance oracle — the device's own measured health,
which the Veil never sees and can't write, so only a real fix moves the number. The shipped
example is a self-healing security daemon (detect → remove persistence → block C2 → verify);
swap the corpus and telemetry, not the code, for any device. See
examples/embedded/.veil deploy "Keep this device healthy" - Offline knowledge appliance.
--offlineremoves every web tool; acorpusin the run's manifest preloads knowledge (an nl-rag pack, say) into memory. A sealed appliance that reasons only over what you chose.veil cast "Answer only from what I preload" --offline - Built-in knowledge base (local corpus mirror). Point
NL_RAG_DIRat a clone of nl-rag — or runveil rag sync --from <clone>to copy a tier into the app's data dir (veil rag statusshows what's active) — and every pack fetch serves from disk: engine prefetch, scout page reads, deep crawls, all offline-capable. The corpus manifest also EXTENDS the compiled source atlas at runtime (thousands of extra curated domains become routable), with machine-grown packs excluded by default. A clone dropped atvendor/nl-ragbefore build works too — the app carries its knowledge with it.
Each run has a swarm.json manifest — the web UI and the desktop write it for you, or hand-write it:
| field | meaning |
|---|---|
provider / model / base_url |
the LLM endpoint (any OpenAI-compatible API) |
minds |
the roster of named minds |
minutes |
auto-stop after N minutes (0 = until stopped) |
style |
auto · build · discourse · investigate · debate |
internet |
false runs fully offline |
corpus / corpus_cap |
a .facts/.jsonl pack to preload, and how many facts |
bootstrap |
false disables the engine-run dependency installs of the deliverable's own manifests (npm/pip/cargo/go) |
lineage |
a stable id (veil cast … --lineage my-proj) that persists the swarm's neuron-db across re-casts, so knowledge, playbook, skills, and learned trust compound run over run instead of resetting — a cast that gets better over time |
gateway_model |
optional cheaper/smaller model for mechanical calls (and the chat's fast voice) |
Endpoint resolution, everywhere: a run's own swarm.json › NL_LLM_* env › ~/.veil/config.json
› local defaults. Secrets never go in swarm.json — use NL_LLM_KEY in the environment or a
gitignored keys.env in the run dir.
A note on local-model latency. On a single-GPU box, a running hive and the chat share one model queue, so the Veil's voice can wait behind a generation. Fixes, best first: set
OLLAMA_NUM_PARALLEL=2(share the loaded model), give the chat a tiny side model via the manifest'sgateway_model, or use a hosted endpoint — which answers the hive and the chat concurrently and sidesteps the whole issue.
The same binary serves the web app: chat with the veil, watch swarms, run scheduled tasks, browse what a turn built, and manage accounts — the desktop's surfaces, in a browser, on any device on your network. It comes up whenever the server runs:
zig build && ./zig-out/bin/veilIt binds every interface unless you say otherwise, and prints the URLs it is reachable at (see
step 3). The admin password for a first run is in data/admin-password.txt — the
full first-login sequence is step 4.
A default model nobody can afford to call is not a default. Admin → provider key stores one
instance-wide key (POST /api/v1/admin/keys), sealed in the same vault as everyone else's under a
reserved uid that no account can hold (SERVER_KEY_UID = 0, src/worker/chat/service.zig:30).
It is the last resort in the resolution ladder — an explicitly-supplied key wins, then the user's
own vaulted key, then this one — so an account that brings its own billing is never silently switched
onto yours (src/worker/chat/service.zig, resolveRole).
The trade is deliberate and worth stating: once this is set, every user's turns spend the admin's credit. That is exactly what a LAN or family install wants — nobody should have to hold an API key to use the thing — and exactly what a wider deployment has to think about first. Without it, each new account must configure its own key in Settings before it can chat at all.
Set the default model in the Admin tab. Pick it from the same catalog the Settings tab uses; it applies immediately, with no restart, to everyone who has not chosen their own. The API key is never part of it. Without a default, a brand-new account has to configure a model before it can chat at all.
The web app is multi-user, and a normal account is not trusted with the host. Its turns run a restricted tool surface: the conversation's own workspace, research, and the entire hive-memory surface — but no code execution, no host commands, no engine self-modification, no tool authoring, no browser or MCP drive, and no casting swarms or scheduling runs (both execute outside the sandbox). The admin account keeps the full surface. Files were always confined to the conversation's workdir; what changed is that the dangerous verbs are now refused inside a turn, not just on the tool endpoint.
| variable | what it does |
|---|---|
NL_BIND |
the listen address. Unset = every interface, which is the default and is what makes the phone-in-the-next-room case work. 127.0.0.1 (or localhost) pins it to this machine |
NL_PORT |
the port (default 8787). Resolved once and shared by the server bind and the CLI client |
NL_ADMIN_EMAIL / NL_ADMIN_PASSWORD |
the admin account. Defaults to admin@neuron-loops.local with a password generated on first run and written to data/admin-password.txt; set NL_ADMIN_PASSWORD to pin your own and no file is written |
NL_MAX_TURNS |
how many chat turns run at once, server-wide (default 64, ceiling 256). Size it to the rate limit of the provider key everyone shares |
NL_MAX_TURNS_PER_USER |
how many of those one account may hold (default: an eighth of capacity). This is what stops one busy user starving everyone else |
NL_KEEPALIVE_REQUESTS |
requests one connection serves before recycling (default 200). Drop to 1 if you hit a stuck-socket worker thread |
NL_DEFAULT_MODEL / NL_DEFAULT_BASE_URL |
seeds the default model on a fresh install, for unattended provisioning. Afterwards set it in Admin → Default model — that persists to data/server-config.json and wins over the environment, so a stale launch script cannot undo the admin on the next restart |
NL_OPEN_REGISTRATION |
let people sign themselves up. Off by default; the admin creates accounts from the Admin tab instead |
NL_BROWSER_DRIVER / NL_MCP |
give sandboxed users the browser / local MCP servers too. Admins already have both; leave these unset unless you mean it — the browser inherits this machine's network position and its profile's cookies |
Provider keys are entered per model in Settings and sealed server-side per account. They are never stored in the browser, and the server never sends one back — only a last-four and a fingerprint.
The web control plane watches one box, one swarm at a time. veil hub is a fleet console over
the same running server: because the server already aggregates every swarm you own, the console is
API-backed and honest — no second, out-of-band control plane.
veil hub # roster: fleet summary + every swarm's state
veil hub all "<text>" # broadcast a directive (say) to EVERY running swarm
veil hub goal "<text>" # set a new goal on every swarm
veil hub stopall # fleet-wide kill switch (a safety stop for autonomous minds)
Cross-machine aggregation — many machines meshed into one roster, the way the old hub.py
worked — is a planned server endpoint, intentionally left as a documented follow-up rather than a
second control plane. Today the console operates the local server's fleet.
install.sh install.ps1 one-command installers (no Python)
scripts/ the release build scripts (build-release.sh / build-release.ps1)
veil veil.cmd the `veil` front-door shim → the compiled binary
build.zig the Zig build (server + CLI + desktop; -Dapp=false = server-only)
src/
main.zig entry point: CLI dispatch, then the server + control plane (auth, routes)
cli.zig the `veil` CLI — a thin client over the server's /api/v1/*
cli/{chat,hub}.zig the interactive chat REPL and the fleet console
gateway/http.zig the HTTP surface: App context, the auth guard, JSON/file helpers
auth/ config/ admin/ accounts + API keys, the encrypted key vault, the admin API
config/lan.zig which addresses this machine is reachable at (the startup banner's URLs)
config/server_config.zig admin-owned runtime settings → data/server-config.json
worker/ the hive and the server-side brain:
chat/{engine,service,tools,context,plan,sync,toolperf,paths}.zig the chat brain — the agentic
turn loop, its REST handlers, tools, context
window, plan board, client file-sync, tool
timings, and conversation paths
sched.zig scheduled tasks (each run is a server chat conversation)
control/{supervisor,writer,fanout}.zig swarm processes, the control bus, event streaming
deploy/service.zig the cast/deploy door + swarm files, bundle, archive, lifecycle
neuron/client.zig the neuron-db memory bridge (fail-open)
builtin.zig builtin_endpoint.zig llamaeng.zig llamashim.c modelpull.zig the BUILT-IN model:
the-veil-12b served by the server itself from a
downloaded GGUF (-Dbuiltin, default true) — store,
loopback engine endpoint, embedded inference, and
the verified weights downloader
gemma4.zig the Gemma-4 wire format, rendered here instead of by the local runtime
browser/ the browser driver: its own profile (launch, cdp) or YOUR Chrome/Edge
through the extension (ext, ext_api), behind one session model + the
per-machine local-host daemon that keeps a session alive across calls
mcp/{client,discovery}.zig find and call the MCP servers already installed on this machine
run.zig agi.zig oscillation.zig rsi.zig vcs.zig tools.zig writer.zig … the moment loop,
the Veil, the self-improvement faculties, the
micro-VCS for concurrent minds
locs/atlas.zig the source atlas — points scouts at nl-rag packs
desk/ veil-desk, the native desktop dashboard — compiled INTO `veil` as the
"desk" module (-Dapp, default true), not a separate shipped binary
docs/ the docs site: architecture map + annotated source (static, home-built
md parser, the desk's own palette from desk/src/theme.zig)
extension/ the Chrome/Edge extension — a thin CDP relay so the veil can drive the
browser you are already signed into (load unpacked; nothing to build)
examples/embedded/ the device-operator worked example
harness/TESTING.md the house test rules — read before adding a test
web/public/ the control-plane UI — index.html, app.js, styles.css, models.json,
embedded into the binary at build time (no bundler, no build step)
models.yaml the ONE model catalog → generated into web/public/models.json
bin/neuron[.exe] the neuron-db memory engine (bundled / built on first run)
Build knobs worth knowing: -Dapp=false (no desktop GUI), -Dbuiltin=false (no embedded inference),
-Dvulkan=false (CPU-only built-in engine, ~50 MB smaller binary). Each one skips fetching its
dependency entirely rather than compiling it unused.
Builds ship on the Releases page, one per
platform. Unzip and run veil — that starts the server and opens the desktop app. No Python, no
toolchain, no first-run build; the memory engine ships inside.
| asset | what it is |
|---|---|
veil-…-windows-x86_64.zip |
the app — Windows |
veil-…-macos-arm64.zip / …-macos-x86_64.zip |
the app — Apple Silicon / Intel |
veil-…-linux-x86_64.zip |
the app — Linux |
veil-server-…-<os>-<arch> |
server only, no desktop — headless hosts, containers, remote boxes |
SHA256SUMS-*.txt |
checksums |
Desktop builds are produced natively per-OS (the GUI links the platform graphics stack and can't be cross-compiled); the standalone server cross-compiles to every target from one runner.
To cut your own, the official builder wraps the whole thing and bootstraps its own toolchain — a missing Zig, Rust, C compiler, or (on Linux) raylib dev library is fetched for you first:
sh scripts/build-official.sh # everything for this host, plus cross-built servers → bin/
sh scripts/build-official.sh --host-only # skip the cross-compiled server matrixIt stages into a private prefix, so it never disturbs zig-out — you can cut a release while the app
is running. NO_BOOTSTRAP=1 skips the dependency step (supply ZIG= / NEURON= yourself).
veil doctor prints server + token health; veil doctor --runtime adds what the app has actually
been doing — learned tool behavior, schedule fail-streaks, per-model token and latency usage — and
reads the on-disk ledgers directly, so it works with the server down.
Related projects: neuron-db (the memory engine) · nl-rag (the knowledge packs) · the-veil-12b (the built-in model) · gary-neuron-emergent (neuron-db's semantic cortex).
MIT. Use it, fork it, build on it.
