This runbook deploys one primary decode node, prefill-compute workers and optional cache-only peers over a private Thunderbolt Bridge. The same services work over LAN/Tailscale with lower endpoint priority.
- dashboard:
https://kakeya.ai/ - health:
https://kakeya.ai/healthz - nodes:
https://kakeya.ai/v1/network/nodes - groups:
https://kakeya.ai/v1/network/groups - token counters:
https://kakeya.ai/v1/network/tokens
Public reads expose aliases, coarse regions and aggregate metrics only. Writes
require X-API-Key.
| Component | Primary | Prefill worker | Cache-only peer |
|---|---|---|---|
| Kakeya RuntimeService / decode | 127.0.0.1:51051 |
no | no |
| Same MLX model loaded | yes | yes | no |
| PrefillWorkerService | no | :53051 |
no |
| PrefillCacheService | runtime/local LRU | co-located | :52051 |
| CapabilityService gossip | enabled | enabled | enabled / pull-only |
| Dashboard/API | 127.0.0.1:8090 |
no | no |
| Cloudflare public edge | kakeya.ai/* Worker |
no | no |
All peers in one inference group must match:
model_id
model_revision
tokenizer_revision / chat template
quantization
KV dtype
cache format version
RoPE hash
layer geometry hash
block size
Changing any field creates a new cache namespace. Never convert or import a best-effort mismatch.
The release launchd asset is:
deploy/launchd/ai.kakeya.grpc-runtime-prefill.plist
It runs Gemma 26B MLX on 51051, queries the peer on 52051, asynchronously
publishes new snapshots, and reports generated/reused token telemetry to the
network API.
Check:
launchctl list | grep ai.kakeya.grpc-runtime-prefill
lsof -nP -iTCP:51051 -sTCP:LISTEN
tail -f ~/.kakeya/grpc-runtime-prefill.logInstall:
openssl rand -hex 24 > ~/.kakeya/network_api_key
chmod 600 ~/.kakeya/network_api_key
bash deploy/install_prefill_network_launchd.shCheck:
launchctl list | grep ai.kakeya.prefill-network
curl -fsS http://127.0.0.1:8090/healthz
curl -fsS http://127.0.0.1:8090/v1/network/nodesUse an isolated venv and copy/sync the repository package. The peer plist is:
deploy/launchd/ai.kakeya.prefill-network-peer.plist
The cache-only plist remains available for rollback or additional RAM-only nodes. In the strict two-Mac decode/prefill profile, allens instead runs the prefill-worker plist below with a co-located cache.
Check from the head over Thunderbolt:
ping -c 3 169.254.27.104
nc -vz 169.254.27.104 52051If nc works but Python/gRPC outbound calls return Errno 65, grant Local
Network access to that Python executable in macOS Privacy & Security. Head→peer
lookup/publish/fetch remains usable while reverse gossip is disabled.
The worker loads the exact same MLX model as the primary, accepts queued prefill-only jobs, writes immutable snapshots into its co-located RAM cache and never serves user decode.
In the strict two-Mac profile, allens-mini runs this role and Primary uses
--prefill-policy remote-required. Workers receive canonical token IDs from
the scheduler; they do not construct their own chat template. The worker stores
the resulting snapshots in its co-located content-addressed cache.
Create a fleet PSK once and copy it to every trusted node:
openssl rand -hex 32 > ~/.kakeya/fleet.psk
chmod 600 ~/.kakeya/fleet.pskInstall the worker:
python -m pip install -r requirements-kakeyalattice.txt
export KAKEYA_WORKER_REPO="$HOME/Kakeya-LLM-Inference-engine"
export KAKEYA_WORKER_PYTHON="$HOME/kakeya-venv/bin/python"
export KAKEYA_WORKER_MODEL="$HOME/kakeya-models/gemma-4-26B-A4B-it-mlx-4bit"
export KAKEYA_CACHE_MODEL_ID="gemma-4-26B-A4B-it-mlx-4bit"
export KAKEYA_MODEL_REVISION="local-4bit-v1"
export KAKEYA_TOKENIZER_REVISION="gemma4-v1"
export KAKEYA_WORKER_NODE_ID="prefill-mini-1"
export KAKEYA_WORKER_BIND="<worker-ip>:53051"
export KAKEYA_WORKER_ADVERTISE="<worker-ip>:53051"
export KAKEYA_LAYER_GEOMETRY_HASH="<same-value-as-primary>"
export KAKEYA_WORKER_SINK="4"
export KAKEYA_WORKER_WINDOW="2048"
export KAKEYA_WORKER_CACHE_GB="8"
export KAKEYA_WORKER_CACHE_MIN_GB="0.25"
export KAKEYA_WORKER_MEMORY_RESERVE_GB="2"
export KAKEYA_WORKER_ADAPTIVE_CACHE="1"
export KAKEYA_CACHE_BLOCK_TOKENS="64"
export KAKEYA_CACHE_FORMAT_VERSION="kakeya-prefill-v3-kl-d4-q38"
export KAKEYA_CACHE_COMPRESSION="kakeyalattice-d4"
export KAKEYA_WORKER_NETWORK="thunderbolt"
export KAKEYA_WORKER_PRIORITY="100"
export KAKEYA_WORKER_RTT_MS="0.55"
export KAKEYA_WORKER_PEER="169.254.187.239:51051"
export KAKEYA_FLEET_PSK_FILE="$HOME/.kakeya/fleet.psk"
export KAKEYA_TENANT_ID="private-fleet"
bash deploy/install_prefill_worker_launchd.shThe primary must use the same compatibility and auth values:
--enable-prefill-cache
--enable-capability-exchange
--peer <worker-ip>:53051
--cache-tenant-id private-fleet
--fleet-psk-file ~/.kakeya/fleet.psk
--cache-compression zlib
--cache-replication-factor 1
For bit-packed KakeyaLattice snapshots, replace the primary compression line
with --cache-compression kakeyalattice-d4 and set
--cache-format-version kakeya-prefill-v3-kl-d4-q38. Both nodes must install
the pinned optional dependency from requirements-kakeyalattice.txt. D4 Q=38
is lossy; live acceptance must validate output quality as well as byte savings.
--cache-peer remains an emergency static override. Normal worker/cache
selection is derived from compatible live capability cards and their TTL/load
metrics.
Create a registration:
KEY="$(cat ~/.kakeya/network_api_key)"
curl -fsS -X POST http://127.0.0.1:8090/v1/network/nodes/register \
-H "Content-Type: application/json" \
-H "X-API-Key: $KEY" \
-d '{"alias":"allens-mini","address":"169.254.27.104:52051","region":"Private","role":"cache"}'Create a paired group:
curl -fsS -X POST http://127.0.0.1:8090/v1/network/groups \
-H "Content-Type: application/json" \
-H "X-API-Key: $KEY" \
-d '{"name":"Thunderbolt Pair","node_ids":["head-mini","allens-mini"]}'Expected invariants:
- both node cards appear within two gossip intervals;
- stale nodes disappear after TTL;
- remote lookup returns the longest available chained-prefix snapshot;
- imported snapshot checksum and compatibility fingerprint match;
- remote failure falls back to local prefill;
- no remote RPC occurs in autoregressive decode;
- a cache miss is assigned to a compatible
PREFILL_COMPUTEworker when its queue+compute+transfer estimate beats blocking the primary; - worker failure/timeout/import rejection resets the verifier and performs full local prefill;
- completed and KV-assisted token counters increase after live calls.
Minimal acceptance:
curl -fsS https://kakeya.ai/healthz
curl -fsS https://kakeya.ai/v1/network/summary
curl -fsS https://kakeya.ai/v1/network/tokens
curl -fsS https://kakeya.ai/v1/network/prefill
PYTHONPATH=.:sdks/python python scripts/verify_remote_prefill_e2e.py \
--address 127.0.0.1:51051 \
--dashboard http://127.0.0.1:8090 \
--tokenizer-id ~/kakeya-models/gemma-4-26B-A4B-it-mlx-4bitBy default the verifier accepts a remote cache hit and requires remote_hits
and tokens_reused to increase. Use --require-worker only when testing the
separate Worker A/B/C path; that mode additionally requires remote_jobs.
Decode throughput is reported separately because all autoregressive decode
remains on the primary.
The logical cross-node KV namespace is available at:
curl -fsS http://127.0.0.1:8090/v1/network/kvfsThe returned kv:// URI and mount table virtualize naming and management only.
Payloads remain in each Mac's physical RAM and are copied into Primary once
before decode; coherent_shared_memory is always false. Mounts are marked
hot for Primary and cold-offload for worker/cache peers. Remote imports are
promoted into the hot LRU, while eviction leaves the cold copy untouched.
Enable the bounded, memory-only first-append capture queue on the primary:
--cache-fill-capture-size 256
The maintenance endpoints require the network API key even when other read endpoints are public. Captured token IDs remain in process memory and are removed when drained; reports contain only salted capture IDs and token counts.
During a maintenance window, start real gRPC chat sessions and run:
PYTHONPATH=.:sdks/python python scripts/fill_prefill_cache_from_live_grpc.py \
--tokenizer-id ~/kakeya-models/gemma-4-26B-A4B-it-mlx-4bit \
--target-one 0.90 \
--target-two 0.95 \
--churn-gb 0.9The harness controls head and allens independently, stops on memory pressure,
publish failures, fallbacks, or /tmp/kakeya-cache-fill.stop, and never expects
resident fleet usage to exceed the configured 1+8 GiB ceiling. Churn is accepted
when bytes_evicted increases while resident bytes remain bounded.
Run from Primary; services are started only if missing and remain running:
bash scripts/run_prefill_architecture_benchmark.sh \
--output-tokens 32 \
--report /tmp/kakeya-prefill-benchmark.jsonThe task runs remote_compute, primary_hot_hit, and
allens_cold_restore, recording client-side append latency, TTFT, decode
latency/tokens-per-second, E2E throughput, and server-side hit/promotion deltas.
Reports never persist prompts, token IDs, cache keys, raw addresses, or user
paths.
Benchmark APIs are public reads and API-key writes:
curl -fsS https://kakeya.ai/v1/network/benchmarks
curl -fsS https://kakeya.ai/v1/network/benchmarks/live
curl -fsS https://kakeya.ai/v1/network/benchmarks/<run-id>The Benchmarks dashboard tab shows live progress, phase comparison, history,
and complete redacted stage details.
Run two logical agents through real multi-round model inference:
bash scripts/run_agent_gan_demo.sh \
--rounds 1 \
--output-tokens 64 \
--report /tmp/kakeya-agent-gan-demo.jsonFor every Generator and Critic turn the task first asks allens to Prefill the complete agent context, then performs the actual inference against the promoted Primary hot snapshot. Terminal output contains the real proposal/critique; persisted reports contain only output length/hash and metrics. The report separates inference-only KV hit rate from whole-workload hit rate including warmup, plus per-agent and aggregate token throughput/latency.
One Generator/Critic cycle is the safe default for the current 16GB allens Gemma worker. Additional rounds grow agent histories and may exceed the strict remote-prefill timeout; increase worker memory/timeout before enabling them.
For an interactive terminal REPL with real-time token output:
bash scripts/run_agent_gan_repl.shType a prompt at prompt>. The terminal prints allens Prefill progress, then
streams generator> and critic> tokens as they arrive, followed by KV hit
rate, decode tok/s, generation latency and E2E tok/s. /quit exits; services
remain running. Reports remain redacted and appear in the Benchmarks tab.
Generation uses 64-token streaming chunks and automatically continues to the
model's EOS. --max-response-tokens defaults to 0 (no client-side cap);
operators may explicitly set a positive cap, which marks the stage incomplete
instead of presenting a truncated answer as successful. The Critic receives the
Generator completion status and must not penalize an honest statement that an
open problem has no accepted proof.
The REPL ignores external SIGTERM; its shell supervisor restarts signal-based
exits. Only /quit, /exit, or EOF is treated as approval to stop.
The REPL input boundary is a strict state machine. Before a goal exists, raw
text sets the initial immutable goal. Afterwards only /continue,
/steer <text>, /new <goal>, and /quit are accepted; bare text and copied
runtime output are rejected. No stdin is read while inference is running.
Completed Generator/Critic text is atomically checkpointed with mode 0600 at
~/.kakeya/agent_gan_state.json, allowing an exact /continue after restart.
--recover-run <id> --recover-log <path> --auto-continue reconstructs a
checkpoint from a complete timestamped run and resumes without using any
partial Terminal echo.
Continuous auto-loop is enabled by default. After every successful
Generator/Critic turn, the READY state internally queues /continue; it does
not wait for Terminal automation. A short boundary window still accepts an
explicit whitelisted command such as /quit or /steer. Any inference
exception pauses auto-loop and preserves the last complete checkpoint.
--no-auto-loop restores manual one-turn-at-a-time operation.
Rigorous mathematical problems discovered outside the running turn can be
queued in ~/.kakeya/agent_gan_critic_inbox.json. The next turn appends every
pending issue to the previous Critic correction for Generator remediation and
to the new Critic steering for verification. The full issue list is timestamped
in the Critic log and recorded by ID/count in benchmark config. Issues remain
pending across failures and are marked consumed only after a complete,
successful Generator/Critic turn.
Mathematical closure is tracked separately in
~/.kakeya/agent_gan_proof_ledger.json. Each unresolved obligation has a stable
ID and is injected through a dedicated PROOF OBLIGATION LEDGER section, never
through Terminal input. Generator must emit one ISSUE_RESPONSE per ID; Critic
must emit a structured ISSUE_VERDICT with status, evidence, and missing lemma.
Missing or structurally weak verdicts remain UNRESOLVED and are carried into
the next auto-loop turn. Infrastructure success does not close proof issues.
Generator output is always streamed in full to Terminal and passed verbatim to
the Gemma Critic. Sampling, truncation, summarization, independent chunk scores,
and semantic fallback are forbidden. A global Critic score is valid only when
review_scope=full, critic_context_tokens=generator_full_tokens, and
critic_omitted_tokens=0. Long Prefill operations emit a heartbeat every 30
seconds; on the 16GB allens worker, full-context Critic Prefill may take 15–25
minutes.
The interactive Critic uses goal_anchored_recursive_gan_v3: numeric scores and blanket
approval are forbidden. It ignores prizes, money, prestige, style, and other
proof-irrelevant facts. It attacks the central stopping claim, recursively
decomposes it into proof obligations, records arguments/counterarguments and
dependencies for every node, and loops until every leaf is either explicitly
derived or a precisely stated open lemma. It then reports the smallest
unresolved frontier and next adversarial step.
The first substantive REPL input becomes an immutable research goal. /continue
feeds the complete previous Generator and Critic outputs into the next
Generator turn; /new <goal> is the only way to switch topics. Inputs beginning
with generator>, critic>, prompt>, [metrics], [allens], errors, or
tracebacks are rejected so Terminal output cannot contaminate the research
objective. Every Critic turn emits an explicit Goal Alignment decision and
restores the anchored proof frontier when drift is detected.
Every REPL session mirrors its complete Terminal transcript to
~/.kakeya/logs/agent_gan_repl.log by default. Log lines use local ISO-8601
timestamps and include user input, immutable goal hash, run ID, Generator and
Critic streaming output, heartbeats, metrics, and explicit
inference-start/inference-complete/inference-failed markers. Writes flush
immediately for post-failure review. Use --log-file PATH to override the
destination.
The worker reserves estimated final-snapshot capacity before model compute,
prevents adaptive shrink from consuming active reservations, then atomically
publishes and leases the final snapshot before adding optional intermediate
boundaries. The 16GB allens deployment uses a 1 GiB cache floor and 0.5 GiB
memory reserve. Capacity failures are rejected before Prefill starts.
Snapshot capacity estimation uses retained KV tokens, not unbounded history:
min(prompt_tokens, sink + window) * estimated_bytes_per_token. Long
auto-loop histories therefore stop increasing the reservation once the MLX
sliding window is full.
The MLX worker exports and compresses only the final chained-prefix snapshot.
The final hash commits every preceding token block, so exporting a growing full
snapshot at every 64-token boundary is redundant and creates quadratic
serialization/compression work for long Critic contexts.
Model compute is segmented independently from the 64-token content-addressed
hash boundary. The default segment is 256 tokens, keeping each allens step below
the five-minute research budget at the measured Prefill rate. Worker job status
updates tokens_computed after every segment; Terminal heartbeats display
tokens, percentage, and ETA.
The worker injects sink + window as an explicit immutable retained-token cap
into PrefillJobStore. Snapshot capacity checks therefore remain window-aware
even if an RPC compatibility object omits its window field; long prompts cannot
fall back to full-length estimates after deployment or process restart.
Karpathy-style optimization lives in autoresearch/prefill/. Humans edit
program.md, the research agent edits only candidate.py, and immutable
prepare.py evaluates full-context correctness plus cold Critic Prefill time.
Candidates are retained only when every semantic/topology constraint passes and
the metric improves; results.tsv records experiments.
autoresearch/prefill/supervisor.py is the top-level runtime: it uses a real
Gemma strategy-agent call to rewrite candidate.py, deploys the selected chunk
strategy to allens, clears both cache tiers, runs one real full-context
Generator/Critic/Proof-Ledger experiment, evaluates the fixed report, compares
the lexicographic proof/prefill baseline, atomically keeps or restores
candidate/state/ledger, appends results.tsv, and repeats. Any proposal,
deployment, telemetry, GAN, evaluator, or restore failure terminates the
supervisor; there is no fallback path.
Run:
bash scripts/run_autoresearch_gan.sh --iterations 1Interactive prompt templates are deterministic and contain no per-run nonce, so repeating the same task can reuse allens cold-tier and Primary hot-tier KV.
The cache is an optimization; inference correctness does not depend on it.
-
Stop the cache services:
launchctl bootout "gui/$(id -u)/ai.kakeya.prefill-network" launchctl bootout "gui/$(id -u)/ai.kakeya.grpc-runtime-prefill"
-
Restart the previous RuntimeService without
--enable-prefill-cache. -
Roll back the Cloudflare Worker deployment with Wrangler versions/deployments.
-
agent.kakeya.airemains available as the direct gateway origin.
In-memory cache entries require no migration or cleanup after rollback.
The live MVP assumes trusted private Macs. Before accepting third-party nodes:
- require mTLS or fleet-PSK authentication;
- bind signed node identity to
node_id; - HMAC prompt-derived block hashes;
- rate-limit registration, lookup and publish;
- cap block size and stream bytes before allocation;
- maintain revocation and audit logs;
- never expose raw prompts, hashes, IPs or cache payloads in the public UI.