Palimpsests / Air-gapped
Air-gapped · on-prem · no network egressAir-gapped LLM inference with a verifiable audit trail
Palimpsests runs the model, its attention state and its log on the machine where the data already sits. No network egress is needed to answer a request — or to verify the record.
What stays local
The model runs through llama.cpp (or a local Ollama daemon); attention state is held and restored in-process; every request lands in a hash-chained log on the same host. Nothing is uploaded and nothing is phoned home. What you export is a record that can be verified without exposing what it recorded.
Time without a trusted clock
Every timestamp carries its trust level — UNKNOWN, UNSYNCED, HW_RTC or NTP_SYNCED. A device that does not know the time says so: an UNKNOWN record carries no wall-clock value at all. Ordering never depends on the clock; it is carried by sequence numbers in the chain.
Anchors without a network
The chain alone proves that nothing inside it changed. Proving that nothing was cut off the end needs the latest head kept somewhere else — an anchor.
| Anchor | Where the head lives |
|---|---|
| File | a file you control, outside the log |
| PKCS#11 token (since 0.11) | a data object on a hardware token the host can read but cannot silently rewrite |
| SCITT transparency service (since 0.11) | one signed statement per published head, only when your policy allows the link — no record content leaves the device |
The token is the mechanism; a claim about a particular deployment depends on the token actually used.
Measured on modest hardware
- Engine: on an Intel Arc integrated GPU (Vulkan), with 1.5B and 7B models, the tool loop runs at parity with a hand-tuned
llama-server— in-process, with no server to run. - Verifier: a one-million-record chain on an Intel Core Ultra 9 185H machine (Windows 11) — working set after
verify()4.77 GB in 0.11, 0.57 GB in 0.12;verify()time 102–107 s, then 43–46 s.
What we don't claim
There is no discrete-GPU run yet: an integrated GPU flatters every mechanism that saves prefill, so these ratios compress where prefill is fast. The audit log's implementation has not been independently pen-tested.