File size: 60,451 Bytes
28c70af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0ab4e0b
 
 
 
 
28c70af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
# Using bankml and Savante

How to install bankml, start it, and use it on your own computer. Everything runs on this machine: the model, the
gate that verifies it, and Savante's page. New to Savante's page and her voice? Read **[playback.md](playback.md)**.
Every flag, environment variable and service setting is in **[install.md](install.md)**; each source module, its
use and its limits, in **[modules/](modules/README.md)**.

This guide covers the latest release, 0.3.6. Sections marked 0.3.7, 0.3.8 or 0.3.9 describe the next release, which
is not yet tagged; the details are in the *Unreleased* sections of [CHANGELOG.md](../CHANGELOG.md).

- [1. Install](#1-install)
- [2. What you need](#2-what-you-need)
- [3. Start, stop, check](#3-start-stop-check)
- [4. Talk to Savante](#4-talk-to-savante)
- [5. Models](#5-models)
- [6. By hand: build, verify, serve](#6-by-hand-build-verify-serve)
- [6a. Ollama's API](#6a-ollamas-api)
- [6b. The C API](#6b-the-c-api)
- [6c. The bankML console: bankML as itself (0.3.7)](#6c-the-bankml-console-bankml-as-itself-037)
- [7. Let others watch (view mode, on the LAN)](#7-let-others-watch-view-mode-on-the-lan)
- [8. The files Savante keeps: `.history`, `.memory`, `.prompt`](#8-the-files-savante-keeps-history-memory-prompt)
- [8a. Proof of data without the data](#8a-proof-of-data-without-the-data)
- [8b. Custom agents from the Savante template](#8b-custom-agents-from-the-savante-template)
- [8c. THOT bundles: the dataset an iNFT points to](#8c-thot-bundles-the-dataset-an-inft-points-to)
- [8d. PostgreSQL: publish an agent, load one back](#8d-postgresql-publish-an-agent-load-one-back)
- [8e. iNFT: mint an agent, load one from a token](#8e-inft-mint-an-agent-load-one-from-a-token)
- [9. Receipts, and how to check an answer](#9-receipts-and-how-to-check-an-answer)
- [10. Savante's canon and the iNFT ledger](#10-savantes-canon-and-the-inft-ledger)
- [11. Testing and the release gate](#11-testing-and-the-release-gate)
- [12. Troubleshooting](#12-troubleshooting)
- [13. Reference: commands, ports, environment](#13-reference-commands-ports-environment)

---

## 1. Install

One command, from a fresh machine to Savante answering on your own CPU:

```sh
git clone https://github.com/cryptoAGI/bankml && cd bankml
chmod +x install.sh
./install.sh
```

`chmod +x install.sh` makes the installer executable. A `git clone` keeps the bit already, so the line is harmless
there; a copy that arrived another way (a zip download, a USB stick, some file systems) may have lost it, and then
`./install.sh` answers *permission denied*. `bash install.sh` works either way.

or, without cloning first (the installer clones bankml to `~/bankml`, or to `$BANKML_DIR`):

```sh
curl -fsSL https://raw.githubusercontent.com/cryptoAGI/bankml/main/install.sh | bash
```

It runs these steps in order, and stops at the first one that fails, saying why:

| step | what it does |
|---|---|
| `check` | Rust 1.99+ (pinned in `rust-toolchain.toml`), Python 3.10+, git, curl, tar, ss, sha256sum; reports AVX2, RAM and free disk |
| `build` | `cargo build --release` and `cargo test --release` |
| `engine` | downloads llama.cpp b11192 (17 MB) and refuses it unless its sha256 is the published one |
| `python` | uses your Python if it has Gradio and numpy; otherwise makes a venv with Gradio 3 |
| `canon` | clones Savante's canon to `~/cryptoAGI/savante` (beside jaimla and luvai) if it is not there; an existing canon is left as it is, and a machine that still has the older `~/savante` keeps using it |
| `model` | imports Bonsai-8B (1.16 GB) if absent, verifies it, and starts `bankml serve` with llama-server behind it |
| `start` | starts Savante on `http://127.0.0.1:7873` |

Run any steps on their own: `./install.sh build`, `./install.sh model start`. Options: `--skip-tests`, `--no-start`,
`--view` (also start view mode, §7), `--voice` (also install Savante's voice; see [playback.md](playback.md)).
Nothing needs sudo. What the installer chose is kept in `~/.local/share/bankml/install.env`, and logs in
`~/.local/share/bankml/logs/`.

## 2. What you need

| what | where it comes from |
|---|---|
| Linux on x86_64 (AVX2 for the fast kernels; without it bankml is correct, just slower) | — |
| Rust 1.99 (pinned in `rust-toolchain.toml`; rustup fetches it) | [rustup](https://rustup.rs) |
| Python 3.10+ | your system |
| about 2 GB of free RAM and 3 GB of free disk | the 1-bit 8B model maps 1.16 GB |

The installer fetches the rest: llama.cpp b11192 (`llama-b11192-bin-ubuntu-x64.tar.gz`, sha256
`34cf6fa5de9da0db3932c78fe15fed2fbca17451e665dac0a4f6a3c8fc881ec7`), Gradio 3 if needed, the model from
[PYTHAI/Bonsai-8B-gguf-fork](https://huggingface.co/PYTHAI/Bonsai-8B-gguf-fork), and Savante's canon (§10). On
another platform, build llama.cpp b11192 yourself and set `BANKML_LLAMA_SERVER` to its `llama-server`.

The ternary model (2.31 GB) also works and gives better answers, but llama.cpp runs it about 5× slower on a CPU (see
[PERFORMANCE.md](PERFORMANCE.md)).

## 3. Start, stop, check

```sh
./install.sh start           # bankml serve (if not up) is started by `model`; this starts Savante's page
./install.sh start --view    # also the read-only page for your network
./install.sh status          # what is running, and which model bankml verified
./install.sh stop            # stops the page(s), bankml serve and llama-server
```

What runs where:

```
 you (browser) ──► Savante's page (127.0.0.1:7873) ──► bankml serve (127.0.0.1:18093) ──► llama-server (127.0.0.1:18092)
                                                        verifies the model file            runs the verified file
 anyone on the LAN ──► view mode (0.0.0.0:7874, read-only)
```

- **bankml serve** is the gate. It refuses to start unless the model passes the guard, its sha256 equals the pin in
  the model's `FORK.json`, and llama-server serves that same file. Every answer carries a receipt (§9).
- **With the native engine** (`--native`; Savante's `auto` engine setting chooses it for the ternary files), there is
  no llama-server: `bankml serve` computes the answer itself and also listens on 18092 with llama-server's endpoints.
- **Savante's page** is reachable only from this computer.
- **View mode** is a read-only page for others on your network (§7).

## 4. Talk to Savante

Open **http://127.0.0.1:7873**, type a question, press **Send**.

- The first answer takes about **2 minutes** on a laptop CPU, while the model reads Savante's instructions. Later
  answers start sooner.
- Every answer is a **draft · not a finding**. The chat shows the answer alone; its clock, receipt (§9) and prompt
  provenance are on the **Admin** tab under *the last answer*, and in `.history` and the Responses tab.
- To ask for a review, say so: *"review: is Bonsai-8B ready to serve mindX?"*. She then answers as FINDINGS /
  VERDICT / RATIONALE / CONDITIONS / RISKS WATCHED. The 1-bit carrier was graded REJECT for review duty in the sAGI
  carrier test, so treat its verdicts as drafts.
- **Stop** cancels an answer; **New session** starts a fresh conversation.

The Interaction tab is only the conversation: the chat, the question field, the clock and the portrait (press it for
the card; **← back to Interaction** or Esc closes it, and every other tab has a **← Interaction** button). The
**Admin** tab holds the settings: the system prompt (`.prompt`, §8), `.memory`, max tokens, temperature, the carrier,
and the engine's CPU threads and RAM budget (**Apply** restarts the engine with them, and restores your previous
settings if it will not start; the engine choice and n-gram speculation are under *advanced*). The other tabs hold
your history, notes, metrics, integrity checks and the verifier. [playback.md](playback.md) walks through the
page, her card and her voice.

## 5. Models

The **Models** tab and `sAGI/models.py` import a model, verify it, and switch to it:

```sh
python3 sAGI/models.py list                 # what is here, pinned or not, and which one is running
python3 sAGI/models.py catalog              # the curated catalogue, and whether each fits this machine
python3 sAGI/models.py import qwen3-1.7b    # a catalogue id, a Hugging Face URL, or ollama:NAME:TAG
python3 sAGI/models.py use Qwen3-1.7B-Q8_0.gguf   # switch (rolls back if the new model fails)
```

An import is kept only if its sha256 equals the one its publisher lists, `bankml guard` passes, and a `FORK.json`
pin is written; `bankml serve` checks that pin before every start. Only open-source models are accepted (Gemma and
Llama are refused), and an import that would not fit this machine's RAM or disk is refused with the numbers.

**Converted models (0.3.4).** SmolLM2-135M-Instruct and mindX's `mindx-gen39` have no published F16 GGUF. They are
made from their pinned safetensors by llama.cpp b11192's own converter, and `sAGI/models.py` pins the result against
the recorded conversion (`CONVERTED` in the file says how each was made):

```sh
python3 llama.cpp/convert_hf_to_gguf.py SmolLM2-135M-Instruct/ --outtype f16 --outfile .models/SmolLM2-135M-Instruct-F16.gguf
#   (b11192, commit 171e8846; the safetensors at HuggingFaceTB/SmolLM2-135M-Instruct@12fd25f7; torch is needed here only)
python3 -c "import sys; sys.path.insert(0, 'sAGI'); import models; print(models.adopt('SmolLM2-135M-Instruct-F16.gguf'))"
#   → hashes the file against the recorded e9aba089…, re-reads the source's LFS sha256 at the revision, writes FORK.json
```

What bankML's own forward pass plays (`bankml serve --native`, `bankml generate`, the C API): Qwen3 in Q1_0 and
Q2_0_g64 (Bonsai-8B, Ternary-Bonsai-8B, Bonsai-1.7B) and the Llama architecture in F16 (SmolLM2-135M-Instruct,
`mindx-gen39`). Anything else is refused with the reason; `/api/tags` lists it with that reason.

## 6. By hand: build, verify, serve

What the installer does, step by step:

```sh
cargo build --release && cargo test --release
target/release/bankml guard  .models/Bonsai-8B-Q1_0.gguf                   # play | refuse (reason) | need_more
target/release/bankml verify .models/Bonsai-8B-Q1_0.gguf --fork FORK.json  # guard, then the sha256 pin
target/release/bankml serve  .models/Bonsai-8B-Q1_0.gguf --fork FORK.json \
    --spawn /path/to/llama-b11192/llama-server --threads 3 --ctx 2048
python3 sAGI/savante.py --mode interact
```

`verify` exits 0 on play, 2 on refuse, 1 on an I/O error, 3 on a truncated file. `serve` can also sit in front of a
llama-server you started (`--upstream 127.0.0.1:18092`); it then checks through `/props` that the server runs the
verified file. It listens on `127.0.0.1:18093` and speaks the OpenAI API:

```sh
curl -s 127.0.0.1:18093/bankml     # what was verified, when, and what is upstream
curl -s 127.0.0.1:18093/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"messages":[{"role":"user","content":"Say hello. /no_think"}],"max_tokens":32}'
```

### JSON mode and grammars (0.3.3)

`/v1/chat/completions` (and the C API's `bankml_chat`) takes what llama-server takes, resolved the way it resolves it:

```sh
B=127.0.0.1:18093; J='Content-Type: application/json'
curl -s $B/v1/chat/completions -H "$J" -d '{"messages": [{"role": "user", "content": "Describe a cat."}],
    "response_format": {"type": "json_object"}, "temperature": 0, "max_tokens": 64}'
curl -s $B/v1/chat/completions -H "$J" -d '{"messages": [{"role": "user", "content": "Is the sun a star?"}],
    "grammar": "root ::= (\"yes\" | \"no\") \".\""}'
echo 'Describe a cat.' | target/release/bankml generate .models/Bonsai-8B-Q1_0.gguf --json --max 64
```

- **`response_format: {"type": "json_object"}`**, and the schemas `{}` and `{"type": "object"}` (in `response_format`
  or the top-level `json_schema`), give **llama-server b11192's grammar for this template**, byte for byte. Its root
  begins with the generation prompt, which the sampler takes in before the first token, as the server does. Under
  that grammar the model may open with a fenced ```` ```json ```` block, and `content` is the JSON value alone. That
  is the object once it closes, or what there is of it when `max_tokens` cuts the answer. A stream sends the
  content's growth.
- **`grammar`**: any GBNF that llama.cpp parses, with root `root`. It is taken as a user grammar, with no prefill, and
  `content` is the text as generated. A grammar llama.cpp would not parse is refused with its parser's reason.
- **Every token is drawn as llama-server draws it under a grammar.** The usual sampler chain runs first. If its token
  breaks the grammar, the logits are masked and the chain runs again, which takes a second draw from the seeded
  generator. Greedy and seeded answers are token-identical to the server (the oracle is in [oracles.md](oracles.md)).
- **Any other JSON schema** (0.3.5): in `response_format` (`json_schema`, or `json_object` with a `schema`), the
  top-level `json_schema`, or Ollama's `format: {…}`, it is converted into the grammar llama-server b11192 builds for
  it on the model's own template (`bankML/schema.rs`; byte-identical on 173 of 173 schemas on each template against
  llama.cpp's own code), and answered as JSON mode is: prefilled, fenced or bare, `content` the value alone (a
  top-level string, number or literal included). The answers are token-identical to llama-server's on every native
  model, greedy and seeded. A schema llama.cpp refuses is refused with its message.
- **An answer cut by `max_tokens`** gets llama-server's content: the value as far as it got (an unfinished escape
  dropped), or, when nothing of the value came yet (`"```json\n"`), the raw text, as the server answers a whole
  request. A stream sends only the value's growth.
- **Refused, with the reason:** a schema llama.cpp b11192 refuses; a `pattern` llama.cpp would turn into a grammar
  that is not UTF-8; a non-object or `null` top-level `json_schema` (llama-server fails those requests);
  `response_format` together with `grammar`; `json_schema` together with `grammar`; a `response_format` type other
  than `text`, `json_object` or `json_schema`.

### The context limit, slots and logprobs (0.3.8)

In the next release (not yet tagged); each is checked against llama-server b11192's own answers
([CHANGELOG.md](../CHANGELOG.md), [modules/serve.md](modules/serve.md), [modules/prompt_cache.md](modules/prompt_cache.md)).

- **At the context limit** the native engine answers as llama-server does with context shift off: a generation that
  reaches `--ctx` stops with `finish_reason: "length"`; a prompt that does not fit is refused at once with HTTP 400
  and llama-server's `exceed_context_size_error` body, which names both counts — so a client can trim and retry.
- **Slots.** Start `bankml serve --native` with `--slot-dir DIR` (created at start). Save the conversation's KV cache
  by name, restore it later — in another process too — and skip recomputing the prompt. Without `--slot-dir` the
  slot actions answer 501, as llama-server's do:

  ```sh
  curl -s 127.0.0.1:18092/slots/0?action=save    -H 'Content-Type: application/json' -d '{"filename":"warm.bin"}'
  curl -s 127.0.0.1:18092/slots/0?action=restore -H 'Content-Type: application/json' -d '{"filename":"warm.bin"}'
  curl -s 127.0.0.1:18092/slots/0?action=erase   -H 'Content-Type: application/json' -d '{}'
  ```

  The file is checked when it is read (the model's sha256 and a sha256 trailer): a file from another model or a
  damaged one is refused and the slot emptied (never half-filled), so the next answer simply recomputes. Filenames
  are plain names inside the directory; a path is refused.
- **Conversations that take turns.** The native engine has one slot and serves requests one at a time, as
  llama-server `-np 1`. Like llama-server, it keeps the states of other conversations in RAM (`BANKML_CACHE_RAM`, MiB;
  0 off, -1 no limit; unset, 8192 MiB but at most a quarter of the memory available at load), so a conversation that comes back after
  another does not recompute its whole history.
- **Your bankML from a web page (0.3.9).** `./install.sh start --space` lets the bankML Space's page
  ([PYTHAI/bankml](https://huggingface.co/spaces/PYTHAI/bankml), *Your own bankML*) talk to your own `bankml serve`
  from your browser: your CPU, your verified model, a receipt on every answer, nothing sent anywhere else. It starts
  serve with `--allow-origin https://pythai-bankml.static.hf.space`, so only that one page gets CORS headers;
  `./install.sh start --no-space` takes it back.
- **A smaller conversation memory (0.3.9, in progress).** `BANKML_CACHE_TYPE=q8_0` keeps the KV cache in q8_0 instead of f16,
  about half the memory, so a long context fits on a small machine. It is llama.cpp's `--cache-type-k q8_0
  --cache-type-v q8_0` exactly, with the Hadamard rotation llama.cpp applies around a quantized cache, and gives the
  same tokens as llama-server so configured (`testing/kv_oracle.py`). A slot saved with one cache type is refused by
  an engine with the other.
- **Logprobs.** `"logprobs": true, "top_logprobs": 5` on `/v1/chat/completions` returns each token's log-probability
  and the five most likely alternatives (`top_logprobs` defaults to 20), the same floats llama-server reports.
  Streamed answers carry each token's entry in the chunk its text makes. `top_logprobs` without `logprobs` is refused
  with llama-server's message. llama-server's own `/completion` endpoint (and its `n_probs`) is not served; use `/v1`.

## 6a. Ollama's API

Since 0.3.1 `bankml serve --native` also speaks Ollama's API, over the same engine, gate and receipts. **Native
only:** bankML answers what its own verified forward pass can do and refuses the rest with HTTP 400
`{"error": "…"}` that says why; it never hands a request to llama-server or to Ollama. What mindX asks of Ollama,
and what is still missing, is in [OLLAMA.md](OLLAMA.md).

```sh
target/release/bankml serve .models/Bonsai-8B-Q1_0.gguf --fork ~/.local/share/bankml/forks/Bonsai-8B-Q1_0.gguf.FORK.json \
    --native --registry --ctx 2048            # --registry [DIR]: every model pinned there, by name; --keep-alive 5m
B=127.0.0.1:18093; J='Content-Type: application/json'
curl -s $B/api/version
curl -s $B/api/tags                          # every pin: name, size, digest = its sha256, details, "bankml": {"native", "reason"}
curl -s $B/api/ps                            # the resident model and its expires_at
curl -s $B/api/show -H "$J" -d '{"model": "bonsai-8b-q1_0"}'
curl -s $B/api/chat -H "$J" -d '{"model": "bonsai-8b-q1_0", "messages": [{"role": "user", "content": "Name a prime."}],
    "stream": false, "options": {"temperature": 0.3, "seed": 7, "num_predict": 32}}'
curl -s $B/api/generate -H "$J" -d '{"model": "bonsai-8b-q1_0", "system": "Answer in one word.", "prompt": "Capital of France?", "stream": false}'
curl -s $B/api/generate -H "$J" -d '{"model": "bonsai-8b-q1_0", "keep_alive": 0}'          # unload now (done_reason "unload")
curl -s $B/api/generate -H "$J" -d '{"model": "ternary-bonsai-8b-q2_0_g64", "keep_alive": "30m"}'  # verify and load (done_reason "load")
```

- **Names.** A model's name is its file's stem, lower-cased: `bonsai-8b-q1_0`, `ternary-bonsai-8b-q2_0_g64`,
  `smollm2-135m-instruct-f16`, `mindx-gen39-f16`.
  - `:latest` and the file name are accepted as aliases.
  - Since 0.3.4 so is the name without its weight-type suffix, when exactly one pin has that base: `mindx-gen39`
    (Ollama's tag for mindX's generation 39), `smollm2-135m-instruct`, `bonsai-1.7b`.
  - A request without a model uses the one `serve` started with.
  - An unknown name answers 404.
- **Where models come from.** Without `--registry`, only the startup model is served. With it, every GGUF pinned in
  `DIR` (default `$BANKML_FORKS`, else `~/.local/share/bankml/forks`) can be asked for. The file is looked for
  beside the startup model and in `DIR`.
- **`/api/tags` lists every pin, honestly.** A pin bankML cannot play is listed with `"native": false` and the reason.
  Asking for one answers 400 with that reason. Two examples:
  - `qwen3-0.6b-q8_0`: Q8_0 weights, O3;
  - `qwen3-0.6b` (Ollama's blob): Q4_K weights, O3.
- **One resident model.** A request for another model drops the resident one, then verifies the new one before it
  loads it: the guard, then the sha256 pin, as at the start.
- **`keep_alive`** is read as Ollama reads it:
  - `"5m"`, `"1h"`, `"1h30m"`, or a number of seconds;
  - `0` unloads after the answer;
  - a negative value keeps the model for good.

  `/api/*` requests that name none get `--keep-alive` (default `5m`). An idle model is dropped when its time runs out,
  and its memory map with it. The OpenAI and llama-server endpoints name no keep-alive. They keep the model resident,
  as 0.3.0 did, and load the startup model again if nothing is resident.
- **An empty prompt only loads or unloads.** That is a `/api/generate` without a `prompt`, or a `/api/chat` without
  `messages`. The answer's `done_reason` is `load` or `unload`.
- **Streaming is on by default**, as in Ollama: NDJSON (`application/x-ndjson`), one object per line, then a final
  `done: true` object. That object carries:
  - `done_reason` (`stop` or `length`);
  - `total_duration`, `load_duration`, `prompt_eval_duration` and `eval_duration`, in nanoseconds;
  - `prompt_eval_count`: the whole prompt;
  - `eval_count`: generated tokens, the end-of-turn token included, as llama-server counts it;
  - `bankml_cache_n`: prompt tokens reused from the cache;
  - `bankml_receipt`, as on `/v1/chat/completions`.
- **Options.**
  - **Honoured:** `temperature`, `top_k`, `top_p`, `min_p`, `seed`, `num_predict` and `stop` (also on `/v1`).
    `top_k` is reproduced from 1 to 128; 0 or a larger value is refused, because llama.cpp sorts larger sets another
    way.
  - **`num_ctx`:** accepted up to the served `--ctx`, refused above it. Under it (0.3.5) a conversation that does not
    fit loses its oldest messages first, as Ollama's do (the last message and the system messages stay); a prompt that
    still does not fit is refused with its size.
  - **Penalties (O2):** `repeat_penalty`, `repeat_last_n`, `presence_penalty` and `frequency_penalty` are honoured as
    llama-server b11192 applies them, token for token on its oracle: the window holds the last `repeat_last_n` tokens
    (default 64) **including the prompt's**, a repeated token's logit is divided by the repeat penalty (multiplied when
    not positive), then `count × frequency + presence` is taken off. `repeat_last_n` below 0 and a repeat penalty of 0
    or less are refused with llama-server's message. Also on `/v1` and as a Modelfile `PARAMETER`.
  - **The rest of the chain (0.3.7, next release; [modules/sampler.md](modules/sampler.md)):** `typical_p` on `/api/*` and `/v1`; on `/v1` (not Ollama options) also
    `top_n_sigma`, `xtc_probability`, `xtc_threshold`, `dynatemp_range`, `dynatemp_exponent`, `dry_multiplier`,
    `dry_base`, `dry_allowed_length`, `dry_penalty_last_n` and `dry_sequence_breakers` — each as llama-server b11192
    applies it, token for token on its oracle, in its default order. Clamped as it clamps (`top_p`, `min_p`, the XTC
    fields to [0, 1]); refused as it refuses (a negative DRY window or allowed length, empty breakers). `mirostat` and a
    custom `samplers` order are refused: they are not reproduced.
  - **Ignored**, because they do not change the answer bankML gives: `num_thread`, `num_batch`, `num_gpu`, `use_mmap`
    and the other resource options.
  - **Refused:** any other option.
- **`format: "json"` is JSON mode (0.3.3).** It is answered by the same grammar llama-server b11192 uses for
  `response_format: {"type": "json_object"}` (below), so `/api/chat` with `format: "json"` gives llama-server's tokens.
  `format: {}` and `format: {"type": "object"}` are the same request. Since 0.3.5 any other schema object is
  answered with the grammar llama-server builds for it (below). `format` with `raw: true` is refused (the grammar
  begins with the template's generation prompt, which a raw prompt does not have).
- **Refused, with the reason:**
  - a `format` schema llama.cpp b11192 refuses (with its message);
  - `tools` and `tool_calls` (O6);
  - `images`;
  - `suffix`, `template`, `context`;
  - `think: true` (the template is rendered with thinking off; `think: false` is accepted);
  - `/api/embed` (O7), `/api/pull`, `/api/push`, and `DELETE /api/delete` of a pinned file (derived models: below).
- **`raw: true`** tokenizes `prompt` as given, with special tokens and no template. `system` is ignored then, as in
  Ollama.
- **The loopback rules still apply.** POST bodies need `Content-Type: application/json` (`curl -d` alone sends a
  form type and gets 415), and `Host` must be a loopback name.

### Derived models and conversion: `bankml create`, `bankml convert` (O5, 0.3.5)

`ollama create`, natively. A derived model is a layer (system prompt, parameters, stop strings, example messages,
licence) over a **pinned** base, written as `<registry>/<name>.MODEL.json`. Weights are never copied. Loading it
verifies the base exactly as a pinned model is verified. The details are in
[OLLAMA.md](OLLAMA.md#what-o5s-first-cut-built-035). mindX's persona layer, made this way, answers token-identically
to llama-server given the same persona (`testing/persona_oracle.py`).

```sh
# mindX's flow, without Ollama: the merged safetensors directory → a pinned GGUF → the persona layer
cat > Modelfile <<'MF'
FROM /home/mindx/mindX/data/godel/ascend/gen39/out/runs/<recipe>/ollama_push/merged
MF
bankml create mindx-gen39 -f Modelfile            # converts (byte-identical to llama.cpp) → mindx-gen39-F16.gguf + its FORK.json
cat > Modelfile.persona <<'MF'
FROM mindx-gen39
SYSTEM """You are mindX generation 39, …"""
PARAMETER stop <|im_end|>
PARAMETER num_ctx 2048
MF
bankml create mindx-gen39 -f Modelfile.persona    # re-created in place: the layer over the same pinned base
# the same over HTTP (bankml serve … --native --registry)
curl -s $B/api/create -H "$J" -d '{"model": "mindx-persona", "from": "mindx-gen39", "system": "You are mindX.", "parameters": {"temperature": 0.7}}'
curl -s $B/api/show   -H "$J" -d '{"model": "mindx-persona"}'     # the Modelfile back, details.parent_model
curl -s $B/api/copy   -H "$J" -d '{"source": "mindx-persona", "destination": "mindx-persona-b"}'
curl -s -X DELETE $B/api/delete -d '{"model": "mindx-persona-b"}'
# conversion alone, with the FORK.json that pins it
bankml convert DIR -o mindx-gen39-F16.gguf --fork mindx-gen39-F16.gguf.FORK.json --source "PYTHAI/mindXascension@4bd31b9d weights/gen39"
```

- **Modelfile.** Ollama's grammar: `"""…"""` blocks, `#` comments, case-insensitive instructions.
  - **Taken:** `FROM`, `SYSTEM`, `PARAMETER` (`temperature`, `top_k`, `top_p`, `min_p`, `seed`, `num_ctx`,
    `num_predict`, `stop`), `MESSAGE`, `LICENSE`, `REQUIRES`.
  - **`TEMPLATE`:** only when it is the base's own.
  - **Refused, with the reason:** `ADAPTER` (merge the LoRA first), other parameters, `quantize`, uploaded `files`.
  - A relative `FROM ./x` is relative to the Modelfile.
- **Where.** `--registry DIR` (default `$BANKML_FORKS`, else `~/.local/share/bankml/forks`).
  - `FROM` a name looks there, and in `--models DIR`.
  - `FROM` a GGUF needs its FORK.json there.
  - `serve --native --registry` lists the derived models and answers `/api/create`. Without `--registry` it refuses,
    because it has nowhere to write.
- **A derived model's name** wins over the suffix-less alias of a pin (`mindx-gen39` is the layer, `mindx-gen39-f16`
  the pin), and `/api/ps` names it, with its digest and parent, as Ollama names the model that loaded its runner.
  A layer's `PARAMETER num_ctx` fits the conversation on `/api` and `/v1` as the request's option does.
- **`bankml convert`.** The Llama architecture (SmolLM2, `mindx-genN`) to GGUF F16, byte-identical to llama.cpp b11192's
  `convert_hf_to_gguf.py --outtype f16`: `e9aba089…` for SmolLM2-135M-Instruct, `6b64c748…` for gen39.
  - llama.cpp names the model from the **directory's name**. A directory called `merged` gives `general.name`
    `Merged`, as llama.cpp's would. `--model-name` sets it, as llama.cpp's flag does.
  - A `README.md` is refused, because llama.cpp would copy the card's metadata in.
  - A tokenizer that llama.cpp has not been seen to name is refused (`testing/convert_oracle.py --chkhsh DIR`
    measures one).

## 6b. The C API

Since 0.3.2 bankML is also a C library, for a program that embeds it rather than calling a server. The full
reference, with the ownership and threading rules and the log formatter, is **[CAPI.md](CAPI.md)**.

```sh
cargo build --release -p bankml-capi       # target/release/libbankml.so, libbankml.a
cc app.c -I capi/include -L target/release -lbankml -Wl,-rpath,$PWD/target/release -o app
```

```c
#include <bankml.h>
static int piece(const char *p, size_t n, void *u) { fwrite(p, 1, n, stdout); return 1; }
int main(void) {
  char *err = NULL, *result = NULL;
  bankml_t *h = bankml_open(".models/Bonsai-8B-Q1_0.gguf", "Bonsai-8B-Q1_0.gguf.FORK.json", 2048, &err);
  if (!h) { fprintf(stderr, "%s\n", err); bankml_free(err); return 2; }   /* refused: the reason, as serve gives it */
  bankml_chat(h, "{\"messages\": [{\"role\": \"user\", \"content\": \"Say hello.\"}], \"max_tokens\": 32}", piece, NULL, &result);
  puts(result);                         /* the /v1/chat/completions object, with usage and bankml_receipt */
  bankml_free(result);
  bankml_close(h);
}
```

`bankml_open` runs the same verification as `bankml serve`, and `bankml_chat` gives the answer and receipt that
`serve --native` gives (the gate's C API oracle compares them, turn by turn). `bankml_set_log` and the printf-style
`bankml_log` send the library's messages to your own callback.

## 6c. The bankML console: bankML as itself (0.3.7)

In the next release (not yet tagged). The design is in [modules/console.md](modules/console.md) and
[modules/metrics.md](modules/metrics.md).

```sh
python3 sAGI/console.py          # http://127.0.0.1:7875 — talks to bankml serve at BANKML_SERVE_LISTEN (127.0.0.1:18093)
```

`--port` moves it; any `--host` other than loopback is refused, because the console can restart the engine.
`./install.sh stop` does not stop it (it stops 7874, 7873, 18093 and 18092); stop it with Ctrl-C.

A simple page with four tabs:

- **Interaction**: ask; the answer streams, and under it the receipt — `✓ answer = receipt` when the sha256 of the text
  you received is the receipt's, the model, the tokens, the time to first token, the generation speed. bankML answers as
  itself (`sAGI/personas/bankml.persona`): each question carries a SELF block measured at that moment from
  `/bankml/usage` and `/bankml/metrics`, so "how many tokens have you used, how fast, how much power" is answered from
  measurement, and "not measured" where it was not.
- **Admin**: three sliders — CPU threads, the RAM budget (weights plus KV cache, which sets the context) and the GPU
  limit (0 = off) — applied as one verified restart of the engine, rolled back if it fails; and D3 charts of what bankML
  measures: tokens per second, TTFT, CPU, memory, GPU busy against its limit, power.
- **Logging**: every exchange with its timings and receipt, refusals included, and the engine's log.
- **Infotags**: the metadata an iNFT publication of the session carries (ERC-721 shape): the engine, the model and its
  sha256, the persona and its doctrine root, the token totals, and an RFC 6962 Merkle root over the exchanges with their
  CID. The exchanges stay on this machine; the root lets a holder check any one of them. Download it as JSON.

Power appears only after `./install.sh power` (opt-in, the one step that uses sudo; see
[install.md §2](install.md#the-steps)); until then it reads "not measured".

## 7. Let others watch (view mode, on the LAN)

```sh
./install.sh start --view          # or: python3 sAGI/view.py --host 0.0.0.0 --port 7874
```

Others open `http://<this computer's LAN address>:7874` (`ip -4 addr` shows it). The page is read-only: the live
testing log, this machine's load, the verified model, the release records, CI, Savante's office and ledger, and her
recordings (see [playback.md](playback.md)). It never shows the content of `.history` or `.memory`, only their
commitments (§8a). It is the Python standard library, not Gradio, because the Gradio installed for the chat has
path-traversal bugs and must stay on this computer.

## 8. The files Savante keeps: `.history`, `.memory`, `.prompt`

Nothing is ever written into the canon (`~/cryptoAGI/savante`). What the UI writes lives in `BANKML_UI_STATE` (default
`~/.local/share/bankml/savante/`):

- **`savante.history`**: one JSON object per exchange (JSONL):

  ```json
  {"ts": 1790633549.123, "sent_at": "2026-09-28T15:12:29.123-0700", "first_token_s": 113.02,
   "response_s": 125.11, "answered_at": "2026-09-28T15:14:34.233-0700", "session": "e2e-test",
   "user": "In one sentence: …", "assistant": "Savante knows …", "assistant_raw": "Savante knows …", "shown": "… (with the footer)",
   "prompt": "persona · system_prompt (canon, ledgered)", "prompt_provenance": "persona.system_prompt · persona sha256 d96556b11989… (ledgered)",
   "receipt": {"bankml": "0.0.7", "model_sha256": "284a335a…", "prompt_tokens": 316, "completion_tokens": 28,
               "ttft_ms": 113000, "wall_ms": 125100, "response_sha256": "80010408…"}}
  ```

  `sent_at` is the press of Send; `first_token_s` and `response_s` are measured from it. The receipt's
  `ttft_ms`/`wall_ms` are measured by bankml serve from when it forwarded the request. `assistant_raw` (0.1.7+) is
  the text exactly as the model wrote it, which is what the receipt hashes; `assistant` is the same text with any
  `<think>` block and surrounding whitespace removed. Each record is appended as one locked, fsynced write; a line
  damaged by a crash is skipped (and counted), never allowed to hide the rest. The history is plain text: back it
  up, grep it, or delete it.
- **`savante.memory`**: one JSON object per note: `{"ts", "at", "text", "sha256", "source": {"kind": "typed" |
  "response", "session", "sent_at", "response_sha256"}}`.
- **`Savante.prompt`**: the Space template's prompt, cached the first time it is chosen.

**The history the model sees.**
- **Without a context limit:** the last 12–17 exchanges, each cut to 4,000 characters. The window starts at 12
  exchanges and moves in steps of six, so between moves each prompt is the previous prompt plus one exchange, and
  the engine's prompt cache reuses all of it.
- **Within the engine's context.** The window must also fit the context the engine actually runs with (`n_ctx` from
  `/props`), counted by its own tokenizer.
  - The start stays where it was last turn while that still fits, so the cache is reused.
  - When it no longer fits, the start jumps so that the newest half of what fits is kept, leaving room for the next
    turns. When only two or three exchanges fit, all of them are kept.
  - The answer's trail (Admin → *the last answer*) then says so: "history: 4 of 13 exchanges fit the engine's
    2048-token context — raise the RAM budget on the Admin tab for more".
  - Exchanges older than the 12–17 window are dropped without a note, as they always were.
- **If the system prompt and the question alone do not fit,** the answer is refused with the numbers.
- On a 2048-token context the persona prompt plus `.memory` leaves room for only a few exchanges; 4096 tokens (about
  0.3 GB more RAM for the 8B model) gives the window room to stay put for several turns. **After an engine
restart**, the first question restores the saved KV of the system prompt (a 51 MB file per model, context and system
prompt, in `~/.local/share/bankml/savante/slots/`), so it skips the system prompt's prefill: on this laptop 132 s
became 15 s, with the same answer.
(A window that slid by one exchange per turn, as before 0.1.8, changed the text right after the system prompt and
forced the whole history to be re-read every turn; over 60 turns the stable window moves 8 times instead of 48.) The
chat's footer (clock, receipt) is never sent to the model.

## 8a. Proof of data without the data

`.history` and `.memory` stay on this computer. What can leave it is a **commitment**:

| part | what |
|---|---|
| leaf | sha256(0x00 ‖ one JSONL line, exactly as written, without its newline) |
| Merkle root | RFC 6962 / RFC 9162: node = sha256(0x01 ‖ left ‖ right), the tree split at the largest power of two below n (an odd node is promoted, never paired with itself). Scheme `rfc6962-sha256`; 0.1.6 and earlier used an unprefixed tree, so their roots differ |
| file sha256 / CIDv1 | of the whole file; CIDv1 raw (0x55), sha2-256, base32 lower: the construction Savante's iNFT ledger uses |

**🔏 proof** (Responses tab) produces, for one exchange, `{"commitment": …, "proof": {"scheme", "record", "records",
"leaf", "path": [hash, …], "merkle_root"}, "verifies": true}`. Give someone the exchange's line and the proof: they
verify it as RFC 9162 §2.1.3.2 describes (the index and the tree size decide at each step which side the sibling is
on) and compare the result with the root **and record count** you published; the pair is the commitment. That proves the exchange is in your history, unchanged,
and reveals nothing else. The same works for any data kept in a browser's local storage: keep the data, publish the
hash or CID. Metrics and the view page show the commitments; the view page never serves a line.

To check a proof by hand, in Python:

```python
import hashlib, json
line = b'...the exact JSONL line...'
p = json.load(open("proof.json"))["proof"]
h = hashlib.sha256(line).hexdigest()
for s in p["path"]:
    a, b = (h, s["hash"]) if s["side"] == "right" else (s["hash"], h)
    h = hashlib.sha256(bytes.fromhex(a) + bytes.fromhex(b)).hexdigest()
print(h == p["merkle_root"])
```

## 8b. Custom agents from the Savante template

**Agents** tab. Savante's canon is the template and is never written. **Derive a new agent** creates, in
`~/.local/share/bankml/agents/<name>/` (`BANKML_AGENTS`):

| file | what |
|---|---|
| `<name>.persona` | the persona (mindX `.persona v1` shape): a new name and identity, Savante's BDI/skills/safety as a starting point; the `token` bindings start empty (Savante's are never copied) |
| `<name>.prompt` | the system prompt, as text: the `.prompt` the chat uses |
| `<name>.agentcard.json` | EIP-721 metadata ∪ ERC-8004 registration-v1; status `not_yet_minted`; `derived_from` names Savante's persona sha256 and doctrine root |
| `<name>.commitments.json` | the ledger: sha256 + CIDv1 of each file, and the **doctrine root** (keccak256 over the same 15 JSON pointers as Savante's ledger) |
| `<name>.history`, `<name>.memory` | the agent's own conversations and notes |

**Use this agent** switches the chat, `.history`, `.memory`, Responses and Metrics to it. **Edit** (custom agents
only) opens the `.persona` as JSON and the `.prompt` as text. **Save and re-ledger** refuses a persona that breaks
the binder's preflight (non-ASCII keys, floats, integers beyond ±2⁵³), lacks a doctrine clause, or **changes one**:
the 15 doctrine clauses (name, source, format, system prompt, mantra, oath, beliefs, skills, safety, embodiment, tool
allowlist …) are fixed when the agent is derived, and their root is recorded then and carried forward. To change them,
derive a new agent. Everything else (the `.prompt` file, voice examples, description, the aivatar) stays editable.
Otherwise it regenerates the card and ledger. An agent whose files do not match its ledger refuses to speak.

The keccak256 is written in pure Python. It reproduces Savante's published doctrine root
(`0x92fe83eb…ae137d0`) from her persona, and it equals pycryptodome's keccak for every input length from 0 to 400
bytes (tested).

## 8c. THOT bundles: the dataset an iNFT points to

**Agents → THOT bundle** builds `<name>.thot.json` to `sagi.thot_manifest/1` (`~/sagi/engine/THOT_MANIFEST.md`):

- **Facets.** Core facets `persona` (the ternary head, leaf 0) and `prompt`, plus declared custom facets
  `x-bankml.agentcard`, `x-bankml.history` and `x-bankml.memory`. The history and memory are committed **by digest
  only**: the manifest holds no conversation text and can be published.
- **Bundle root.** `bundle_root` = keccak256 over `label ‖ 0x1f ‖ sha256_hex ‖ 0x1e`, core facets in registry order,
  then custom facets by name.
- **Merkle root.** A 64-leaf keccak256 tree, padded with `keccak256(b"")`, with unsorted pairs and no prefixes.
- **Identity.** `thot:<sha256>`, a CIDv1, `thot-<cid>`, and `contentRoot` (keccak256), all over the canonical
  manifest without its identity block.
- **Lineage.** Genesis is generation 1. Every facet change (a new chat, a new note, an edited prompt) makes the next
  build generation n+1, with `parent` = the previous manifest's CID. A re-bind with no change keeps the generation.
  `relations` records `derived_from` Savante's bundle.
- **Rung.** Each agent directory is its own git repository, so the rung's locator
  (`localhost/<user>/<name>@<commit>`) names real bytes. History and memory are git-ignored and listed as
  `locator_lacks`.

The builder is checked against the spec's own test vectors: Savante @1fcca89, Jaimla @8b57ccf and LuvAI @0c1eef7 give
the same bundle roots, Merkle roots and identity CIDs. Savante's current generation-9 manifest verifies with no findings.

## 8d. PostgreSQL: publish an agent, load one back

**Agents → PostgreSQL**. The connection is `BANKML_PG_DSN` (default `dbname=bankml`, the local socket, your role).
One-time setup, which needs sudo:

```sh
sudo -u postgres psql -c "CREATE ROLE $USER LOGIN CREATEDB" -c "CREATE DATABASE bankml OWNER $USER"
sudo -u postgres psql -d bankml -c "CREATE EXTENSION IF NOT EXISTS vector"
```

- **Publish the agent in use.** This rebuilds and verifies its THOT bundle, then upserts `bankml_agents`: persona
  (the exact bytes, plus a jsonb copy for queries), prompt, card, ledger, manifest, THOT CID, contentRoot and
  generation.
  - **Off by default:** `.history` and `.memory` lines go in (`bankml_exchanges`, `bankml_memory`, exact line text,
    sequence and sha256) only when *include .history and .memory lines* is ticked. Otherwise the database holds only
    their commitments, inside the manifest.
- **Load.** Rebuilds a published agent as a local one. It is **refused** unless the stored manifest's identity
  recomputes and every restored file (persona, prompt, and history and memory when included) re-hashes to the
  manifest's digest. A tampered row cannot be loaded.
- **Vectors.** `bankml_exchanges.embedding` is `vector(1024)` (bge-m3's width, as mindX uses), indexed with DiskANN
  when pgvectorscale is installed and HNSW (pgvector) otherwise. It is filled with bge-m3 vectors when private lines are
  published and bge-m3 can run (see [embedding.md](embedding.md)); otherwise it stays empty.
- **Safety.** Values travel to psql as COPY data into a temporary table; no value is ever part of SQL text.
  `testing/test_connectors.py` runs every step against a throwaway cluster (initdb in a temp dir, pgvector, its own
  socket), including a tampered prompt, a tampered history line and an SQL-injection string, all handled.

## 8e. iNFT: mint an agent, load one from a token

**Agents → iNFT**, for the agent in use. Savante herself is not minted from here: her verdict on minting is DEFER.

The contract is the house ERC-7857 `iNFT_7857` (DeltaVerse iNFT4), and an open (unsealed) agent is minted with:

```
mintOpenAgent(address to, bytes32 contentRoot, string storageURI, bytes32 metadataRoot,
              uint256 dimensions, uint8 parallelUnits, string tokenURI)        — MINTER_ROLE only
```

| argument | from the agent |
|---|---|
| `contentRoot` | its THOT manifest's `identity.contentRoot` (keccak256 of the canonical manifest). The contract accepts a content root once, ever |
| `metadataRoot` | keccak256 of the agent card's canonical bytes |
| `storageURI`, `tokenURI` | `local://thot/<cid>` and `local://card/<cid>` until the bundle and card are stored somewhere (the rung stays `referenced`) |
| `dimensions`, `parallelUnits` | one of 8 … 1048576 (768 by default), and 1 |

- **Plan and simulate** builds the calldata and runs it as an `eth_call` from the minter. You get the token id it
  would receive, or the decoded revert (`AccessControlUnauthorizedAccount`, `ContentRootAlreadyMinted`,
  `InvalidDimension`, …).
- **Unsigned transaction** gives the transaction (chain id, to, data, gas estimate) and the same call as a
  `cast send … --account <your keystore>` line. **You sign it**, in your own wallet: bankml never signs on a
  public chain.
- **Mint on the local devnet** sends only when the chain id is 31337 (anvil). On any other chain it refuses.
- **Load the agent from this token** reads `getPayload`, `tokenURI`, `ownerOf` and `openMint`, finds the agent whose
  THOT generation has that `contentRoot` (every manifest is archived by CID in `<agent>/thot/`), verifies its
  bundle, and walks the THOT parent links from the current generation back to the minted one. The genesis stays
  attached to the token forever; later generations are proven to descend from it.
- **Start a local devnet** runs anvil (127.0.0.1:8545, chain 31337), deploys `iNFT_7857` from its compiled artifact
  (`DeltaVerse/deploy/iNFT4/out`) with account 0 as admin and minter, and fills in the fields. The whole path then
  runs on this computer.

Where it stands: `iNFT_7857` is **not deployed on any public chain**, its audit (`iNFT4/audit.json`) is not cleared,
and a real mint needs a wallet holding MINTER_ROLE. `testing/test_chain.py` runs the whole path on a throwaway
anvil: deploy, simulate (and the refusal without the role), unsigned transaction, mint, read-back, the double-mint
refusal, load, two more generations and load again, and the refusal to send off a devnet.

## 9. Receipts, and how to check an answer

Every answer from `bankml serve` carries:

| field | meaning |
|---|---|
| `bankml` | the version that served it |
| `engine` | what computed it: `llama.cpp b11192 llama-server (loopback), behind bankml P0`, or with `--native` `bankML <version> native: its own forward pass, …` (and the GPU when one took part) |
| `model_sha256`, `guard` | the file that was verified and pinned, and the guard's verdict |
| `prompt_tokens`, `completion_tokens` | the engine's own counts |
| `ttft_ms`, `wall_ms` | time to first token and total, at the gateway (`ttft_ms` is `null` on a non-streamed answer in 0.3.6; the next release sets it there too) |
| `response_sha256` | sha256 of the answer text exactly as the model produced it (`choices[0]`, UTF-8; an unpaired `\u` surrogate counts as U+FFFD) |
| `request_sha256` | (0.1.7+) sha256 of the request body bankml forwarded: which prompt this answer is to |
| `signed` | always `false`: a receipt is not a signature (below) |

The UI recomputes the answer's sha256 and shows ✓ when it matches. A mismatch shows `(≠ received!)`. To check an
answer yourself, hash `assistant_raw` from `.history` (older records: `assistant`) and compare it with
`receipt.response_sha256`.

**What a receipt proves, and what it does not.** It binds an answer to the request it answered and to the sha256 of
the model file bankml verified before serving, and it catches any change to the text afterwards. It is **not
signed**: anyone can write a receipt, so it shows integrity between you and your own bankml serve, not to a third
party who does not trust your machine. bankml serve also re-checks, before each answer, that the model file is the
one it verified (device, inode, size and modification time) and that the engine still serves that file; if either
changed, it refuses to answer rather than issue a receipt for weights it did not verify.

## 10. Savante's canon and the iNFT ledger

Savante is defined by her canon, `~/cryptoAGI/savante` (github.com/cryptoAGI/savante), bound for an iNFT by the ledger
`savante.commitments.json`. The ledger commits by sha256 (and CIDv1) to:
- the persona;
- the charter;
- the sagi skill;
- six `sAGI.*` facets;
- the agent card;
- the image;
- the thot bundle;
- a keccak256 **doctrine root** over 15 clauses an owner may not edit.

At start the UI re-hashes every one of those files and shows the result in **Integrity**. If the persona does not
verify, the UI refuses to speak as Savante. The **Verifier** tab runs the full offline check. Nothing here mints: the
card's status is `not_yet_minted`, and the decision to mint belongs to the owner's signature. Point the UI at another
checkout with `SAVANTE_CANON=/path`.

## 11. Testing and the release gate

```sh
cargo test --release                                             # offline
BANKML_GGML_LIB=/path/to/llama-b11192 testing/release_gate.sh    # everything → testing/results/<version>.txt
testing/live.sh "title" <command…>                               # one step, shown live in view mode
```

The gate runs, in order:
1. the build, the unit tests and the end-to-end CLI tests;
2. clippy (`-D warnings`);
3. the licence headers (`testing/spdx_check.py`);
4. the Python suites: the guard, the UI data layer, the PostgreSQL connector (a throwaway cluster), the iNFT path (a
   throwaway anvil devnet), the model importer (with a real carrier on spare ports);
5. the Rust and Python guards agreeing on every case; the C API's printf oracle (`bankml_log` against libc
   `snprintf`, from C);
6. every oracle: bankml's kernels must equal llama.cpp's compiled kernels, bit for bit, on every weight of the real
   models;
7. the A/B speed tests, the prefill tile, the memory floor and the whole-token budgets;
8. the conversation oracles: `serve --native` through `/api/chat` and `/v1`, and the C API's `bankml_chat` against
   `serve --native` (ternary) and llama-server's record (1-bit);
9. the live oracles that start `serve --native` and compare it with llama-server's answers: JSON mode and schemas, the
   persona layer, the penalties, and in the next release the samplers, the context limit, slots, sessions, logprobs
   and the q8_0 KV cache.

A speed counts only if every oracle passed on the same code. See [testing/README.md](../testing/README.md), and
[oracles.md](oracles.md) for what each oracle compares.

## 12. Troubleshooting

| symptom | cause | fix |
|---|---|---|
| chat says *Cannot reach bankml serve* | serve not running | `./install.sh status`, then `./install.sh model` (§3) |
| the installer stops at a step | that step's prerequisite is missing | it says which; fix it and run `./install.sh` again (steps already done are quick) |
| serve prints `refuse: … != pinned` | the file is not the one the fork pinned | re-download; `bankml sha256 FILE` |
| serve prints `upstream … serves X, not the verified Y` | the llama-server on that port runs another file | restart it with the verified file, or use `--spawn` |
| serve prints `upstream … not healthy` | llama-server not up (a cold load can take a minute) | wait; check its log |
| the first answer takes minutes | CPU prefill of the ~300-token system prompt (≈3 tok/s on a laptop) | expected; later turns reuse the cache |
| *refused: savante.persona does not verify* | the canon was edited or is incomplete | `git -C ~/cryptoAGI/savante status`; run the Verifier |
| view mode says *swap full* | the machine is under memory pressure | close other work before measuring |
| view from another computer does not load | firewall, or bound to 127.0.0.1 | `--host 0.0.0.0`; allow TCP 7874 |

## 13. Reference: commands, ports, environment

| command | what |
|---|---|
| `./install.sh [step …] [--skip-tests] [--no-start] [--view] [--voice]` | install and start (§1); steps: check build engine python canon model voice start stop status, and (0.3.7, opt-in, sudo) `power [--remove]` ([install.md §2](install.md#the-steps)) |
| `bankml guard FILE [--engine mainline\|prism] [--json]` | header check: play / refuse / need_more |
| `bankml sha256 FILE` | the file's sha256 |
| `bankml pin FILE --fork FORK.json` | sha256 against the fork's record |
| `bankml verify FILE --fork FORK.json [--engine mainline\|prism] [--json]` | guard, then pin |
| `bankml serve FILE --fork FORK.json [--upstream H:P \| --spawn BIN] [--listen H:P] [--threads N] [--ctx N] [--spec-ngram] [--slot-dir DIR]` | the gate in front of llama-server (n-gram speculation opt-in; `--slot-dir` becomes llama-server's `--slot-save-path`) |
| `bankml chat-template MODEL.gguf < messages.json` | the prompt a conversation becomes, as llama.cpp's `/apply-template` (P3; byte-identical on its oracle; tools and assistant prefills refused) |
| `BANKML_GPU=off` · `BANKML_GPU_SHARE=0.3` | the GPU worker (0.2.14): a verified card takes a calibrated share of every 1-bit matrix's rows; `off` disables it, a number overrides the share |
| `BANKML_GPU_LIMIT=0.8` | the GPU limiter (0.3.7, next release): bankML's GPU buffers stay within this share of the card's heap (of the RAM they could use, on an integrated card) and the card rests so it is busy at most this share of the time; what does not fit or arrives while it rests runs on the CPU, the same bits. Each matrix shape also decides, measured in the pipeline, whether the card's share pays (on the Vega 3 most shapes choose the CPU) |
| `bankml serve FILE --fork FORK.json --native [--listen H:P] [--upstream H:P] [--ctx N] [--registry [DIR]] [--keep-alive DUR] [--slot-dir DIR]` | answers from bankML's own forward pass (0.3.0); since 0.3.1 also Ollama's API (§6a) and, with `--registry`, any pinned model by name, one resident at a time; token-identical to llama-server on its oracle; also serves llama-server's endpoints on the engine address, so Savante reaches it unchanged. `--slot-dir` (0.3.8, next release): llama-server's slot save/restore (§6). Savante's Admin tab **engine** setting (under *advanced*) chooses it: `auto` (bankML for the ternary files), `native`, `llama.cpp` |
| `bankml create NAME -f Modelfile [--registry DIR] [--models DIR]` | O5 (0.3.5): a derived model, a layer (SYSTEM, PARAMETER, stop, MESSAGE, LICENSE) over a pinned base, `FROM` a name, a pinned GGUF or a safetensors directory (converted); written as `DIR/NAME.MODEL.json`, no weights copied (§6a) |
| `bankml convert DIR -o OUT.gguf [--model-name N] [--fork F --source S] [--ignore-model-card]` | O5 (0.3.5): Llama safetensors → GGUF F16, byte-identical to llama.cpp b11192's `convert_hf_to_gguf.py --outtype f16`; `--fork` writes the pin with every input's sha256 |
| `bankml gpu --verify` | runs the bit-exact kernel oracle on every usable card (0.2.13); a card that fails is named and never used |
| `bankml gpu [--remote]` | every video card found (Vulkan, merged with `/sys/class/drm`) and which bankml will use; `--remote` adds the GPUs Hugging Face rents (22 NVIDIA flavors with card counts and prices; listed, never started); `BANKML_GPU=off` turns the component off, `BANKML_GPU=0,2` picks cards (0.2.12; the GPU kernels are the next steps) |
| `bankml generate MODEL.gguf [--max N] [--sample [--temp T] [--top-k K] [--top-p P] [--min-p P] [--seed S]] < messages.json` (or plain text) | bankml's own forward pass, greedy, or with `--sample` llama-server's sampler chain (the model's defaults unless given; same seed, same tokens as llama-server, 0.2.11), streamed (P3, 0.2.7): token-identical to llama-server b11192 on its oracle for the 1-bit and ternary models, prompts of any length and contexts of any length (all three of ggml's CPU attention kernels, since 0.2.10); `BANKML_LLAMA_THREADS` (default 3) must equal the `-t` of the llama.cpp being matched, because its long-context decode kernel chunks by thread; `BANKML_THREADS` sets the threads. Ternary: 2.3–2.4 tokens/s against llama-server's 0.30; 1-bit: 1.8 against 2.8 |
| `bankml tokenize MODEL.gguf [--no-special] < text` | token ids, as llama.cpp's `/tokenize` (P3's tokenizer; token-identical on its oracle) |
| `bankml usage [PID …]` | memory, cores, and each process's resident memory and CPU % (bankml's psutil, from `/proc`); `bankml serve` answers the same at `GET /bankml/usage`, with (0.3.7, next release) package watts when RAPL is readable (`./install.sh power`), each GPU's busy %, VRAM and GTT, and the GPU limiter's state |
| `GET /bankml/metrics` | (0.3.7, next release) bankML's own measurements of its last 256 answers: TTFT, prompt and generation tokens per second, grammar time, energy and joules per token where measured; totals |
| `python3 sAGI/savante.py --mode interact [--port 7873]` | talk to Savante (loopback) |
| `python3 sAGI/view.py [--host 0.0.0.0] [--port 7874]` | the read-only page for the LAN |
| `python3 sAGI/console.py [--port 7875]` | (0.3.7, next release) the bankML console, loopback only (§6c) |
| `python3 sAGI/models.py list \| catalog \| search Q \| import ID\|URL\|ollama:NAME:TAG \| use FILE \| first-run` | the model importer and carrier switch |
| `python3 sAGI/speak.py [--prune]` | render every voice clip, write both exports (`--prune`: drop unused clips) |

| port | service |
|---|---|
| 18092 | llama-server b11192 (upstream, loopback) |
| 18093 | bankml serve (loopback) |
| 7873 | interact mode (loopback) |
| 7874 | view mode (LAN) |
| 7875 | the bankML console (loopback; 0.3.7, next release) |

The variables below are the ones a user meets most. The complete list, with where each is read, is in
[install.md §6](install.md#6-environment-variables-the-complete-list).

| variable | default | meaning |
|---|---|---|
| `SAVANTE_CANON` | `~/cryptoAGI/savante` | Savante's canon (read-only) |
| `BANKML_SERVE` | `http://127.0.0.1:18093` | where the UI finds bankml serve |
| `BANKML_UI_STATE` | `~/.local/share/bankml/savante` | `.history` and caches |
| `BANKML_REPO` | the checkout | where view mode reads `testing/` |
| `BANKML_AGENTS` | `~/.local/share/bankml/agents` | custom agents (one git repository each) |
| `BANKML_PG_DSN` | `dbname=bankml` | PostgreSQL for publishing and loading agents |
| `RAGE_PATH` | `~/mindX/mindx/godel/mindxtrain/hf/space_ui` | where the ragebar finds mindX's `rage.py` (built-in BM25 otherwise) |
| `BANKML_GGML_LIB` | — | llama.cpp b11192 release dir, for the oracles |
| `BANKML_THREADS` | pool: all cores; budgets: `1,2,3,4` (Q1_0), `1,3` (Q2_0) | the thread pool's size, or a comma list of thread counts for the decode budgets |
| `BANKML_NO_SHANI` | unset | set to hash with the portable SHA-256 instead of the CPU's SHA extensions |
| `BANKML_OLLAMA_MODELS` | `/usr/share/ollama/.ollama/models` | the local Ollama store the importer adopts from |
| `BANKML_PIPER`, `BANKML_PRONUNCIATION` | `~/.local/share/bankml/piper`, the house table | Savante's voice engine and pronunciation table |
| `BANKML_ESPEAK`, `BANKML_ESPEAK_VOICES` | the DeltaVerse vendor dirs | the eSpeak fallback voice |
| `BANKML_OLLAMA`, `BANKML_EMBED_MODEL`, `BANKML_EMBED_KEEP_ALIVE`, `BANKML_EMBED_NEED_GB` | see [embedding.md](embedding.md) | the embedding model |
| `BANKML_MODELS` | `.models` in the checkout | where imported models go |
| `BANKML_FORKS` | `~/.local/share/bankml/forks` | FORK.json pins, one per imported model |
| `BANKML_LLAMA_SERVER` | `~/sAGI/bonsai/llama-b11192/llama-server` | the engine the carrier spawns |
| `BANKML_BIN` | `target/release/bankml` | the bankml binary the importer and switch call |
| `BANKML_FIRST_RUN` | `1` | `0` stops interact from starting the Bonsai-8B carrier by itself |
| `BANKML_CTX`, `BANKML_THREADS_SERVE` | `2048`, `3` | the engine's context and threads when the importer starts the carrier |
| `BANKML_SERVE_LISTEN`, `BANKML_UPSTREAM` | `127.0.0.1:18093`, `127.0.0.1:18092` | the ports the switch manages |
| `BANKML_VOICE_DIR`, `BANKML_EXPORT_DIR` | `sAGI/voice/cache`, `sAGI/voice/export` | voice clips and the two exports |
| `BANKML_VOICE_ASYNC` | `1` | `0` stops the UI rendering missing clips in the background |
| `BANKML_CACHE_RAM` | 8192 MiB, at most a quarter of the memory available at load | (0.3.8, next release) the native engine's host prompt cache in MiB, as llama-server's `--cache-ram` (0 off, -1 no limit) |
| `BANKML_CACHE_TYPE` | `f16` | (0.3.9, in progress) `q8_0` keeps the native KV cache as llama.cpp's `--cache-type-k/v q8_0` (§6) |