bankml / docs /modules /create.md
Gregory-L's picture
bankML: the whole source (github.com/cryptoAGI/bankml @ 12ae409) and its page, with the bankML persona; the live engine (Dockerfile, hf/start.sh) ready for Docker hardware
28c70af verified
|
Raw History Blame Contribute Delete
10.7 kB

bankML/create.rs — bankml create: derived models over pinned bases

Summary

create.rs is Ollama's ollama create over bankML's gate (O5, 0.3.5). A derived model is a manifest that layers configuration on a pinned base: a system prompt, parameters, stop strings, example messages, a licence. It never copies weights. Loading one verifies the base exactly as a pinned model is verified (the guard, then the sha256 pin of its FORK.json), then applies the layer.

It exists for mindX's flow: mindXtrain merges each generation into safetensors, and mindX's promote.py layers a persona with ollama create. bankml create takes the same Modelfile. FROM a safetensors directory converts it with convert.rs and pins the result, so the persona becomes a verified layer over a pinned GGUF.

The module holds Ollama's Modelfile parser, the manifest format, the derived-model store, the layer's application to Ollama and OpenAI requests, the HTTP handlers for /api/create, /api/delete and /api/copy, and the CLI.

Technical usage

The manifest

<registry>/<name>.MODEL.json   {kind, bankml, name, created_at, base: {name, file, sha256, fork, path},
                                layer: {base_sha256, system, parameters, stop, template_sha256, license, messages, requires},
                                digest: "sha256:…"}

digest is the sha256 of the layer's content (the base's sha256 and the layer), not of the name or the time. The same Modelfile over the same base gives the same digest, and a copy keeps it. A manifest whose digest is not its content's is refused when read: "edited by hand, refused — create it again". A manifest whose base pin has changed is refused too. It is written atomically (a .json.part file, then a rename).

The Modelfile (parse_modelfile, Spec::from_modelfile)

The parser is Ollama's parser/parser.go, state for state: case-insensitive instructions; # opens a comment only at a line's start; a value runs to the end of the line; "…" and """…""" may span lines, with no escapes. The vendored reference for the format is mindX's docs/ollama/setup/modelfile.md. quote() writes values back as Ollama's ollama show --modelfile does, and they round-trip.

instruction handling
FROM exactly one: a registry name (pinned or derived), its suffix-less alias, a pinned GGUF path, or a safetensors directory (converted to <name>-F16.gguf and pinned with every input's sha256)
SYSTEM, TEMPLATE the last wins; TEMPLATE only when it is the base's own chat template (bankML renders that template byte-identically to llama.cpp; a different one would change every prompt)
PARAMETER one of PARAMS (below) or stop; the last value wins, stop accumulates
MESSAGE role content role system, user or assistant; recorded and applied, as Ollama applies them
LICENSE, REQUIRES recorded
ADAPTER refused

FROM a derived model inherits its layer: new values win, stop is replaced as a list, licences add up. A relative FROM ./x on the CLI is relative to the Modelfile.

PARAMS

pub const PARAMS: [(&str, bool); 13] = [("temperature", false), ("top_k", true), ("top_p", false), ("min_p", false),
    ("seed", true), ("num_ctx", true), ("num_predict", true), ("repeat_penalty", false), ("repeat_last_n", true),
    ("presence_penalty", false), ("frequency_penalty", false), ("typical_p", false), ("min_keep", true)];

The bool says whether the value is an integer; the order is the order a manifest keeps. The four penalties joined in 0.3.6 (O2); typical_p and its min_keep in 0.3.7, when the engine came to reproduce typical-p. Floats are parsed at 32 bits, as Ollama parses them. Range checks at create: num_ctx ≥ 1; num_predict ≥ −2; repeat_penalty > 0; frequency_penalty and presence_penalty any finite value (negative too, as llama.cpp takes them); top_k and seed unchecked; the rest ≥ 0. A PARAMETER outside PARAMS is refused with one of three reasons: a sampler not reproduced (tfs_z, mirostat, mirostat_eta, mirostat_tau, penalize_newline); a resource option that belongs to bankml serve; or not a parameter bankML knows.

Applying the layer (Ollama's server/routes.go)

  • apply_ollama(&self, req, chat): the layer's parameters become defaults under the request's options, overridden key by key (stop as a whole list). For chat, the model's MESSAGEs go before the request's, and its SYSTEM goes first unless the request's first message is a system message. For a non-raw generate, the conversation goes in bankml_messages: the request's system if any, else the model's, then the MESSAGEs, then the prompt. A raw prompt gets neither.
  • apply_openai(&self, req): the same message rule (Ollama's OpenAI endpoint goes through the same chat handler); the parameters fill absent top-level fields (num_predict as max_tokens); num_ctx is not a field but fits the conversation (num_ctx()).

Resolution

resolve(rs, model) tries a pin's exact name first, then a derived model's name, then the registry's suffix-less alias (0.3.5). So promote.py's re-create in place (FROM mindx-gen39 as mindx-gen39, where mindx-gen39 is the alias of the one pin mindx-gen39-f16) answers as itself with the layer, and mindx-gen39-f16 gives the pin. FROM resolves the suffix-less alias the same way, and only when exactly one pin has it. A derived model loads through its base's entry, so it and its base share one residency.

HTTP (needs serve --native --registry)

curl -s $B/api/create -H "$J" -d '{"model": "mindx-persona", "from": "mindx-gen39", "system": "You are mindX.", "parameters": {"temperature": 0.7}}'
curl -s $B/api/create -H "$J" -d '{"model": "p", "modelfile": "FROM mindx-gen39\nPARAMETER repeat_penalty 1.3"}'
curl -s $B/api/copy   -H "$J" -d '{"source": "mindx-persona", "destination": "mindx-persona-b"}'
curl -s -X DELETE $B/api/delete -H "$J" -d '{"model": "mindx-persona-b"}'
  • /api/create: {model, modelfile} or Ollama's structured {model, from, system, template, license, parameters, messages}; status lines streamed as NDJSON unless stream: false. Without --registry: 400, nowhere to write.
  • /api/delete: derived models only; a pin is refused ("bankml does not delete pins over HTTP").
  • /api/copy: the same content and digest under a new name; a pin's copy is an empty layer over it.
  • /api/show of a derived model (show_json): the reconstructed Modelfile, parameters, system, license, messages, details.parent_model, and bankml.layer with the digest and the base's sha256 and pin.

CLI

bankml create NAME -f Modelfile [--registry DIR] [--models DIR] prints Ollama's status lines and exits 0, or bankml create: refuse: … and 2. The registry defaults to $BANKML_FORKS, else ~/.local/share/bankml/forks.

How it is verified

  • Unit tests: modelfile_parses_as_ollama (mindX's own Modelfiles, comments, CRLF, BOM, quoting, errors), quote_round_trips_through_the_parser, spec_keeps_what_is_reproduced_and_refuses_the_rest, manifest_digest_is_the_content, the_layer_applies_as_ollama_applies_it, create_layers_over_a_pinned_base, http_create_show_copy_delete_round_trip.
  • testing/cli.rs: create_writes_a_layer_over_a_pin.
  • oracle_persona_layer (#[ignore], gate; testing/persona_oracle.py): promote.py's persona Modelfile created two ways, FROM the merged safetensors directory and FROM mindx-gen39 in place. Asked the user's turns alone, each must give llama-server b11192's tokens for the same GGUF given the persona as the system message. CHANGELOG 0.3.5: 27 / 27 each. Live: persona_oracle_live through /api/chat, /api/generate, /v1 and /api/ps.

Advantages and efficiency

  • No weights copied. A layer is a small JSON file over a pinned GGUF; its digest is content-addressed, as Ollama's content addressing does. Every load of a derived model verifies its base (the guard, then the pin).
  • mindX's Modelfiles unchanged. The parser is Ollama's state machine, so promote.py's files create the same layer, and /api/show gives back a Modelfile that recreates it.
  • Efficient. Tamper detection is one sha256 over a short text. A derived model and its base share one load. A conversion never overwrites an existing pin.
  • Rust practice. No external crates; atomic manifest writes; RwLock poisoning recovered; errors carry the reason as Result<_, String>.
  • Next (docs/OLLAMA.md, docs/TODO.md): ADAPTER (LoRA merging) is a later O-phase; promote.py --to bankml in mindX and a bankml backend for mindXtrain (O5, proposed).

Limitations

  • ADAPTER: "bankML does not merge LoRA adapters yet (a later O-phase); merge it into the base weights … then FROM the merged safetensors directory".
  • TEMPLATE other than the base's own: Ollama's are Go templates; bankML renders the GGUF's Jinja byte-identically.
  • /api/create files, adapters, quantize: refused, each with its reason.
  • More than one FROM; a name a pinned file already has; a FROM file not pinned in the registry directory.
  • Model names: letters, digits, ., _, -, at most 128, no tag but latest, not ending in .gguf.
  • A layer num_ctx above the server's --ctx is accepted at create (with a warning) and refused per request.
  • The 0.3.7 samplers that are not Ollama options (DRY, XTC, top-n-σ, dynamic temperature) are not Modelfile parameters; a request gives them through /v1.

Design notes

  • bankml create is phase O5 of ../OLLAMA.md, first built in 0.3.5; since 0.3.5 a conversion under create never overwrites an existing pin, nor its FORK.json wherever the file is.
  • is_print approximates Go's strconv.IsPrint (no control characters, no space other than U+0020, no format characters), which the parser uses as Ollama's does.
  • A create changes what requests resolve to, so POST /api/create takes the engine's run lock, as a model load does.
  • Parameter floats are stored at 32 bits, as Ollama parses them, and written in the shortest form that reads back.

See also