Download docs/modules/create.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 10.7 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/create.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/modules/create.md
-
curl -L -o create.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/create.md
bankML/create.rs — bankml create: derived models over pinned bases
Summary
create.rs is Ollama's ollama create over bankML's gate (O5, 0.3.5). A derived model is a manifest that
layers configuration on a pinned base: a system prompt, parameters, stop strings, example messages, a licence.
It never copies weights. Loading one verifies the base exactly as a pinned model is verified (the guard, then the
sha256 pin of its FORK.json), then applies the layer.
It exists for mindX's flow: mindXtrain merges each generation into safetensors, and mindX's promote.py layers a
persona with ollama create. bankml create takes the same Modelfile. FROM a safetensors directory converts it
with convert.rs and pins the result, so the persona becomes a verified layer over a pinned GGUF.
The module holds Ollama's Modelfile parser, the manifest format, the derived-model store, the layer's application to
Ollama and OpenAI requests, the HTTP handlers for /api/create, /api/delete and /api/copy, and the CLI.
Technical usage
The manifest
<registry>/<name>.MODEL.json {kind, bankml, name, created_at, base: {name, file, sha256, fork, path},
layer: {base_sha256, system, parameters, stop, template_sha256, license, messages, requires},
digest: "sha256:…"}
digest is the sha256 of the layer's content (the base's sha256 and the layer), not of the name or the time. The
same Modelfile over the same base gives the same digest, and a copy keeps it. A manifest whose digest is not its
content's is refused when read: "edited by hand, refused — create it again". A manifest whose base pin has changed
is refused too. It is written atomically (a .json.part file, then a rename).
The Modelfile (parse_modelfile, Spec::from_modelfile)
The parser is Ollama's parser/parser.go, state for state: case-insensitive instructions; # opens a comment only at
a line's start; a value runs to the end of the line; "…" and """…""" may span lines, with no escapes. The
vendored reference for the format is mindX's docs/ollama/setup/modelfile.md. quote()
writes values back as Ollama's ollama show --modelfile does, and they round-trip.
| instruction | handling |
|---|---|
FROM |
exactly one: a registry name (pinned or derived), its suffix-less alias, a pinned GGUF path, or a safetensors directory (converted to <name>-F16.gguf and pinned with every input's sha256) |
SYSTEM, TEMPLATE |
the last wins; TEMPLATE only when it is the base's own chat template (bankML renders that template byte-identically to llama.cpp; a different one would change every prompt) |
PARAMETER |
one of PARAMS (below) or stop; the last value wins, stop accumulates |
MESSAGE role content |
role system, user or assistant; recorded and applied, as Ollama applies them |
LICENSE, REQUIRES |
recorded |
ADAPTER |
refused |
FROM a derived model inherits its layer: new values win, stop is replaced as a list, licences add up. A relative
FROM ./x on the CLI is relative to the Modelfile.
PARAMS
pub const PARAMS: [(&str, bool); 13] = [("temperature", false), ("top_k", true), ("top_p", false), ("min_p", false),
("seed", true), ("num_ctx", true), ("num_predict", true), ("repeat_penalty", false), ("repeat_last_n", true),
("presence_penalty", false), ("frequency_penalty", false), ("typical_p", false), ("min_keep", true)];
The bool says whether the value is an integer; the order is the order a manifest keeps. The four penalties joined
in 0.3.6 (O2); typical_p and its min_keep in 0.3.7, when the engine came to reproduce typical-p. Floats are parsed at 32 bits, as Ollama parses them. Range checks at create:
num_ctx ≥ 1; num_predict ≥ −2; repeat_penalty > 0; frequency_penalty and presence_penalty any finite value
(negative too, as llama.cpp takes them); top_k and seed unchecked; the rest ≥ 0. A PARAMETER outside PARAMS
is refused with one of three reasons: a sampler not reproduced (tfs_z, mirostat, mirostat_eta,
mirostat_tau, penalize_newline); a resource option that belongs to bankml serve; or not a parameter bankML
knows.
Applying the layer (Ollama's server/routes.go)
apply_ollama(&self, req, chat): the layer's parameters become defaults under the request'soptions, overridden key by key (stopas a whole list). For chat, the model'sMESSAGEs go before the request's, and itsSYSTEMgoes first unless the request's first message is a system message. For a non-raw generate, the conversation goes inbankml_messages: the request'ssystemif any, else the model's, then theMESSAGEs, then the prompt. Arawprompt gets neither.apply_openai(&self, req): the same message rule (Ollama's OpenAI endpoint goes through the same chat handler); the parameters fill absent top-level fields (num_predictasmax_tokens);num_ctxis not a field but fits the conversation (num_ctx()).
Resolution
resolve(rs, model) tries a pin's exact name first, then a derived model's name, then the registry's suffix-less
alias (0.3.5). So promote.py's re-create in place (FROM mindx-gen39 as mindx-gen39, where mindx-gen39 is the
alias of the one pin mindx-gen39-f16) answers as itself with the layer, and mindx-gen39-f16 gives the pin. FROM
resolves the suffix-less alias the same way, and only when exactly one pin has it. A derived model loads through its base's entry, so it and its base share one residency.
HTTP (needs serve --native --registry)
curl -s $B/api/create -H "$J" -d '{"model": "mindx-persona", "from": "mindx-gen39", "system": "You are mindX.", "parameters": {"temperature": 0.7}}'
curl -s $B/api/create -H "$J" -d '{"model": "p", "modelfile": "FROM mindx-gen39\nPARAMETER repeat_penalty 1.3"}'
curl -s $B/api/copy -H "$J" -d '{"source": "mindx-persona", "destination": "mindx-persona-b"}'
curl -s -X DELETE $B/api/delete -H "$J" -d '{"model": "mindx-persona-b"}'
/api/create:{model, modelfile}or Ollama's structured{model, from, system, template, license, parameters, messages}; status lines streamed as NDJSON unlessstream: false. Without--registry: 400, nowhere to write./api/delete: derived models only; a pin is refused ("bankml does not delete pins over HTTP")./api/copy: the same content and digest under a new name; a pin's copy is an empty layer over it./api/showof a derived model (show_json): the reconstructed Modelfile,parameters,system,license,messages,details.parent_model, andbankml.layerwith the digest and the base's sha256 and pin.
CLI
bankml create NAME -f Modelfile [--registry DIR] [--models DIR] prints Ollama's status lines and exits 0, or
bankml create: refuse: … and 2. The registry defaults to $BANKML_FORKS, else ~/.local/share/bankml/forks.
How it is verified
- Unit tests:
modelfile_parses_as_ollama(mindX's own Modelfiles, comments, CRLF, BOM, quoting, errors),quote_round_trips_through_the_parser,spec_keeps_what_is_reproduced_and_refuses_the_rest,manifest_digest_is_the_content,the_layer_applies_as_ollama_applies_it,create_layers_over_a_pinned_base,http_create_show_copy_delete_round_trip. testing/cli.rs:create_writes_a_layer_over_a_pin.oracle_persona_layer(#[ignore], gate;testing/persona_oracle.py): promote.py's persona Modelfile created two ways,FROMthe merged safetensors directory andFROM mindx-gen39in place. Asked the user's turns alone, each must give llama-server b11192's tokens for the same GGUF given the persona as the system message. CHANGELOG 0.3.5: 27 / 27 each. Live:persona_oracle_livethrough/api/chat,/api/generate,/v1and/api/ps.
Advantages and efficiency
- No weights copied. A layer is a small JSON file over a pinned GGUF; its digest is content-addressed, as Ollama's content addressing does. Every load of a derived model verifies its base (the guard, then the pin).
- mindX's Modelfiles unchanged. The parser is Ollama's state machine, so promote.py's files create the same layer,
and
/api/showgives back a Modelfile that recreates it. - Efficient. Tamper detection is one sha256 over a short text. A derived model and its base share one load. A conversion never overwrites an existing pin.
- Rust practice. No external crates; atomic manifest writes;
RwLockpoisoning recovered; errors carry the reason asResult<_, String>. - Next (docs/OLLAMA.md, docs/TODO.md):
ADAPTER(LoRA merging) is a later O-phase;promote.py --to bankmlin mindX and a bankml backend for mindXtrain (O5, proposed).
Limitations
ADAPTER: "bankML does not merge LoRA adapters yet (a later O-phase); merge it into the base weights … then FROM the merged safetensors directory".TEMPLATEother than the base's own: Ollama's are Go templates; bankML renders the GGUF's Jinja byte-identically./api/createfiles,adapters,quantize: refused, each with its reason.- More than one
FROM; a name a pinned file already has; aFROMfile not pinned in the registry directory. - Model names: letters, digits,
.,_,-, at most 128, no tag butlatest, not ending in.gguf. - A layer
num_ctxabove the server's--ctxis accepted at create (with a warning) and refused per request. - The 0.3.7 samplers that are not Ollama options (DRY, XTC, top-n-σ, dynamic temperature) are not Modelfile
parameters; a request gives them through
/v1.
Design notes
bankml createis phase O5 of ../OLLAMA.md, first built in 0.3.5; since 0.3.5 a conversion undercreatenever overwrites an existing pin, nor its FORK.json wherever the file is.is_printapproximates Go'sstrconv.IsPrint(no control characters, no space other than U+0020, no format characters), which the parser uses as Ollama's does.- A create changes what requests resolve to, so
POST /api/createtakes the engine's run lock, as a model load does. - Parameter floats are stored at 32 bits, as Ollama parses them, and written in the shortest form that reads back.