Mikrodev commited on
Commit
504d2f2
·
verified ·
1 Parent(s): fc52946

Upload Modelfile with huggingface_hub

Browse files
Files changed (1) hide show
  1. Modelfile +10 -14
Modelfile CHANGED
@@ -10,26 +10,22 @@
10
  # 2) ollama create stcoder-gemma4-12b:<quant> -f Modelfile e.g. ollama create stcoder-gemma4-12b:q8_0 -f Modelfile
11
  # 3) ollama run stcoder-gemma4-12b:<quant> "Motor starts 5 seconds after the start button; stop and E-stop drop it."
12
  #
13
- # Requires Ollama 0.5 or newer.
 
 
 
 
 
 
 
 
14
  # num_ctx 8192 matches the sequence length this model was fine-tuned at.
15
  # Published numbers were measured greedily (temperature 0, seed 42) at num_ctx 16384 / num_predict 8192. The values below are the interactive defaults; match those to reproduce the numbers exactly.
16
- # Gemma turn format, not ChatML. Do not mix the two.
17
- # Q8_0 crashed on a 16 GiB card in our own testing at this context size; drop to Q6_K if you hit an out-of-memory or a load failure.
18
  # Q8_0: crashed on a 16 GiB card in our own testing at 8k context - if that happens, drop to Q6_K.
19
 
20
  FROM ./gemma4_12b-tc.q8_0.gguf
21
 
22
- TEMPLATE """<start_of_turn>user
23
- {{ if .System }}{{ .System }}
24
-
25
- {{ end }}{{ .Prompt }}<end_of_turn>
26
- <start_of_turn>model
27
- {{ .Response }}<end_of_turn>
28
- """
29
-
30
- PARAMETER stop "<end_of_turn>"
31
- PARAMETER stop "<start_of_turn>"
32
-
33
  PARAMETER temperature 0.2
34
  PARAMETER top_p 0.95
35
  PARAMETER top_k 20
 
10
  # 2) ollama create stcoder-gemma4-12b:<quant> -f Modelfile e.g. ollama create stcoder-gemma4-12b:q8_0 -f Modelfile
11
  # 3) ollama run stcoder-gemma4-12b:<quant> "Motor starts 5 seconds after the start button; stop and E-stop drop it."
12
  #
13
+ # Requires Ollama 0.32 or newer; verified on 0.32.5.
14
+ #
15
+ # No TEMPLATE line here, deliberately: Ollama uses its built-in `gemma4` renderer, which matches the chat template stored inside the GGUF
16
+ # - the template of the tokenizer this model was trained with, so it cannot drift out of
17
+ # sync with the weights. We measured this: adding a TEMPLATE to this file does not change
18
+ # the prompt the model receives (identical prompt token counts with, without, and with a
19
+ # deliberately wrong template). With llama.cpp directly, pass --jinja so llama-cli and
20
+ # llama-server use that same embedded template.
21
+ #
22
  # num_ctx 8192 matches the sequence length this model was fine-tuned at.
23
  # Published numbers were measured greedily (temperature 0, seed 42) at num_ctx 16384 / num_predict 8192. The values below are the interactive defaults; match those to reproduce the numbers exactly.
24
+ # Gemma 4 support in Ollama is recent. If `ollama create` fails on this GGUF, update Ollama before anything else.
 
25
  # Q8_0: crashed on a 16 GiB card in our own testing at 8k context - if that happens, drop to Q6_K.
26
 
27
  FROM ./gemma4_12b-tc.q8_0.gguf
28
 
 
 
 
 
 
 
 
 
 
 
 
29
  PARAMETER temperature 0.2
30
  PARAMETER top_p 0.95
31
  PARAMETER top_k 20