Janus-35B-HERETIC / Modelfile
FoolDev's picture
Claude Opus 5 (1M context)
release 0.9.8: fix the system block, correct 0.9.7's render claim, land audit top tier
ee4c9b0
Raw History Blame Contribute Delete
12.5 kB
FROM ./Janus-35B-A3B.Q4_K_M.gguf
# The bundled GGUF carries no MTP / NextN block. The base
# (llmfan46/Qwen3.6-35B-A3B-uncensored-heretic) publishes its GGUFs already
# MTP-clean β€” blk.0…blk.39 only, block_count 40, no nextn_predict_layers key
# (read from the bundled file's own header, 2026-09-18) β€” so scripts/strip_mtp.py
# is a defensive no-op here, not a required build step. The separate
# llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GGUF variant
# does keep the MTP head (its Q4_K_M has block_count 41, NextN block at index
# 40) and would need the strip; this repo does not ship it. If you want the MTP
# head, run the upstream safetensors under vLLM/SGLang, or take that variant's
# GGUF.
# Chat template β€” Qwen 3.6 ChatML in Ollama Go-template form, with the
# tool-calling blocks Ollama's capability detector looks for. Without a
# TEMPLATE that references .Tools and .ToolCalls, an Ollama that uses this
# template anyway (OLLAMA_GO_TEMPLATE=1 - per Ollama's source, not tested live -
# or one that predates the template comparison described below) rejects any
# request carrying a `tools` array with `<model> does not support tools`;
# Ollama 0.33.3 would instead switch to the GGUF's embedded template. Same template as the dense
# 27B sibling (FoolDev/Thanatos-27B-HERETIC) β€” the ChatML wire format is
# identical across Qwen 3.6 and 3.8.
#
# Thinking IS replayed across turns: every assistant message that carries
# reasoning renders <think>...</think>, earlier turns included, where Qwen's
# stock condition kept only the turn in progress (a tool-call chain included).
# The client must send the reasoning back - per Ollama's source, /api/chat reads
# each assistant message's `thinking` field and /v1/chat/completions reads only
# `reasoning` (it drops `reasoning_content` and `thinking` silently). Every
# retained trace stays in the prompt and prefill grows to match, which is the
# slow part on a CPU-only host. The prompt-token counts that used to sit here
# were measured on the dense 27B this repo shipped in the interim; they are
# removed rather than carried over, pending re-measurement.
#
# The thinking block in the assistant branch is load-bearing for a second
# reason. Ollama infers a Go template's "thinking" capability from a `.Thinking` field inside
# `range .Messages` wrapped in <think>/</think>, and Ollama 0.33.3 and 0.34.0 (verified;
# the selection code is unchanged through 0.34.2, checked in source)
# compare that capability list with the GGUF's embedded Jinja template at load
# time, switching to the embedded template if that one advertises more. Delete
# the block and Ollama switches: in one test on the dense 27B this repo shipped
# in the interim a single tool call then came back twice
# (a separate raw generation showed the model drafting the call inside <think>
# first; that the parser matched such a draft is inferred, not confirmed). Per Ollama's source (not tested live), a
# setup forcing this template (OLLAMA_GO_TEMPLATE=1) would also reject thinking
# requests with HTTP 400. To stop replaying earlier turns' reasoning, change the
# condition back to Qwen's stock
# `(and $.IsThinkSet (and .Thinking (or $last (gt $i $lastUserIdx))))` - the
# $lastUserIdx loop at the top of the template is kept for it; do not delete the
# block.
#
# Two more properties of this template are read from its PARSE TREE, not its
# output, so no render test can see them break. Ollama takes the tool-call tag
# from the first `{{ if }}` whose condition names .ToolCalls and then the first
# literal text in that block, so (a) the `range .ToolCalls` body must start with
# text, never an action - that was 0.9.5 - and (b) no earlier condition may name
# .ToolCalls, which is why the assistant branch aliases it as
# `{{ $calls := .ToolCalls }}` and branches on the variable. The thinking tags
# come from the first and last nodes of the list holding {{ .Thinking }}, both of
# which must be literal text, so nothing conditional may end that block - the
# blank-line condition sits after it. scripts/check_go_template.py emulates both
# of Ollama's readers and asserts what they derive.
#
# Reasoning effort: the block at the top of the template injects an effort
# instruction keyed on Ollama's think level ($.ThinkLevel), not the request
# value. It is Ollama-only now β€” this repo no longer ships chat_template.jinja,
# and the Qwen 3.6 base's own embedded template, which governs the llama.cpp
# path from here on, has no reasoning_effort handling at all.
# Ollama folds `reasoning_effort` into four levels first:
# high/xhigh/max/ultra -> high or max (xhigh line), low/minimal -> low (low
# line), medium and unset -> medium (no line), none -> thinking off. The
# outer `if $.ThinkLevel` keeps the template rendering on Ollama builds that
# predate the field.
TEMPLATE """{{- $lastUserIdx := -1 -}}
{{- range $idx, $msg := .Messages -}}
{{- if eq $msg.Role "user" }}{{ $lastUserIdx = $idx }}{{ end -}}
{{- end }}
{{- $effort := "" }}
{{- if $.ThinkLevel }}
{{- if or (eq $.ThinkLevel "high") (eq $.ThinkLevel "max") }}{{ $effort = "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer." }}
{{- else if eq $.ThinkLevel "low" }}{{ $effort = "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration." }}
{{- end }}
{{- end }}
{{- if or .System .Tools $effort }}<|im_start|>system
{{ if $effort }}{{ $effort }}{{ if or .System .Tools }}
{{ end }}{{ end }}{{ if .System }}{{ .System }}{{ if .Tools }}
{{ end }}{{ end }}
{{- if .Tools }}# Tools
You may call one or more functions to assist with the user query.
You are provided with function signatures within <tools></tools> XML tags:
<tools>
{{- range .Tools }}
{"type": "function", "function": {{ json .Function }}}
{{- end }}
</tools>
For each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:
<tool_call>
{"name": <function-name>, "arguments": <args-json-object>}
</tool_call>
{{- end -}}<|im_end|>
{{ end }}
{{- $prevIsTool := false }}
{{- range $i, $_ := .Messages }}
{{- $rest := slice $.Messages $i }}
{{- $last := eq (len $rest) 1 -}}
{{- $nextIsTool := and (gt (len $rest) 1) (eq (index $rest 1).Role "tool") -}}
{{- if eq .Role "user" }}<|im_start|>user
{{ .Content }}<|im_end|>
{{ else if eq .Role "assistant" }}{{ $calls := .ToolCalls }}{{ $blank := or .Content (not $calls) }}<|im_start|>assistant
{{ if $.IsThinkSet -}}
<think>
{{ .Thinking }}
</think>
{{ end -}}
{{ if and $.IsThinkSet $blank }}
{{ end -}}
{{ if .Content }}{{ .Content }}{{ if $calls }}
{{ end }}{{ end }}
{{- if .ToolCalls }}
{{- range .ToolCalls }}
<tool_call>
{"name": "{{ .Function.Name }}", "arguments": {{ .Function.Arguments }}}
</tool_call>
{{- end }}
{{- end }}{{ if not $last }}<|im_end|>
{{ end }}
{{- else if eq .Role "tool" }}
{{- if $prevIsTool }}
{{ else }}<|im_start|>user
{{ end }}<tool_response>
{{ .Content }}
</tool_response>
{{- if not $nextIsTool }}<|im_end|>
{{ end }}
{{- end }}
{{- $prevIsTool = eq .Role "tool" }}
{{- if and (ne .Role "assistant") $last }}<|im_start|>assistant
{{ if and $.IsThinkSet (not $.Think) -}}
<think>
</think>
{{ else -}}
<think>
{{ end -}}
{{ end }}
{{- end }}"""
# Sampling tuned for reasoning + general use. See README "Recommended sampling"
# for creative/RP alternatives.
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER top_k 0
PARAMETER repeat_penalty 1.05
PARAMETER num_ctx 262144
# Stop tokens. Without these, Ollama only honors <|im_end|> from the GGUF
# metadata; the model occasionally emits <|endoftext|> instead and Ollama
# keeps generating past it (synthesising a fake new user turn). Listing
# both β€” plus <|im_start|> as a belt-and-braces guard against the same
# loop β€” keeps responses cleanly terminated. Same fix the sibling
# (FoolDev/Thanatos-27B-HERETIC) shipped in commit 6672746.
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
PARAMETER stop "<|im_start|>"
SYSTEM """You are Janus, a precise and capable assistant for reasoning, writing, coding, and long-form dialogue.
Behavior rules:
- Answer the user's actual request directly.
- Be accurate, complete, and structured.
- Think before answering, but do not get stuck in repetitive loops or meta-commentary.
- If the request is ambiguous or incomplete, state what is missing and make the smallest reasonable assumption needed to continue.
- If the user wants creative writing, preserve tone, continuity, and character consistency.
- If the user wants analysis or technical help, prefer concrete steps, examples, and decisions over fluff.
- Finish with a usable answer, not just planning."""
# Hardware notes
# --------------
# This Q4_K_M is 21,233,608,512 bytes on disk β€” 21.23 GB decimal, 19.78 GiB.
# Measured 2026-09-18 on Ollama 0.33.3's CPU backend, from llama.cpp's own
# allocation lines, in an isolated store at num_ctx 8192 / 32768 / 65536:
# weights 19.77 GiB (CPU model buffer 5361.34 MiB + CPU_REPACK
# 14878.12 MiB); constant
# KV cache only 10 of the 40 layers are full-attention (indices
# 3, 7, ... 39), and llama.cpp logs the cache as "10
# layers": 20,480 B/token at f16 = 160 / 640 / 1280 MiB
# at 8K / 32K / 64K, i.e. 0.625 GiB per 32K, and
# 10,880 B/token at q8_0 = 340 / 680 MiB at 32K / 64K,
# i.e. 0.332 GiB per 32K. Exactly linear in num_ctx ->
# 5.0 GiB at the 262144 default with f16, 2.66 with q8_0
# recurrent state 62.81 MiB for the 30 Gated-DeltaNet layers (R f32 2.81
# + S f32 60.00) β€” identical at every context
# compute buffer 160.04 / 208.04 / 544.07 MiB at 8K / 32K / 64K; grows
# faster than linearly, so it is not extrapolated
# totals 20.7 GiB at 32768 and 21.6 GiB at 65536 (both sums of
# measured parts); ~25.4 GiB at the 262144 default with
# the f16 cache and ~23.0 GiB with
# OLLAMA_KV_CACHE_TYPE=q8_0 β€” those two carry the 64K
# compute buffer forward, so treat them as floors.
# No YaRN rope-scaling is baked in this GGUF, so output
# past ~262K degrades; treat 1.01M as an advertised
# ceiling.
#
# This is a 35B-A3B MoE across 40 layers: 256 experts, 8 routed per token plus
# one shared expert, ~34.7B parameters total and ~3B active per token. The active
# count cuts compute per token, not the memory floor β€” every expert has to be
# resident, so budget for the whole 19.77 GiB of weights.
#
# Working configurations (rows assume the 262144 default: ~25.4 GiB with the f16
# cache, ~23.0 GiB with q8_0):
# βœ“ Single H100 80GB / A100 80GB β€” full GPU offload
# βœ“ RTX 5090 32GB / RTX 4090 24GB + 32GB RAM β€” partial offload
# βœ“ Mac Studio M2/M3 Ultra 64GB+ β€” unified memory
# βœ“ Linux box with 48GB+ RAM (CPU-only) β€” CPU inference
# ⚠ 32GB hosts (e.g. ASUS ROG Flow Z13) β€” set OLLAMA_KV_CACHE_TYPE=q8_0
# or lower num_ctx
#
# Throughput, measured 2026-09-18 on CPU only (Ryzen AI Max+ 395, Ollama 0.33.3,
# isolated store): ./scripts/bench.sh aggregated 24.97 tok/s over its three-prompt
# mix (4,543 tokens / 181,890 ms; 25.88 / 25.57 / 24.87 individually), about 5x the
# 4.97 tok/s the dense 27B managed on the same CPU β€” that is the ~3B active
# parameters per token. OLLAMA_KV_CACHE_TYPE=q8_0 with OLLAMA_FLASH_ATTENTION=1
# aggregated 24.85 tok/s over the same mix, 0.5% below f16 β€” it buys memory, not
# speed. Generation only:
# this is a reasoning-first model, so an answer costs many more tokens than its
# visible length suggests.
#
# To run on a 32 GB unified-memory laptop, override these in your local
# Modelfile copy (or via `/set parameter` in the interactive `ollama run` REPL):
# PARAMETER num_ctx 4096
# PARAMETER num_batch 256
#
# If you have β‰₯48 GB RAM but want partial GPU offload, set:
# PARAMETER num_gpu 24 # offload most layers (model has 40)