Text Generation
GGUF
qwen36
Mixture of Experts
conversational
multimodal
agent
ollama
heretic
uncensored
reasoning
distillation
Instructions to use FoolDev/Janus-35B-HERETIC with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use FoolDev/Janus-35B-HERETIC with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: llama cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: llama cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Use Docker
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use FoolDev/Janus-35B-HERETIC with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FoolDev/Janus-35B-HERETIC" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FoolDev/Janus-35B-HERETIC", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Ollama
How to use FoolDev/Janus-35B-HERETIC with Ollama:
ollama run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Unsloth Desktop
- Pi
How to use FoolDev/Janus-35B-HERETIC with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "FoolDev/Janus-35B-HERETIC:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use FoolDev/Janus-35B-HERETIC with Docker Model Runner:
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Lemonade
How to use FoolDev/Janus-35B-HERETIC with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull FoolDev/Janus-35B-HERETIC:Q4_K_M
Run and chat with the model
lemonade run user.Janus-35B-HERETIC-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use FoolDev/Janus-35B-HERETIC with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default FoolDev/Janus-35B-HERETIC:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use FoolDev/Janus-35B-HERETIC with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "FoolDev/Janus-35B-HERETIC:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download Modelfile from FoolDev/Janus-35B-HERETIC: direct link, hf CLI and curl.
- Browser
- Download file 12.5 kB
-
https://huggingface.co/FoolDev/Janus-35B-HERETIC/resolve/main/Modelfile
- Command line
-
hf download hf://FoolDev/Janus-35B-HERETIC/Modelfile
-
curl -L -o Modelfile https://huggingface.co/FoolDev/Janus-35B-HERETIC/resolve/main/Modelfile
12.5 kB
| FROM ./Janus-35B-A3B.Q4_K_M.gguf | |
| # The bundled GGUF carries no MTP / NextN block. The base | |
| # (llmfan46/Qwen3.6-35B-A3B-uncensored-heretic) publishes its GGUFs already | |
| # MTP-clean β blk.0β¦blk.39 only, block_count 40, no nextn_predict_layers key | |
| # (read from the bundled file's own header, 2026-09-18) β so scripts/strip_mtp.py | |
| # is a defensive no-op here, not a required build step. The separate | |
| # llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GGUF variant | |
| # does keep the MTP head (its Q4_K_M has block_count 41, NextN block at index | |
| # 40) and would need the strip; this repo does not ship it. If you want the MTP | |
| # head, run the upstream safetensors under vLLM/SGLang, or take that variant's | |
| # GGUF. | |
| # Chat template β Qwen 3.6 ChatML in Ollama Go-template form, with the | |
| # tool-calling blocks Ollama's capability detector looks for. Without a | |
| # TEMPLATE that references .Tools and .ToolCalls, an Ollama that uses this | |
| # template anyway (OLLAMA_GO_TEMPLATE=1 - per Ollama's source, not tested live - | |
| # or one that predates the template comparison described below) rejects any | |
| # request carrying a `tools` array with `<model> does not support tools`; | |
| # Ollama 0.33.3 would instead switch to the GGUF's embedded template. Same template as the dense | |
| # 27B sibling (FoolDev/Thanatos-27B-HERETIC) β the ChatML wire format is | |
| # identical across Qwen 3.6 and 3.8. | |
| # | |
| # Thinking IS replayed across turns: every assistant message that carries | |
| # reasoning renders <think>...</think>, earlier turns included, where Qwen's | |
| # stock condition kept only the turn in progress (a tool-call chain included). | |
| # The client must send the reasoning back - per Ollama's source, /api/chat reads | |
| # each assistant message's `thinking` field and /v1/chat/completions reads only | |
| # `reasoning` (it drops `reasoning_content` and `thinking` silently). Every | |
| # retained trace stays in the prompt and prefill grows to match, which is the | |
| # slow part on a CPU-only host. The prompt-token counts that used to sit here | |
| # were measured on the dense 27B this repo shipped in the interim; they are | |
| # removed rather than carried over, pending re-measurement. | |
| # | |
| # The thinking block in the assistant branch is load-bearing for a second | |
| # reason. Ollama infers a Go template's "thinking" capability from a `.Thinking` field inside | |
| # `range .Messages` wrapped in <think>/</think>, and Ollama 0.33.3 and 0.34.0 (verified; | |
| # the selection code is unchanged through 0.34.2, checked in source) | |
| # compare that capability list with the GGUF's embedded Jinja template at load | |
| # time, switching to the embedded template if that one advertises more. Delete | |
| # the block and Ollama switches: in one test on the dense 27B this repo shipped | |
| # in the interim a single tool call then came back twice | |
| # (a separate raw generation showed the model drafting the call inside <think> | |
| # first; that the parser matched such a draft is inferred, not confirmed). Per Ollama's source (not tested live), a | |
| # setup forcing this template (OLLAMA_GO_TEMPLATE=1) would also reject thinking | |
| # requests with HTTP 400. To stop replaying earlier turns' reasoning, change the | |
| # condition back to Qwen's stock | |
| # `(and $.IsThinkSet (and .Thinking (or $last (gt $i $lastUserIdx))))` - the | |
| # $lastUserIdx loop at the top of the template is kept for it; do not delete the | |
| # block. | |
| # | |
| # Two more properties of this template are read from its PARSE TREE, not its | |
| # output, so no render test can see them break. Ollama takes the tool-call tag | |
| # from the first `{{ if }}` whose condition names .ToolCalls and then the first | |
| # literal text in that block, so (a) the `range .ToolCalls` body must start with | |
| # text, never an action - that was 0.9.5 - and (b) no earlier condition may name | |
| # .ToolCalls, which is why the assistant branch aliases it as | |
| # `{{ $calls := .ToolCalls }}` and branches on the variable. The thinking tags | |
| # come from the first and last nodes of the list holding {{ .Thinking }}, both of | |
| # which must be literal text, so nothing conditional may end that block - the | |
| # blank-line condition sits after it. scripts/check_go_template.py emulates both | |
| # of Ollama's readers and asserts what they derive. | |
| # | |
| # Reasoning effort: the block at the top of the template injects an effort | |
| # instruction keyed on Ollama's think level ($.ThinkLevel), not the request | |
| # value. It is Ollama-only now β this repo no longer ships chat_template.jinja, | |
| # and the Qwen 3.6 base's own embedded template, which governs the llama.cpp | |
| # path from here on, has no reasoning_effort handling at all. | |
| # Ollama folds `reasoning_effort` into four levels first: | |
| # high/xhigh/max/ultra -> high or max (xhigh line), low/minimal -> low (low | |
| # line), medium and unset -> medium (no line), none -> thinking off. The | |
| # outer `if $.ThinkLevel` keeps the template rendering on Ollama builds that | |
| # predate the field. | |
| TEMPLATE """{{- $lastUserIdx := -1 -}} | |
| {{- range $idx, $msg := .Messages -}} | |
| {{- if eq $msg.Role "user" }}{{ $lastUserIdx = $idx }}{{ end -}} | |
| {{- end }} | |
| {{- $effort := "" }} | |
| {{- if $.ThinkLevel }} | |
| {{- if or (eq $.ThinkLevel "high") (eq $.ThinkLevel "max") }}{{ $effort = "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer." }} | |
| {{- else if eq $.ThinkLevel "low" }}{{ $effort = "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration." }} | |
| {{- end }} | |
| {{- end }} | |
| {{- if or .System .Tools $effort }}<|im_start|>system | |
| {{ if $effort }}{{ $effort }}{{ if or .System .Tools }} | |
| {{ end }}{{ end }}{{ if .System }}{{ .System }}{{ if .Tools }} | |
| {{ end }}{{ end }} | |
| {{- if .Tools }}# Tools | |
| You may call one or more functions to assist with the user query. | |
| You are provided with function signatures within <tools></tools> XML tags: | |
| <tools> | |
| {{- range .Tools }} | |
| {"type": "function", "function": {{ json .Function }}} | |
| {{- end }} | |
| </tools> | |
| For each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags: | |
| <tool_call> | |
| {"name": <function-name>, "arguments": <args-json-object>} | |
| </tool_call> | |
| {{- end -}}<|im_end|> | |
| {{ end }} | |
| {{- $prevIsTool := false }} | |
| {{- range $i, $_ := .Messages }} | |
| {{- $rest := slice $.Messages $i }} | |
| {{- $last := eq (len $rest) 1 -}} | |
| {{- $nextIsTool := and (gt (len $rest) 1) (eq (index $rest 1).Role "tool") -}} | |
| {{- if eq .Role "user" }}<|im_start|>user | |
| {{ .Content }}<|im_end|> | |
| {{ else if eq .Role "assistant" }}{{ $calls := .ToolCalls }}{{ $blank := or .Content (not $calls) }}<|im_start|>assistant | |
| {{ if $.IsThinkSet -}} | |
| <think> | |
| {{ .Thinking }} | |
| </think> | |
| {{ end -}} | |
| {{ if and $.IsThinkSet $blank }} | |
| {{ end -}} | |
| {{ if .Content }}{{ .Content }}{{ if $calls }} | |
| {{ end }}{{ end }} | |
| {{- if .ToolCalls }} | |
| {{- range .ToolCalls }} | |
| <tool_call> | |
| {"name": "{{ .Function.Name }}", "arguments": {{ .Function.Arguments }}} | |
| </tool_call> | |
| {{- end }} | |
| {{- end }}{{ if not $last }}<|im_end|> | |
| {{ end }} | |
| {{- else if eq .Role "tool" }} | |
| {{- if $prevIsTool }} | |
| {{ else }}<|im_start|>user | |
| {{ end }}<tool_response> | |
| {{ .Content }} | |
| </tool_response> | |
| {{- if not $nextIsTool }}<|im_end|> | |
| {{ end }} | |
| {{- end }} | |
| {{- $prevIsTool = eq .Role "tool" }} | |
| {{- if and (ne .Role "assistant") $last }}<|im_start|>assistant | |
| {{ if and $.IsThinkSet (not $.Think) -}} | |
| <think> | |
| </think> | |
| {{ else -}} | |
| <think> | |
| {{ end -}} | |
| {{ end }} | |
| {{- end }}""" | |
| # Sampling tuned for reasoning + general use. See README "Recommended sampling" | |
| # for creative/RP alternatives. | |
| PARAMETER temperature 1.0 | |
| PARAMETER top_p 0.95 | |
| PARAMETER top_k 0 | |
| PARAMETER repeat_penalty 1.05 | |
| PARAMETER num_ctx 262144 | |
| # Stop tokens. Without these, Ollama only honors <|im_end|> from the GGUF | |
| # metadata; the model occasionally emits <|endoftext|> instead and Ollama | |
| # keeps generating past it (synthesising a fake new user turn). Listing | |
| # both β plus <|im_start|> as a belt-and-braces guard against the same | |
| # loop β keeps responses cleanly terminated. Same fix the sibling | |
| # (FoolDev/Thanatos-27B-HERETIC) shipped in commit 6672746. | |
| PARAMETER stop "<|im_end|>" | |
| PARAMETER stop "<|endoftext|>" | |
| PARAMETER stop "<|im_start|>" | |
| SYSTEM """You are Janus, a precise and capable assistant for reasoning, writing, coding, and long-form dialogue. | |
| Behavior rules: | |
| - Answer the user's actual request directly. | |
| - Be accurate, complete, and structured. | |
| - Think before answering, but do not get stuck in repetitive loops or meta-commentary. | |
| - If the request is ambiguous or incomplete, state what is missing and make the smallest reasonable assumption needed to continue. | |
| - If the user wants creative writing, preserve tone, continuity, and character consistency. | |
| - If the user wants analysis or technical help, prefer concrete steps, examples, and decisions over fluff. | |
| - Finish with a usable answer, not just planning.""" | |
| # Hardware notes | |
| # -------------- | |
| # This Q4_K_M is 21,233,608,512 bytes on disk β 21.23 GB decimal, 19.78 GiB. | |
| # Measured 2026-09-18 on Ollama 0.33.3's CPU backend, from llama.cpp's own | |
| # allocation lines, in an isolated store at num_ctx 8192 / 32768 / 65536: | |
| # weights 19.77 GiB (CPU model buffer 5361.34 MiB + CPU_REPACK | |
| # 14878.12 MiB); constant | |
| # KV cache only 10 of the 40 layers are full-attention (indices | |
| # 3, 7, ... 39), and llama.cpp logs the cache as "10 | |
| # layers": 20,480 B/token at f16 = 160 / 640 / 1280 MiB | |
| # at 8K / 32K / 64K, i.e. 0.625 GiB per 32K, and | |
| # 10,880 B/token at q8_0 = 340 / 680 MiB at 32K / 64K, | |
| # i.e. 0.332 GiB per 32K. Exactly linear in num_ctx -> | |
| # 5.0 GiB at the 262144 default with f16, 2.66 with q8_0 | |
| # recurrent state 62.81 MiB for the 30 Gated-DeltaNet layers (R f32 2.81 | |
| # + S f32 60.00) β identical at every context | |
| # compute buffer 160.04 / 208.04 / 544.07 MiB at 8K / 32K / 64K; grows | |
| # faster than linearly, so it is not extrapolated | |
| # totals 20.7 GiB at 32768 and 21.6 GiB at 65536 (both sums of | |
| # measured parts); ~25.4 GiB at the 262144 default with | |
| # the f16 cache and ~23.0 GiB with | |
| # OLLAMA_KV_CACHE_TYPE=q8_0 β those two carry the 64K | |
| # compute buffer forward, so treat them as floors. | |
| # No YaRN rope-scaling is baked in this GGUF, so output | |
| # past ~262K degrades; treat 1.01M as an advertised | |
| # ceiling. | |
| # | |
| # This is a 35B-A3B MoE across 40 layers: 256 experts, 8 routed per token plus | |
| # one shared expert, ~34.7B parameters total and ~3B active per token. The active | |
| # count cuts compute per token, not the memory floor β every expert has to be | |
| # resident, so budget for the whole 19.77 GiB of weights. | |
| # | |
| # Working configurations (rows assume the 262144 default: ~25.4 GiB with the f16 | |
| # cache, ~23.0 GiB with q8_0): | |
| # β Single H100 80GB / A100 80GB β full GPU offload | |
| # β RTX 5090 32GB / RTX 4090 24GB + 32GB RAM β partial offload | |
| # β Mac Studio M2/M3 Ultra 64GB+ β unified memory | |
| # β Linux box with 48GB+ RAM (CPU-only) β CPU inference | |
| # β 32GB hosts (e.g. ASUS ROG Flow Z13) β set OLLAMA_KV_CACHE_TYPE=q8_0 | |
| # or lower num_ctx | |
| # | |
| # Throughput, measured 2026-09-18 on CPU only (Ryzen AI Max+ 395, Ollama 0.33.3, | |
| # isolated store): ./scripts/bench.sh aggregated 24.97 tok/s over its three-prompt | |
| # mix (4,543 tokens / 181,890 ms; 25.88 / 25.57 / 24.87 individually), about 5x the | |
| # 4.97 tok/s the dense 27B managed on the same CPU β that is the ~3B active | |
| # parameters per token. OLLAMA_KV_CACHE_TYPE=q8_0 with OLLAMA_FLASH_ATTENTION=1 | |
| # aggregated 24.85 tok/s over the same mix, 0.5% below f16 β it buys memory, not | |
| # speed. Generation only: | |
| # this is a reasoning-first model, so an answer costs many more tokens than its | |
| # visible length suggests. | |
| # | |
| # To run on a 32 GB unified-memory laptop, override these in your local | |
| # Modelfile copy (or via `/set parameter` in the interactive `ollama run` REPL): | |
| # PARAMETER num_ctx 4096 | |
| # PARAMETER num_batch 256 | |
| # | |
| # If you have β₯48 GB RAM but want partial GPU offload, set: | |
| # PARAMETER num_gpu 24 # offload most layers (model has 40) | |