How to use from
OpenClaw
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "FoolDev/Janus-35B-HERETIC:Q4_K_M" \
  --custom-provider-id llama-cpp \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links
Janus-35B banner

License Base Model Architecture Quant Sibling

Buy me a coffee — support Janus-35B

Janus-35B

Flagship Reasoning. Sparse Footprint. Uncensored. llmfan46's Heretic abliteration of Qwen 3.6 35B-A3B, repackaged with Claude Fable 5 in the teacher slot.

Architecture: Qwen 3.6 35B-A3B (MoE) | Total Params: 35B | Active Params: 3B | Base: Heretic (llmfan46) | Teacher: Claude Fable 5 | Type: Distilled + Abliterated MoE LLM

A personal fork of llmfan46/Qwen3.6-35B-A3B-uncensored-heretic — an uncensored Heretic-style abliteration of Qwen/Qwen3.6-35B-A3B, the 35B-total / 3B-active mixture-of-experts multimodal base — repackaged as Janus-35B with Claude Fable 5 reasoning data in the teacher slot. Refusal-trained behavior is dialed back at the base layer.

TL;DR

One-liner via Hugging Face (pulls a GGUF + this repo's root-level template / system / params files, including the tool-calling template — HF's Ollama bridge ingests those three files, not Modelfile):

ollama run hf.co/FoolDev/Janus-35B-HERETIC               # default ~19 GB Q4_K_M
ollama run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M        # same blob, explicit tag

Or build locally (uses this repo's Modelfile, kept in sync with the three bridge files):

git clone https://huggingface.co/FoolDev/Janus-35B-HERETIC && cd Janus-35B-HERETIC
ollama create janus -f Modelfile && ollama run janus

After either path, ollama show janus lists completion, tools, and thinking under Capabilities. Hardware: the default num_ctx is 1000000 — a ~1M ceiling above the 262144 native window (YaRN is not baked into this GGUF, so context past ~262K degrades) — so trim it down to fit your host (see Hardware requirements).

What's here

File Use
Janus-35B-A3B.Q4_K_M.gguf Recommended default, ~19 GB
Modelfile Ollama wrapper for local builds (ollama create janus -f Modelfile) — overrides the GGUF's embedded template with one that exposes .Tools / .ToolCalls to Ollama's capability detector.
template, system, params Used by HF's Ollama bridge when users ollama run hf.co/FoolDev/Janus-35B-HERETIC directly. The bridge does not read Modelfile (see HF Ollama docs); it ingests these three root-level files instead. Kept in sync with the Modelfile's TEMPLATE / SYSTEM / PARAMETER directives.
scripts/build.sh Pulls a GGUF from llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-GGUF (default Q4_K_M) and runs ollama create janus. The bundled Q4_K_M is already this Heretic quant; use this to build other quants locally.
scripts/check_bridge_sync.py Run before pushing a Modelfile / template / system / params edit to verify the four configurations remain in sync. Exits 0 if in sync, 1 with a per-key diff if not.
scripts/smoke_test.sh Integration smoke test against a running Ollama daemon: server reachable, model loaded, tools capability present, chat round-trip, and no control-token leakage. TOOLS_TEST=1 adds a tool-call round-trip. Defaults to MODEL=janus.
scripts/bench.sh Measures tok/s from Ollama's eval_count / eval_duration over a short/medium/long prompt mix (with a discarded warmup). Defaults to MODEL=janus.
scripts/load_bundle.sh Loads the bundled Janus-35B-A3B.Q4_K_M.gguf into Ollama as a local janus tag without an upstream pull (smudges the LFS pointer via hf download if needed, checks the arch is qwen35moe).
scripts/fetch_vision.sh Downloads the vision projector (Qwen3.6-35B-A3B-uncensored-heretic-mmproj-BF16.gguf) from the Heretic GGUF repo for llama.cpp image input (Ollama vision is broken upstream — see Vision).
examples/ Ready-to-run Python clients for Ollama, Transformers, and llama-cpp-python (text, tools, and vision — see examples/README.md)

GGUF-only release. Pull the Heretic safetensors from llmfan46/Qwen3.6-35B-A3B-uncensored-heretic if you need the transformers tree (or the vanilla pre-Heretic base from Qwen/Qwen3.6-35B-A3B).

Bundled blob status: the bundled Janus-35B-A3B.Q4_K_M.gguf is the Heretic Q4_K_M quant (from llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-GGUF), qwen35moe-stamped and verified against the Architecture below (40 layers, 256 experts, vocab 248,320). It serves the uncensored Heretic behavior directly; ./scripts/build.sh remains the path for other quants.

Architecture

animated MoE routing visualization: 16x16 grid of 256 expert dots with 8 lit at any time, cycling through 8 routing patterns

  • Qwen 3.6, 35B total / 3B active, MoE (256 experts, 8 activated per token)
  • 40 layers, 10 × (3 × DeltaNet → MoE / 1 × Gated Attention → MoE)
  • 262 144 native context (extensible to ~1 M with YaRN, but YaRN is not enabled in the bundled GGUF)
  • Vision + video supported by upstream (mmproj not included in this release)
  • Vocab 248,320

Quick start

llama.cpp / LM Studio

Drop the GGUF into your loader of choice. The chat template is embedded in the GGUF metadata, so llama.cpp's --chat-template auto and LM Studio's GGUF auto-detection handle plain conversation correctly.

Ollama

The chat template baked into the GGUF is not sufficient on Ollama — it lacks the .Tools / .ToolCalls blocks Ollama's capability detector requires, so a naive ollama pull reports does not support tools and rejects any request carrying a tools array. Two paths fix this:

# A. Pull straight from HF (uses the root-level template/system/params files):
ollama run hf.co/FoolDev/Janus-35B-HERETIC               # default tag, ~19 GB Q4_K_M
ollama run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M        # same blob, explicit tag
# Note: HF's Ollama bridge does NOT read Modelfile; it reads template/system/params.

# B. Build locally (uses Modelfile, which is kept in sync with the three above):
ollama create janus -f Modelfile && ollama run janus

After either path, ollama show janus should list completion, tools, and thinking under Capabilities.

Inference examples

Once the model is loaded (via ollama run janus, lms server, or llama-server), all the standard OpenAI-compatible clients work. Examples assume the loader is listening on http://localhost:11434 (Ollama default) — adjust the port for LM Studio (:1234) or llama.cpp (:8080). Runnable versions of everything below live in examples/.

curl

curl -s http://localhost:11434/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "janus",
    "messages": [
      {"role": "system", "content": "You are Janus, a precise reasoning assistant."},
      {"role": "user", "content": "Sketch an algorithm to detect cycles in a directed graph."}
    ],
    "temperature": 0.6,
    "max_tokens": 800
  }' | jq -r '.choices[0].message.content'

Python (openai-compat)

from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ignored")

resp = client.chat.completions.create(
    model="janus",
    messages=[
        {"role": "user", "content": "Write a haiku about a stack overflow."}
    ],
    temperature=0.8,
    top_p=0.95,
)
print(resp.choices[0].message.content)

Streaming

stream = client.chat.completions.create(
    model="janus",
    messages=[{"role": "user", "content": "Explain RoPE briefly."}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content or ""
    print(delta, end="", flush=True)

Recommended sampling

Use temp top_p top_k repeat_penalty
Default (Fable-matched) 1.0 0.95 0 1.05
Tighter reasoning 0.6 0.95 20 1.05
Creative / RP 0.8 0.95 40 1.02

The shipped default is Fable-matched — warm (temperature 1.0), no top_k, with top_p 0.95 + repeat_penalty 1.05 kept as loop insurance. Drop to the reasoning row for tighter, more deterministic output; lower temperature (0.4–0.6) and bump repeat_penalty to 1.08 if it loops inside <think> tags.

System prompt

You are Janus, a precise and capable assistant for reasoning, writing, coding, and long-form dialogue.

Behavior rules:
- Answer the user's actual request directly.
- Be accurate, complete, and structured.
- Think before answering, but do not get stuck in repetitive loops or meta-commentary.
- If the request is ambiguous or incomplete, state what is missing and make the smallest reasonable assumption needed to continue.
- If the user wants creative writing, preserve tone, continuity, and character consistency.
- If the user wants analysis or technical help, prefer concrete steps, examples, and decisions over fluff.
- Finish with a usable answer, not just planning.

Vision

The Qwen 3.6 base supports image (and video) input via a separate mmproj projector. The full multimodal stack is:

Janus-35B-A3B.Q4_K_M.gguf                              (~19 GB, the text decoder)
Qwen3.6-35B-A3B-uncensored-heretic-mmproj-BF16.gguf    (~903 MB, the vision projector)

The projector and other-quant text decoders live at llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-GGUF (BF16 mmproj only). For the vanilla pre-Heretic projector in F16/F32, see unsloth/Qwen3.6-35B-A3B-GGUF (mmproj-F16.gguf). This repo intentionally does not redistribute either; ./scripts/fetch_vision.sh pulls the projector into the repo root.

Loader compatibility

Loader Text Vision (mmproj) Notes
llama.cpp (llama-mtmd-cli, llama-server --mmproj) Reference path. Upstream has the qwen35moe arch entry.
llama-cpp-python See examples/llama_cpp_vision.py.
Ollama 0.24 Text inference works: Ollama's Go engine has the qwen35 / qwen35moe arch entries. Vision (mmproj) is still broken: the C++ llama.cpp fallback that Ollama switches to when an mmproj is attached lacks those entries. ollama create accepts a dual-FROM (text + mmproj) and ollama show reports vision capability — but the first inference request fails with error loading model architecture: unknown model architecture: 'qwen35moe', and once mmproj is attached this blocks text inference too. See ollama/ollama#14575 (open — the earlier #15898 was closed as its duplicate, and the sync PR #15899 was closed unmerged).
LM Studio Uses upstream llama.cpp directly.

Vision via llama.cpp

# Fetch the projector first (into the repo root):
./scripts/fetch_vision.sh                    # Qwen3.6-35B-A3B-uncensored-heretic-mmproj-BF16.gguf

# A. HTTP via llama-server (the easiest path):
llama-server \
  -m Janus-35B-A3B.Q4_K_M.gguf \
  --mmproj Qwen3.6-35B-A3B-uncensored-heretic-mmproj-BF16.gguf \
  --host 127.0.0.1 --port 8765 -c 8192 -ngl 99
# then POST OpenAI-style chat completions with an image_url content block —
# e.g. {"type":"image_url","image_url":{"url":"data:image/jpeg;base64,..."}}

# B. CLI via llama-mtmd-cli (one-shot). It's a separate cmake target, so a
#    selective build can skip it; a plain `cmake --build build` produces it.
llama-mtmd-cli \
  -m Janus-35B-A3B.Q4_K_M.gguf \
  --mmproj Qwen3.6-35B-A3B-uncensored-heretic-mmproj-BF16.gguf \
  --image photo.jpg \
  -p "Describe this image."

# C. Python via llama-cpp-python:
python examples/llama_cpp_vision.py \
  --gguf Janus-35B-A3B.Q4_K_M.gguf \
  --mmproj Qwen3.6-35B-A3B-uncensored-heretic-mmproj-BF16.gguf \
  --image /path/to/photo.jpg \
  --prompt "What is in this image?"

Until the Ollama upstream issue is fixed, treat Ollama as text-only for this model. The bundled Q4_K_M decoder pairs with the projector directly — the mmproj is family-wide for Qwen 3.6 35B-A3B, so no separate text download is needed for vision.

Hardware requirements

This is a ~19 GB Q4_K_M GGUF. Ollama's runtime footprint is roughly 2× the model file (weights mmap + compute graph), plus a KV cache that scales ~2 GB per 32K (q8_0). The default num_ctx is 1000000 — a ~1M ceiling above the 262144 native window — so KV alone is ~61 GB for ~99 GB total (theoretical, extrapolated from the ~2 GB/32K rule). This GGUF ships no YaRN rope-scaling (rope.freq_base 10M, no rope.scaling), so positions past the 262144 native window use untrained RoPE and output degrades — treat 1M as an advertised ceiling and keep real work within ~262K. Most hosts must override num_ctx down: e.g. the 262144 native window → ~16 GB KV / ~53 GB total, or num_ctx 32768 → ~2 GB KV / ~39 GB total. 32 GB hosts fit the model by trimming ctx + batch (see Z13 row in the table).

How to override it: ollama run has no -o flag, and OLLAMA_CONTEXT_LENGTH only sets a default that the baked num_ctx overrides — so set it per-session from the interactive prompt. The model loads lazily on the first message, so /set applies before the default context is allocated:

ollama run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
>>> /set parameter num_ctx 4096
>>> /set parameter num_batch 256

Programmatic callers pass the same via the API options field: "options": {"num_ctx": 4096, "num_batch": 256}.

Hardware Status
≥48 GB RAM (CPU-only) Works, ~3-6 tok/s
Single H100/A100 80 GB Works, full offload, ~30+ tok/s
RTX 4090 24 GB / 5090 32 GB + 32 GB RAM Works, partial offload, ~15-25 tok/s
Mac Studio M2/M3 Ultra 64 GB+ unified Works, ~20+ tok/s
32 GB unified-memory laptops (Ryzen AI Max+, Apple M-series) Works with num_ctx ≤ 4096 and num_batch ≤ 256 to fit the compute graph; the 1M default OOMs (override num_ctx down). Measured 28.71 tok/s on ASUS ROG Flow Z13 GZ302EA at Q4_K_M (Radeon 8060S iGPU via ROCm gfx1151).

Reaching a coherent ~1M context (opt-in YaRN). The bundled GGUF ships no YaRN rope-scaling, so the 1M default degrades past the 262144 native window (see above). Ollama has no rope knob, so for a genuinely coherent long context run the GGUF under llama.cpp with YaRN enabled:

llama-server -m Janus-35B-A3B.Q4_K_M.gguf \
  --rope-scaling yarn --yarn-orig-ctx 262144 --rope-scale 3.8 -c 1000000

--rope-scale 3.8 ≈ 1000000 / 262144; use a smaller factor for a smaller window. Static YaRN rescales all prompts, so enable it only when you actually need > 262K — it slightly degrades short-context quality otherwise.

Chat template

The model uses the standard Qwen 3.x ChatML format with <|im_start|> / <|im_end|> role markers. The template is embedded in the GGUF metadata for plain conversation use, but Ollama users should rely on the TEMPLATE block in the included Modelfile — that version exposes the tool-calling scaffolding Ollama's capability detector requires (the embedded template alone is insufficient; see Ollama above).

Plain conversation

<|im_start|>system
You are Janus, a precise and capable assistant…<|im_end|>
<|im_start|>user
What is the time complexity of mergesort?<|im_end|>
<|im_start|>assistant

With reasoning trace

When the model decides to think, the assistant turn contains a <think>…</think> block followed by the visible answer:

<|im_start|>assistant
<think>
The user is asking about mergesort. Mergesort divides the array, recursively sorts each half, then merges. The recurrence T(n) = 2T(n/2) + O(n) solves to O(n log n).
</think>

Mergesort runs in **O(n log n)** time in the worst, average, and best cases. The recurrence is T(n) = 2T(n/2) + O(n), which solves to Θ(n log n) by the master theorem.<|im_end|>

Most clients (Open WebUI, LibreChat, etc.) hide the <think> block by default and show only the final answer. If your client doesn't, set its "show reasoning" toggle off.

Disabling thinking

This is a reasoning-first model — it opens a <think> block by default. For a direct answer with no reasoning trace (simple or latency-sensitive calls), turn thinking off:

ollama run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M --think=false

or send "think": false on /api/chat. With thinking off the model skips the reasoning trace and answers straight into content; with it on (the default) reasoning is emitted into the thinking field.

Tool / function calling

The wire format depends on which path you take. Both are valid — the model adapts to whichever format the system prompt specifies.

Ollama path (this repo's Modelfile). The TEMPLATE advertises tools inside <tools>…</tools> and asks the model to reply in JSON-in-XML — the form Ollama's tool-call extractor parses into a structured tool_calls array on /api/chat and /v1/chat/completions:

<tool_call>
{"name": "get_weather", "arguments": {"city": "Tokyo"}}
</tool_call>

Embedded-jinja path (llama.cpp, llama-cpp-python, LM Studio). The Qwen 3.6 native chat template baked into the GGUF instructs the model to emit a more verbose XML form. This is the shape you'll see if you talk to llama-server or LM Studio directly:

<tool_call>
<function=get_weather>
<parameter=city>
Tokyo
</parameter>
</function>
</tool_call>

Pick the parser shape that matches your loader. Don't mix.

Example (Ollama, OpenAI-compatible API)

from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ignored")

resp = client.chat.completions.create(
    model="janus",
    messages=[
        {"role": "user", "content": "Call get_weather for Tokyo. Respond ONLY with the tool call."}
    ],
    tools=[{
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get current weather for a city",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
                "required": ["city"],
            },
        },
    }],
    temperature=0.3,
)
print(resp.choices[0].message.tool_calls)
# [ToolCall(id='call_xxx', type='function',
#           function=Function(name='get_weather', arguments='{"city":"Tokyo"}'))]

Tips

  • Use direct prompts ("Call X for Y") rather than soft hints ("Use the tool"). The model thinks before committing to a call, and weak prompts can exhaust num_predict inside the <think> block before the call is emitted.
  • Allow at least num_predict: 1024 (or max_tokens: 1024) for tool-calling turns, more if the schemas are large.
  • The Modelfile's JSON-in-XML format is what Ollama's tool-call extractor understands; if you swap loaders, swap the parser to match (see "Embedded-jinja path" above).

Known limitations

  • No mmproj in this release. The base Qwen3.6 supports image and video input via a separate mmproj file, which is not included here. Text-only inference works out of the box; multimodal inference requires fetching Qwen3.6-35B-A3B-uncensored-heretic-mmproj-BF16.gguf (or equivalent) from upstream — run ./scripts/fetch_vision.sh and see Vision for the full path.
  • Quantization-induced quality loss. Q4_K_M is a strong general-purpose quant but does measurably degrade math and code accuracy compared to BF16. If you need maximum quality, run the upstream safetensors on a GPU that fits BF16 (~70 GB).
  • MoE expert utilization is uneven. Stock Qwen3.6-35B-A3B routes 8 of 256 experts per token. On narrow domains (e.g. only one programming language) a small subset of experts dominates; load-balance loss was a training-time concern, not a runtime guarantee.
  • Thinking traces can loop. Like most reasoning-distilled models, Janus-35B occasionally gets stuck repeating itself inside <think> tags. Mitigations: lower temperature to 0.4-0.6, raise repeat_penalty to 1.08, or set a <think>-token budget cap if your loader supports it.
  • Large tool-call arguments can be dropped. Ollama's JSON-in-XML tool format makes the model JSON-escape the entire arguments object inline; for a big/complex payload (e.g. a file's content in a write_file call) the model can fail to escape it, so the field arrives undefined and the call fails. Qwen's native <function=…><parameter=…> format (raw values, no escaping) was tested as a fix but parses unreliably through Ollama, so the template deliberately keeps JSON-in-XML. Mitigation: write large files in smaller pieces per call.
  • Uncensored base — not aligned with any specific safety policy. This is a personal repackage of an open-weight base whose refusal behavior has been abliterated away (the llmfan46 Heretic base). There is no RLHF refusal layer; the model will attempt most requests, so downstream safety is entirely the operator's responsibility.
  • No formal evaluation in this card. Numbers in the hardware table are estimates, not measured. If you produce real benchmarks (MMLU, HumanEval, etc.) and want them included, file a PR.

Related models

Model Size Notes
llmfan46/Qwen3.6-35B-A3B-uncensored-heretic 35B / 3B active Immediate base. Uncensored Heretic abliteration of Qwen 3.6 35B-A3B; transformers-native safetensors.
llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-GGUF 35B / 3B active Heretic GGUFs — pull other quants here; the bundled Q4_K_M is already this Heretic quant.
llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved 35B / 3B active Same Heretic base but keeps the MTP head for vLLM / SGLang speculative decoding.
Qwen/Qwen3.6-35B-A3B 35B / 3B active Upstream pre-Heretic base model. transformers-native multimodal weights.
FoolDev/Thanatos-27B-HERETIC 27B dense Dense sibling on the llmfan46/Qwen3.6-27B-uncensored-heretic-v2 Heretic base. Same teacher (Fable 5), same dataset family, smaller memory footprint, no MoE quirks. (The older FoolDev/Thanatos-27B and Thanatos-27B-Heretic slugs now 307 to this path.)
Crownelius/Crow-9B-HERETIC-4.6 9B dense Heretic-flavored fine-tune on a smaller 9B Qwen base. Useful as a fast first-pass model when 35B is too heavy for the host.

Credits

License inherited from upstream: Apache-2.0.

Downloads last month
216
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FoolDev/Janus-35B-HERETIC

Finetuned
(2)
this model

Datasets used to train FoolDev/Janus-35B-HERETIC

Space using FoolDev/Janus-35B-HERETIC 1