Instructions to use FoolDev/Janus-35B-HERETIC with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use FoolDev/Janus-35B-HERETIC with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: llama cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: llama cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Use Docker
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use FoolDev/Janus-35B-HERETIC with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FoolDev/Janus-35B-HERETIC" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FoolDev/Janus-35B-HERETIC", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Ollama
How to use FoolDev/Janus-35B-HERETIC with Ollama:
ollama run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Unsloth Desktop
- Pi
How to use FoolDev/Janus-35B-HERETIC with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "FoolDev/Janus-35B-HERETIC:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use FoolDev/Janus-35B-HERETIC with Docker Model Runner:
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Lemonade
How to use FoolDev/Janus-35B-HERETIC with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull FoolDev/Janus-35B-HERETIC:Q4_K_M
Run and chat with the model
lemonade run user.Janus-35B-HERETIC-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use FoolDev/Janus-35B-HERETIC with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default FoolDev/Janus-35B-HERETIC:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use FoolDev/Janus-35B-HERETIC with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "FoolDev/Janus-35B-HERETIC:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
release 0.9.10: check.sh's pyflakes hint named a command that fails on this host
Browse filesThe skip line said `pip install pyflakes`. On a PEP 668 distro (Arch, Debian
12+, Fedora) that is refused outright, and PEP 668's own error message then
steers the reader to pipx -- which cannot satisfy this check at all: it runs
`python3 -m pyflakes`, so pyflakes must be importable by the SYSTEM python3,
and a venv or pipx install leaves the check skipping forever.
The hint now states that constraint and gives both routes that work:
pip install --user --break-system-packages pyflakes # ~/.local, per python version
<package manager> python-pyflakes # e.g. pacman -S python-pyflakes
Also: the card's scripts/check.sh row listed the lint steps incompletely --
naming pyflakes or shellcheck but not both -- and never said either is skipped
when absent, with the tail reporting the skip count rather than a green
all-passed.
With pyflakes present this repo now reports "all 8 checks passed" with nothing
skipped, and pyflakes finds no issues in any tracked python file. The [~] path
was still exercised for this release by hiding the user site directory
(PYTHONNOUSERSITE=1), so the new text is what an affected reader actually sees.
Text-only release; blob untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- CHANGELOG.md +27 -0
- CITATION.cff +1 -1
- README.md +1 -1
- scripts/check.sh +11 -1
|
@@ -8,6 +8,33 @@ track the **tooling and documentation**, not the underlying base model.
|
|
| 8 |
|
| 9 |
## [Unreleased]
|
| 10 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
## [0.9.9] β 2026-09-19
|
| 12 |
|
| 13 |
Second tier of the adversarial audit whose top tier shipped in 0.9.8. No template change,
|
|
|
|
| 8 |
|
| 9 |
## [Unreleased]
|
| 10 |
|
| 11 |
+
## [0.9.10] β 2026-09-19
|
| 12 |
+
|
| 13 |
+
### Fixed
|
| 14 |
+
|
| 15 |
+
- **`check.sh`'s pyflakes hint named a command that fails on the maintainer's own host.**
|
| 16 |
+
It said `pip install pyflakes`, which a PEP 668 distro (Arch, Debian 12+, Fedora) refuses
|
| 17 |
+
outright β and PEP 668's own error message then steers you to `pipx`, which **cannot**
|
| 18 |
+
satisfy this check: it runs `python3 -m pyflakes`, so pyflakes has to be importable by
|
| 19 |
+
the *system* interpreter, and a venv or pipx install leaves the check skipping forever.
|
| 20 |
+
The hint now says that, and gives both working routes: a `--user
|
| 21 |
+
--break-system-packages` install, or the distro package.
|
| 22 |
+
- The card's `scripts/check.sh` row listed the lint steps incompletely (it named
|
| 23 |
+
`pyflakes` or `shellcheck` but not both) and did not mention that either one is skipped
|
| 24 |
+
when absent, with the tail reporting the count rather than a green all-passed.
|
| 25 |
+
|
| 26 |
+
### Notes
|
| 27 |
+
|
| 28 |
+
- With pyflakes present this repo now reports **all 8 checks passed** with nothing
|
| 29 |
+
skipped, and pyflakes finds no issues in any tracked python file. The `[~]` skip path was
|
| 30 |
+
still exercised for this release by hiding the user site directory
|
| 31 |
+
(`PYTHONNOUSERSITE=1`), so the new hint is the text an affected reader actually sees.
|
| 32 |
+
- One gotcha the hint's `--user` route carries, worth knowing before you pick it: it
|
| 33 |
+
installs under `~/.local/lib/python3.X/site-packages`, which is scoped to that python
|
| 34 |
+
version. A distro upgrade to a new python minor silently returns this check to "skipped".
|
| 35 |
+
The distro package does not have that problem.
|
| 36 |
+
- Text-only release; blob untouched, sha unchanged.
|
| 37 |
+
|
| 38 |
## [0.9.9] β 2026-09-19
|
| 39 |
|
| 40 |
Second tier of the adversarial audit whose top tier shipped in 0.9.8. No template change,
|
|
@@ -1,5 +1,5 @@
|
|
| 1 |
cff-version: 1.2.0
|
| 2 |
-
version: 0.9.
|
| 3 |
date-released: "2026-09-19"
|
| 4 |
title: "Janus-35B: A Sparse Distillation Wrapper for llmfan46's Qwen 3.6 35B-A3B Uncensored Heretic"
|
| 5 |
message: "If you use this model card or its accompanying files, please cite as below."
|
|
|
|
| 1 |
cff-version: 1.2.0
|
| 2 |
+
version: 0.9.10
|
| 3 |
date-released: "2026-09-19"
|
| 4 |
title: "Janus-35B: A Sparse Distillation Wrapper for llmfan46's Qwen 3.6 35B-A3B Uncensored Heretic"
|
| 5 |
message: "If you use this model card or its accompanying files, please cite as below."
|
|
@@ -140,7 +140,7 @@ OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_FLASH_ATTENTION=1 ollama serve
|
|
| 140 |
| `template`, `system`, `params` | Used by HF's Ollama bridge when users `ollama run hf.co/FoolDev/Janus-35B-HERETIC` directly. The bridge does **not** read `Modelfile` (see [HF Ollama docs](https://huggingface.co/docs/hub/en/ollama)); it ingests these three root-level files instead. Kept in sync with the `Modelfile`'s `TEMPLATE` / `SYSTEM` / `PARAMETER` directives. |
|
| 141 |
| `scripts/build.sh` | Pulls a GGUF from `llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-GGUF` (default Q4_K_M), runs it through `strip_mtp.py` (a no-op for these quants, kept as a guard), then runs `ollama create janus`. Quants published there: BF16, Q3_K_L, Q3_K_M, Q4_K_M, Q4_K_S, Q5_K_M, Q5_K_S, Q6_K, Q8_0. The bundled Q4_K_M *is* that repo's Q4_K_M; use this to build the others locally. |
|
| 142 |
| `scripts/check_bridge_sync.py` | Run before pushing a `Modelfile` / `template` / `system` / `params` edit to verify the four configurations remain in sync. Exits 0 if in sync, 1 with a per-key diff if not. |
|
| 143 |
-
| `scripts/check.sh` | Local lint: `bash -n`, `shellcheck`, `py_compile`, footgun-grep, `Modelfile`-vs-bridge-files sync and the Go `template` guard (`make check`) |
|
| 144 |
| `scripts/check_go_template.py` | Guards the Go `template`: Ollama's thinking detection, the condition that replays earlier turns' reasoning, the tool round trip, JSON tool signatures, and the reasoning-effort mapping β the xhigh and low arms and the medium/unset default (check 8 in `check.sh`) |
|
| 145 |
| `scripts/live_check.sh` | Live end-to-end checks in an isolated, CPU-only Ollama (own port and model store; it refuses to run if Ollama reports a GPU): template selection, one tool call, string-argument replay, an earlier turn's reasoning replayed and a live tool chain's kept, every `reasoning_effort` value and the effort lines β 17 checks (`make live-check`) |
|
| 146 |
| `scripts/verify_arch.py` | Cross-checks the README Architecture bullets (layer count, hidden size, expert counts and widths, native context, vocab) plus the underlying forward-pass structure (Gated Attention + Gated DeltaNet dims, partial RoPE, full-attention interval) against the bundled GGUF's `qwen35moe` metadata. Run on demand (`python3 scripts/verify_arch.py`; needs `pip install gguf`); reads the 21 GB GGUF (LFS smudge required β on an un-smudged pointer it now says so and exits 2) and exits non-zero on any mismatch. `block_count` is checked against **40**, this model's natural depth β the MTP-Preserved variant this repo does not ship reports 41. |
|
|
|
|
| 140 |
| `template`, `system`, `params` | Used by HF's Ollama bridge when users `ollama run hf.co/FoolDev/Janus-35B-HERETIC` directly. The bridge does **not** read `Modelfile` (see [HF Ollama docs](https://huggingface.co/docs/hub/en/ollama)); it ingests these three root-level files instead. Kept in sync with the `Modelfile`'s `TEMPLATE` / `SYSTEM` / `PARAMETER` directives. |
|
| 141 |
| `scripts/build.sh` | Pulls a GGUF from `llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-GGUF` (default Q4_K_M), runs it through `strip_mtp.py` (a no-op for these quants, kept as a guard), then runs `ollama create janus`. Quants published there: BF16, Q3_K_L, Q3_K_M, Q4_K_M, Q4_K_S, Q5_K_M, Q5_K_S, Q6_K, Q8_0. The bundled Q4_K_M *is* that repo's Q4_K_M; use this to build the others locally. |
|
| 142 |
| `scripts/check_bridge_sync.py` | Run before pushing a `Modelfile` / `template` / `system` / `params` edit to verify the four configurations remain in sync. Exits 0 if in sync, 1 with a per-key diff if not. |
|
| 143 |
+
| `scripts/check.sh` | Local lint: `bash -n`, `shellcheck`, `pyflakes`, `py_compile`, footgun-grep, `Modelfile`-vs-bridge-files sync and the Go `template` guard (`make check`). `shellcheck` and `pyflakes` are skipped with a `[~]` line if absent, and the tail reports how many were skipped rather than a green all-passed. |
|
| 144 |
| `scripts/check_go_template.py` | Guards the Go `template`: Ollama's thinking detection, the condition that replays earlier turns' reasoning, the tool round trip, JSON tool signatures, and the reasoning-effort mapping β the xhigh and low arms and the medium/unset default (check 8 in `check.sh`) |
|
| 145 |
| `scripts/live_check.sh` | Live end-to-end checks in an isolated, CPU-only Ollama (own port and model store; it refuses to run if Ollama reports a GPU): template selection, one tool call, string-argument replay, an earlier turn's reasoning replayed and a live tool chain's kept, every `reasoning_effort` value and the effort lines β 17 checks (`make live-check`) |
|
| 146 |
| `scripts/verify_arch.py` | Cross-checks the README Architecture bullets (layer count, hidden size, expert counts and widths, native context, vocab) plus the underlying forward-pass structure (Gated Attention + Gated DeltaNet dims, partial RoPE, full-attention interval) against the bundled GGUF's `qwen35moe` metadata. Run on demand (`python3 scripts/verify_arch.py`; needs `pip install gguf`); reads the 21 GB GGUF (LFS smudge required β on an un-smudged pointer it now says so and exits 2) and exits non-zero on any mismatch. `block_count` is checked against **40**, this model's natural depth β the MTP-Preserved variant this repo does not ship reports 41. |
|
|
@@ -97,7 +97,17 @@ if (( ${#PY_FILES[@]} )); then
|
|
| 97 |
fi
|
| 98 |
done
|
| 99 |
else
|
| 100 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 101 |
SKIPPED=$((SKIPPED + 1))
|
| 102 |
fi
|
| 103 |
fi
|
|
|
|
| 97 |
fi
|
| 98 |
done
|
| 99 |
else
|
| 100 |
+
# The hint matters more than it looks: this check runs `python3 -m pyflakes`,
|
| 101 |
+
# so pyflakes has to be importable by the SYSTEM python3. A venv or a pipx
|
| 102 |
+
# install - which is exactly what PEP 668's own error message suggests -
|
| 103 |
+
# leaves this check skipping forever. And on a PEP 668 distro (Arch,
|
| 104 |
+
# Debian 12+, Fedora) the plain pip command below is refused outright.
|
| 105 |
+
yellow "[~] pyflakes not installed (skip). This check runs 'python3 -m pyflakes',"
|
| 106 |
+
yellow " so it must be importable by the SYSTEM python3 - a venv or pipx install"
|
| 107 |
+
yellow " does not satisfy it. Install with: pip install pyflakes"
|
| 108 |
+
yellow " PEP 668 distro (Arch, Debian 12+, Fedora), either:"
|
| 109 |
+
yellow " pip install --user --break-system-packages pyflakes # ~/.local, per python version"
|
| 110 |
+
yellow " <your package manager> python-pyflakes # e.g. pacman -S python-pyflakes"
|
| 111 |
SKIPPED=$((SKIPPED + 1))
|
| 112 |
fi
|
| 113 |
fi
|