Image-Text-to-Text
Transformers
GGUF
qwen36
Mixture of Experts
conversational
multimodal
agent
heretic
uncensored
Instructions to use FoolDev/Janus-35B-HERETIC with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FoolDev/Janus-35B-HERETIC with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="FoolDev/Janus-35B-HERETIC") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("FoolDev/Janus-35B-HERETIC", device_map="auto") - llama-cpp-python
How to use FoolDev/Janus-35B-HERETIC with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="FoolDev/Janus-35B-HERETIC", filename="Janus-35B-A3B.Q4_K_M.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use FoolDev/Janus-35B-HERETIC with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: llama cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: llama cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf FoolDev/Janus-35B-HERETIC:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Use Docker
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use FoolDev/Janus-35B-HERETIC with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FoolDev/Janus-35B-HERETIC" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FoolDev/Janus-35B-HERETIC", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- SGLang
How to use FoolDev/Janus-35B-HERETIC with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FoolDev/Janus-35B-HERETIC" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FoolDev/Janus-35B-HERETIC", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FoolDev/Janus-35B-HERETIC" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FoolDev/Janus-35B-HERETIC", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Ollama
How to use FoolDev/Janus-35B-HERETIC with Ollama:
ollama run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Unsloth Studio
How to use FoolDev/Janus-35B-HERETIC with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for FoolDev/Janus-35B-HERETIC to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for FoolDev/Janus-35B-HERETIC to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for FoolDev/Janus-35B-HERETIC to start chatting
- Pi
How to use FoolDev/Janus-35B-HERETIC with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "FoolDev/Janus-35B-HERETIC:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use FoolDev/Janus-35B-HERETIC with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default FoolDev/Janus-35B-HERETIC:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use FoolDev/Janus-35B-HERETIC with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FoolDev/Janus-35B-HERETIC:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "FoolDev/Janus-35B-HERETIC:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use FoolDev/Janus-35B-HERETIC with Docker Model Runner:
docker model run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M
- Lemonade
How to use FoolDev/Janus-35B-HERETIC with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull FoolDev/Janus-35B-HERETIC:Q4_K_M
Run and chat with the model
lemonade run user.Janus-35B-HERETIC-Q4_K_M
List all available models
lemonade list
| # Janus-35B — smoke test against a running Ollama daemon. | |
| # | |
| # Verifies: | |
| # 1. The Ollama server is reachable. | |
| # 2. The target model is loaded / loadable. | |
| # 3. The model exposes the `tools` capability (Modelfile TEMPLATE wired). | |
| # 4. A single chat round-trip succeeds and produces non-empty output. | |
| # 5. No chat-template control tokens leak into the response. | |
| # 6. (TOOLS_TEST=1) An end-to-end tool-call round-trip emits a structured | |
| # tool_calls array with the expected name and arguments. Off by default | |
| # because it costs ~5-10 sec of inference; on for comprehensive runs. | |
| # | |
| # Usage: | |
| # ./scripts/smoke_test.sh # fast checks only | |
| # TOOLS_TEST=1 ./scripts/smoke_test.sh # add tool-call round-trip | |
| # MODEL=hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M ./scripts/smoke_test.sh | |
| # HOST=http://localhost:11434 ./scripts/smoke_test.sh | |
| # | |
| # Requires: curl, jq. | |
| set -euo pipefail | |
| MODEL="${MODEL:-janus}" | |
| HOST="${HOST:-http://localhost:11434}" | |
| PROMPT="${PROMPT:-Reply with the single word: OK}" | |
| require() { | |
| if ! command -v "$1" >/dev/null 2>&1; then | |
| echo "[!] missing dependency: $1" >&2; exit 1 | |
| fi | |
| } | |
| require curl | |
| require jq | |
| echo "[*] host: ${HOST}" | |
| echo "[*] model: ${MODEL}" | |
| # ---- 1. Server up? ---------------------------------------------------------- | |
| if ! curl -fsS "${HOST}/api/tags" >/dev/null; then | |
| echo "[!] Ollama not reachable at ${HOST}. Is 'ollama serve' running?" >&2 | |
| exit 1 | |
| fi | |
| echo "[+] server reachable" | |
| # ---- 2. Model present? ------------------------------------------------------ | |
| # Match case-insensitively: Ollama 0.24 normalizes model names at lookup but | |
| # preserves whatever case was first registered on disk (e.g. an earlier | |
| # session may leave a `Janus-35B:latest` manifest dir behind even when a | |
| # build was invoked with TAG=janus). The exact tag the user typed is still | |
| # valid for `ollama run` — the comparison just needs to be case-folded to | |
| # match. | |
| if ! curl -fsS "${HOST}/api/tags" | jq -e --arg m "${MODEL}" '.models[] | select((.name | ascii_downcase) | startswith($m | ascii_downcase))' >/dev/null; then | |
| echo "[!] Model '${MODEL}' not found. Build it first:" >&2 | |
| echo " ./scripts/build.sh # Q4_K_M" >&2 | |
| echo " ./scripts/build.sh Q3_K_M # smaller quant" >&2 | |
| echo " ./scripts/load_bundle.sh # load this repo's qwen35moe bundle" >&2 | |
| exit 1 | |
| fi | |
| echo "[+] model present" | |
| # ---- 3. Capability guard ---------------------------------------------------- | |
| # The Modelfile TEMPLATE must expose .Tools / .ToolCalls so Ollama lists | |
| # `tools` under capabilities. Without it, /api/chat with a tools array returns | |
| # 400 "does not support tools" even though plain chat works. Catches Modelfile | |
| # regressions that strip or break the TEMPLATE. | |
| CAPS="$(curl -fsS "${HOST}/api/show" -H 'Content-Type: application/json' \ | |
| -d "$(jq -n --arg m "${MODEL}" '{name: $m}')" | jq -r '.capabilities[]?')" | |
| if ! grep -qx -- 'tools' <<<"${CAPS}"; then | |
| echo "[!] model missing capability: tools" >&2 | |
| echo " Modelfile likely missing TEMPLATE that references .Tools / .ToolCalls." >&2 | |
| echo "----- present capabilities -----" >&2 | |
| echo "${CAPS:-<none>}" >&2 | |
| echo "--------------------------------" >&2 | |
| exit 1 | |
| fi | |
| echo "[+] capabilities include: tools" | |
| # ---- 4. Round-trip ---------------------------------------------------------- | |
| echo "[*] sending test prompt..." | |
| RESP="$(curl -fsS "${HOST}/api/chat" \ | |
| -H 'Content-Type: application/json' \ | |
| -d "$(jq -n --arg m "${MODEL}" --arg p "${PROMPT}" '{ | |
| model: $m, | |
| messages: [{role:"user", content:$p}], | |
| stream: false | |
| }')" | jq -r '.message.content // empty')" | |
| if [[ -z "${RESP}" ]]; then | |
| echo "[!] empty response from model" >&2 | |
| exit 1 | |
| fi | |
| # Token-leakage guard: if any of the chat-template control tokens show up | |
| # verbatim in the response, the Modelfile stop-token list is broken and the | |
| # model is bleeding past EOS. | |
| LEAKED=() | |
| for tok in '<|im_start|>' '<|im_end|>' '<|endoftext|>'; do | |
| if grep -qF -- "${tok}" <<<"${RESP}"; then | |
| LEAKED+=("${tok}") | |
| fi | |
| done | |
| if (( ${#LEAKED[@]} )); then | |
| echo "[!] response contains raw control tokens: ${LEAKED[*]}" >&2 | |
| echo " Modelfile likely missing PARAMETER stop directives." >&2 | |
| echo "----- model said -----" >&2 | |
| echo "${RESP}" >&2 | |
| echo "----------------------" >&2 | |
| exit 1 | |
| fi | |
| echo "[+] round-trip OK" | |
| echo "----- model said -----" | |
| echo "${RESP}" | |
| echo "----------------------" | |
| # ---- 5. Tool-call round-trip (opt-in via TOOLS_TEST=1) ---------------------- | |
| # | |
| # Capability advertisement (step 3) only checks the TEMPLATE references | |
| # .Tools / .ToolCalls. It does NOT check the model actually emits a parseable | |
| # tool call. A regression in the prompt scaffolding (e.g. the system-prompt | |
| # instructions inside the TEMPLATE going stale) can leave capabilities | |
| # reported correctly but tool calls failing — the assistant prose-describes | |
| # the call instead of emitting <tool_call>{...}</tool_call>. This block sends | |
| # a tools-array request, parses .message.tool_calls, and asserts the shape | |
| # matches. | |
| if [[ "${TOOLS_TEST:-0}" == "1" ]]; then | |
| echo "[*] tool-call round-trip..." | |
| TOOL_RESP="$(curl -fsS "${HOST}/api/chat" \ | |
| -H 'Content-Type: application/json' \ | |
| -d "$(jq -n --arg m "${MODEL}" '{ | |
| model: $m, | |
| messages: [{role:"user", content:"Call get_weather for Tokyo. Respond ONLY with the tool call."}], | |
| tools: [{ | |
| type: "function", | |
| function: { | |
| name: "get_weather", | |
| description: "Get the weather for a city", | |
| parameters: { | |
| type: "object", | |
| properties: {city: {type: "string"}}, | |
| required: ["city"] | |
| } | |
| } | |
| }], | |
| stream: false, | |
| options: {num_predict: 1024, temperature: 0.3} | |
| }')")" | |
| TC_COUNT="$(jq -r '.message.tool_calls // [] | length' <<<"${TOOL_RESP}")" | |
| if [[ "${TC_COUNT}" -lt 1 ]]; then | |
| echo "[!] model did not emit a tool call" >&2 | |
| echo "----- response -----" >&2 | |
| echo "${TOOL_RESP}" | jq . >&2 | |
| echo "--------------------" >&2 | |
| exit 1 | |
| fi | |
| TC_NAME="$(jq -r '.message.tool_calls[0].function.name // empty' <<<"${TOOL_RESP}")" | |
| TC_CITY="$(jq -r '.message.tool_calls[0].function.arguments.city // empty' <<<"${TOOL_RESP}")" | |
| if [[ "${TC_NAME}" != "get_weather" ]]; then | |
| echo "[!] unexpected tool name: '${TC_NAME}' (wanted 'get_weather')" >&2 | |
| exit 1 | |
| fi | |
| if [[ "${TC_CITY,,}" != "tokyo" ]]; then | |
| echo "[!] unexpected city argument: '${TC_CITY}' (wanted 'Tokyo' case-insensitive)" >&2 | |
| exit 1 | |
| fi | |
| echo "[+] tool-call round-trip OK (name=${TC_NAME} city=${TC_CITY})" | |
| fi | |