Instructions to use CloudGoat/Mephisto-4B-v2.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CloudGoat/Mephisto-4B-v2.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="CloudGoat/Mephisto-4B-v2.1") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("CloudGoat/Mephisto-4B-v2.1") model = AutoModelForMultimodalLM.from_pretrained("CloudGoat/Mephisto-4B-v2.1", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use CloudGoat/Mephisto-4B-v2.1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf CloudGoat/Mephisto-4B-v2.1:Q8_0 # Run inference directly in the terminal: llama cli -hf CloudGoat/Mephisto-4B-v2.1:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf CloudGoat/Mephisto-4B-v2.1:Q8_0 # Run inference directly in the terminal: llama cli -hf CloudGoat/Mephisto-4B-v2.1:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf CloudGoat/Mephisto-4B-v2.1:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf CloudGoat/Mephisto-4B-v2.1:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf CloudGoat/Mephisto-4B-v2.1:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf CloudGoat/Mephisto-4B-v2.1:Q8_0
Use Docker
docker model run hf.co/CloudGoat/Mephisto-4B-v2.1:Q8_0
- LM Studio
- Jan
- vLLM
How to use CloudGoat/Mephisto-4B-v2.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "CloudGoat/Mephisto-4B-v2.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CloudGoat/Mephisto-4B-v2.1", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/CloudGoat/Mephisto-4B-v2.1:Q8_0
- SGLang
How to use CloudGoat/Mephisto-4B-v2.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "CloudGoat/Mephisto-4B-v2.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CloudGoat/Mephisto-4B-v2.1", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "CloudGoat/Mephisto-4B-v2.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CloudGoat/Mephisto-4B-v2.1", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Ollama
How to use CloudGoat/Mephisto-4B-v2.1 with Ollama:
ollama run hf.co/CloudGoat/Mephisto-4B-v2.1:Q8_0
- Unsloth Studio
How to use CloudGoat/Mephisto-4B-v2.1 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for CloudGoat/Mephisto-4B-v2.1 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for CloudGoat/Mephisto-4B-v2.1 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for CloudGoat/Mephisto-4B-v2.1 to start chatting
- Pi
How to use CloudGoat/Mephisto-4B-v2.1 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf CloudGoat/Mephisto-4B-v2.1:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "CloudGoat/Mephisto-4B-v2.1:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use CloudGoat/Mephisto-4B-v2.1 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf CloudGoat/Mephisto-4B-v2.1:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "CloudGoat/Mephisto-4B-v2.1:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use CloudGoat/Mephisto-4B-v2.1 with Docker Model Runner:
docker model run hf.co/CloudGoat/Mephisto-4B-v2.1:Q8_0
- Lemonade
How to use CloudGoat/Mephisto-4B-v2.1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull CloudGoat/Mephisto-4B-v2.1:Q8_0
Run and chat with the model
lemonade run user.Mephisto-4B-v2.1-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use CloudGoat/Mephisto-4B-v2.1 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf CloudGoat/Mephisto-4B-v2.1:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default CloudGoat/Mephisto-4B-v2.1:Q8_0
Run Hermes
hermes
- Atomic Chat
Mephisto-4B-v2.1
Abstract
Mephisto-4B-v2.1 is a 4B-parameter agentic language model constructed via multi-stage model merging using mergekit. The model integrates reasoning, coding, and agentic capabilities into a single 4B checkpoint based on the Qwen3.5-4B architecture, targeting performance exceeding Mephisto-4B-v2 on Japanese and agentic benchmarks while maintaining deployment feasibility on consumer hardware (RTX 3060 12GB).
Methodology
Merge Pipeline
The model is constructed through a three-stage merging pipeline:
Stage 1: Reasoning-Code Fusion (NuSLERP)
Normalized spherical linear interpolation (NuSLERP) on task vectors derived from a common base model (Qwen3.5-4B). This stage fuses a reasoning-distilled model (Opus-reasoning distillation) with a code-specialized model.
| Component | Source Model | Weight |
|---|---|---|
| Reasoning | Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled-v2 | 1.0 |
| Code | Jackrong/Qwopus3.5-4B-Coder | 0.7 |
Configuration:
merge_method: nuslerp
base_model: Qwen/Qwen3.5-4B
models:
- model: Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled-v2
parameters: {weight: 1.0}
- model: Jackrong/Qwopus3.5-4B-Coder
parameters: {weight: 0.7}
parameters:
nuslerp_flatten: true
nuslerp_row_wise: false
dtype: bfloat16
Stage 2: Agent Foundation (DARE-TIES 3-way)
DARE-TIES merging with density-based pruning and sign consensus. This stage integrates three complementary agentic models into a unified foundation.
| Component | Source Model | Weight | Density |
|---|---|---|---|
| Reasoning | BAAI/AREX-Turbo | 1.0 | 0.5 |
| Agent/Planning | InternScience/Agents-A1-4B | 1.0 | 0.5 |
| General/Japanese | Jackrong/Qwopus3.5-4B-v3 | 0.6 | 0.5 |
Configuration:
merge_method: dare_ties
base_model: Qwen/Qwen3.5-4B
models:
- model: BAAI/AREX-Turbo
parameters: {weight: 1.0, density: 0.5}
- model: InternScience/Agents-A1-4B
parameters: {weight: 1.0, density: 0.5}
- model: Jackrong/Qwopus3.5-4B-v3
parameters: {weight: 0.6, density: 0.5}
parameters:
lambda: 0.8
density: 0.5
dtype: bfloat16
Stage 3: Final Integration (FrankenMerge / Passthrough)
Layer-wise model selection (FrankenMerge) via Passthrough merge. This stage structurally allocates layers by functional specialization:
| Layer Range | Source | Functional Role |
|---|---|---|
| 0–15 | Qwen/Qwen3.5-4B (Instruct) | Embeddings, shallow features, Japanese language modeling |
| 16–23 | Stage 2 Output (Agent Foundation) | Planning, function calling, agentic reasoning |
| 24–31 | Stage 1 Output (Reasoning + Code) | Deep reasoning, code generation, complex logic |
Configuration:
merge_method: passthrough
slices:
- sources:
- model: Qwen/Qwen3.5-4B
layer_range: [0, 16]
- sources:
- model: merged/s2
layer_range: [16, 24]
- sources:
- model: merged/s1
layer_range: [24, 32]
Source Models
| Model | Role | Type | License |
|---|---|---|---|
| Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled-v2 | Reasoning distillation | Instruct | Apache-2.0 |
| Jackrong/Qwopus3.5-4B-Coder | Code / Tool Use | Instruct | Apache-2.0 |
| BAAI/AREX-Turbo | General reasoning | Base/Instruct | Apache-2.0 |
| InternScience/Agents-A1-4B | Agent / Planning / FC | Instruct | Apache-2.0 |
| Jackrong/Qwopus3.5-4B-v3 | Japanese / General | Instruct | Apache-2.0 |
| Qwen/Qwen3.5-4B | Base anchor (Instruct) | Instruct | Apache-2.0 |
All source models are Apache-2.0 licensed, enabling commercial use of the merged artifact.
Benchmarks (In Progress)
| Benchmark | Status | Category |
|---|---|---|
| ELYZA-tasks-100 | In Progress | Japanese instruction following |
| MT-Bench-JP | In Progress | Japanese conversation quality |
| JGLUE | In Progress | Japanese NLU |
| JMMLU | In Progress | Japanese knowledge |
| HumanEval | In Progress | Code generation |
| BFCL (Function Calling) | In Progress | Function calling |
| JSB (Safety) | In Progress | Safety alignment |
Benchmark evaluation is currently in progress. Target thresholds are based on Mephisto-4B-v2 baseline and merge methodology expectations.
Quantization
| Format | Size | Command |
|---|---|---|
| Q8_0 | 4.5 GB | llama.cpp convert_hf_to_gguf.py --outtype q8_0 --no-mtp |
| Q4_K_M | ~2.6 GB | llama.cpp convert_hf_to_gguf.py --outtype q4_k_m --no-mtp |
Note: The --no-mtp flag is required to exclude Multi-Token Prediction (MTP) modules present in Qwen3.5-series models, which are incompatible with llama.cpp inference.
Usage
HF Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Mephisto-4B-v2.1", torch_dtype="bfloat16", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("Mephisto-4B-v2.1")
llama.cpp (Recommended)
./llama-cli -m Mephisto-4B-v2.1_q8_0.gguf -ngl 99 -c 4096 -p "ユーザー: こんにちは
アシスタント: "
Hardware Requirements
| Component | Specification |
|---|---|
| GPU | RTX 3060 12GB+ (for Q8_0 with -ngl 99) |
| RAM | 16 GB+ system RAM |
| Disk | ~12 GB (HF + GGUF) |
License
This model inherits licenses from its source models. All source models are Apache-2.0 licensed, permitting commercial use. Users should verify individual source model licenses for their specific use cases.
Citation
If you use this model in your research, please cite:
@misc{mephisto-4b-v2.1,
title={Mephisto-4B-v2.1: Multi-Stage Merged Agentic Model},
author={CloudGoat},
year={2024},
note={Constructed via mergekit multi-stage merging: NuSLERP + DARE-TIES + FrankenMerge}
}
Model Card Version: 1.0
Last Updated: 2024-08-05
Merge Tool: mergekit (NuSLERP + DARE-TIES + Passthrough)
Base Architecture: Qwen3.5-4B (32 layers, 2560 hidden dim, 16/4 attention heads)
- Downloads last month
- 66