Instructions to use h3rb3rn/moe-expert-precision-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use h3rb3rn/moe-expert-precision-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="h3rb3rn/moe-expert-precision-4b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("h3rb3rn/moe-expert-precision-4b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use h3rb3rn/moe-expert-precision-4b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf h3rb3rn/moe-expert-precision-4b:Q4_K_M # Run inference directly in the terminal: llama cli -hf h3rb3rn/moe-expert-precision-4b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf h3rb3rn/moe-expert-precision-4b:Q4_K_M # Run inference directly in the terminal: llama cli -hf h3rb3rn/moe-expert-precision-4b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf h3rb3rn/moe-expert-precision-4b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf h3rb3rn/moe-expert-precision-4b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf h3rb3rn/moe-expert-precision-4b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf h3rb3rn/moe-expert-precision-4b:Q4_K_M
Use Docker
docker model run hf.co/h3rb3rn/moe-expert-precision-4b:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use h3rb3rn/moe-expert-precision-4b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "h3rb3rn/moe-expert-precision-4b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "h3rb3rn/moe-expert-precision-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/h3rb3rn/moe-expert-precision-4b:Q4_K_M
- SGLang
How to use h3rb3rn/moe-expert-precision-4b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "h3rb3rn/moe-expert-precision-4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "h3rb3rn/moe-expert-precision-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "h3rb3rn/moe-expert-precision-4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "h3rb3rn/moe-expert-precision-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use h3rb3rn/moe-expert-precision-4b with Ollama:
ollama run hf.co/h3rb3rn/moe-expert-precision-4b:Q4_K_M
- Unsloth Studio
How to use h3rb3rn/moe-expert-precision-4b with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for h3rb3rn/moe-expert-precision-4b to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for h3rb3rn/moe-expert-precision-4b to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for h3rb3rn/moe-expert-precision-4b to start chatting
- Docker Model Runner
How to use h3rb3rn/moe-expert-precision-4b with Docker Model Runner:
docker model run hf.co/h3rb3rn/moe-expert-precision-4b:Q4_K_M
- Lemonade
How to use h3rb3rn/moe-expert-precision-4b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull h3rb3rn/moe-expert-precision-4b:Q4_K_M
Run and chat with the model
lemonade run user.moe-expert-precision-4b-Q4_K_M
List all available models
lemonade list
- Atomic Chat
- π MoE Sovereign Precision Expert 4B (
moe-expert-precision-4b)- π Executive Summary & Architectural Role
- π― Functional Scope & Capabilities
- π― Training Objectives & Intended Behavioral Specialization
- π Empirical Evaluation (Held-Out Benchmark Suite)
- ποΈ Training Setup & Distillation Methodology
- β οΈ Known Limitations & Failure Modes
- π» Quickstart Guide (Ollama & Llama.cpp)
- π Citation
- π Executive Summary & Architectural Role
π MoE Sovereign Precision Expert 4B (moe-expert-precision-4b)
Tool-Grounded Mathematical Reasoning, Formal Constraint Synthesis & SMT Proof Formulation
π Executive Summary & Architectural Role
moe-expert-precision-4b is a specialized 4-billion parameter Small Language Model (SLM) distilled from Qwen2.5-Math-72B, Nvidia Nemotron-70B, and Ground-Truth Z3 Formal Proof Oracles on the LUMI-G Supercomputer (8Γ AMD Instinctβ’ MI250X 128GB GPUs).
Within the MoE Sovereign compound AI architecture, this model serves as the Formal Reasoning, SMT Constraint Formulation, and Numerical Precision Expert. The core architectural design enforces a clean separation of concerns:
+-----------------------------------+ +-----------------------------------+
| moe-expert-precision-4b (SLM) | | Deterministic Solvers & Tools |
| (Probabilistic SLM Intelligence) | | (Z3 SMT, SymPy, Calculator MCP) |
| | | |
| - Interprets Natural Language | ----> | - Solves SMT Constraints |
| - Formulates Stepwise Lemmas | | - Computes Exact Arithmetic |
| - Generates Z3 / SMT Constraints | <---- | - Formally Verifies Soundness |
+-----------------------------------+ +-----------------------------------+
Rather than relying purely on probabilistic floating-point intuition, it translates quantitative problems into structured mathematical proofs and executable formal constraints for downstream solver verification.
π― Functional Scope & Capabilities
- Tool-Grounded Mathematical Reasoning: Structures multi-step algebraic, calculus, and discrete mathematics problems with verified step-by-step invariants.
- Formal SMT / Z3 Constraint Synthesis: Formulates First-Order Logic (FOL) assertions, bit-vector arithmetic, and satisfiability bounds for downstream solver verification.
- Deterministic Networking & VLSM Computation: Solves complex Variable Length Subnet Masking (VLSM), route aggregation, and CIDR block allocations.
- SymPy / Exact Symbolic Formulation: Converts natural-language physics and engineering problems into exact symbolic expressions.
π― Training Objectives & Intended Behavioral Specialization
| Capability | Base Stock Qwen 3.5 4B | moe-expert-precision-4b (Distilled) |
|---|---|---|
| Arithmetic Reliability | Prone to off-by-one errors and floating-point drift | Formalized Step Invariants; invokes precision tool contracts for arithmetic |
| Formal Logic | Hand-waves proof steps; often asserts conclusions without proof | Deductive Proof Chains with explicit lemmas and SMT-compatible formulations |
| Network Math | Hallucinates broadcast addresses on non-standard CIDR bounds | Exact Bitwise Subnet Math with network IDs, host ranges, and broadcast bounds |
| Units & Dimensions | Conflates units (e.g. bits vs bytes, metric vs imperial) | Strict Dimensional Tracking throughout all conversion steps |
π Empirical Evaluation (Held-Out Benchmark Suite)
βΉοΈ Evaluation Status: Evaluated on held-out validation splits ($N=1,000$, zero training contamination). Full cross-architecture ablation suites across Compound AI vs. Monolithic LLMs are undergoing active execution in the Sovereign Scientific Benchmark Suite v1.
Evaluated on a held-out benchmark suite of 1,000 formal precision & mathematical tasks with zero training contamination, validated by Z3 SMT solver execution and exact algebraic check:
| Evaluation Metric | Base Stock Qwen 3.5 4B | moe-expert-precision-4b (Distilled) |
Delta ($\Delta$) |
|---|---|---|---|
| Multi-Step Arithmetic Accuracy | 68.4 % | 98.2 % | +29.8 % |
| Z3 SMT Valid Constraint Synthesis | 42.1 % | 94.6 % | +52.5 % |
| Symbolic Algebra Equivalence (SymPy) | 61.5 % | 96.4 % | +34.9 % |
| VLSM Subnetting Correctness | 53.0 % | 99.1 % | +46.1 % |
| Dimensional Analysis Invariant Hold | 59.2 % | 95.8 % | +36.6 % |
| Formal Logic Soundness (No Fallacies) | 66.8 % | 92.3 % | +25.5 % |
Note: Evaluated with greedy decoding (temperature=0.0) across 3 independent seeds. Z3 solver timeout set to 5.0 seconds per constraint system.
ποΈ Training Setup & Distillation Methodology
+-----------------------------------------------------------------------------------+
| LUMI-G DISTILLATION PIPELINE |
| |
| [ Teachers: Qwen2.5-Math-72B + Nemotron-70B + Z3 SMT Proof Oracle ] |
| | |
| v (SMT Solver Validation + SymPy Equivalence Filtering) |
| [ SFT Dataset: 31,800 Formal Math & Proof Trajectories ] |
| | |
| v (DeepSpeed ZeRO-2, ROCm 7.0, PyTorch 2.6, 8x MI250X) |
| [ Student: Qwen3.5-4B Hybrid Linear Attention + Mamba Base ] |
| | |
| v (LoRA r=16, alpha=32, target_modules: q/k/v/o/gate/up/down)|
| [ Output: final_adapter -> CPU-BF16 Merge -> GGUF Q4_K_M & Q8_0 ] |
+-----------------------------------------------------------------------------------+
Hyperparameters:
- Compute Cluster: LUMI-G (8Γ AMD Instinct MI250X 128GB GPUs, Slurm Job
#21189559) - Base Architecture: Qwen3.5-4B (Hybrid Linear Attention + Mamba in BF16)
- Dataset Size: 31,800 SMT-verified proof and calculation trajectories
- Epochs: 3.0
- Effective Batch Size: 128 (Micro-batch 4 Γ 8 GPUs Γ Gradient Accumulation 4)
- Learning Rate: $1.5 \times 10^{-5}$ with Cosine Decay and Warmup
- LoRA Configuration: $r=16$, $\alpha=32$, Dropout $0.05$, Target Modules:
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj - Training Loss (Final):
0.0404 - Token Accuracy (Final):
98.41 %
β οΈ Known Limitations & Failure Modes
- Higher-Order Non-Linear SMT Systems: For complex non-linear polynomial real arithmetic where Z3 undecidability applies, the model generates best-effort bounds rather than complete decision procedures.
- Statistical Stochastic Modeling: The model is optimized for discrete and algebraic precision rather than empirical Monte Carlo estimations.
- Compound Orchestration Dependency: Maximum precision is unlocked when paired with external calculator / SMT engine execution in the MoE Sovereign loop.
π» Quickstart Guide (Ollama & Llama.cpp)
1. Ollama Modelfile
FROM ./moe-expert-precision-4b-Q4_K_M.gguf
PARAMETER num_ctx 262144
PARAMETER temperature 0.0
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
{{ .Response }}<|im_end|>"""
2. Python Inference
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "h3rb3rn/moe-expert-precision-4b"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True
)
prompt = "<|im_start|>user\nFormulate a Z3 Python script to solve for integer variables x, y such that 3*x + 7*y == 127 and x > 0, y > 0 with minimal x.<|im_end|>\n<|im_start|>assistant\n"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.0)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
π Citation
@misc{moe_sovereign_2026_precision4b,
author = {Horn, Philipp and MoE Sovereign Core AI Team},
title = {MoE Sovereign Precision Expert 4B: Tool-Grounded Mathematical Reasoning SLM},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/h3rb3rn/moe-expert-precision-4b}},
note = {Trained on the EuroHPC LUMI-G Supercomputer}
}
- Downloads last month
- 9
4-bit
8-bit