Text Generation
Transformers
Safetensors
GGUF
English
causal-lm
qwen2.5
reasoning
code-generation
Mixture of Experts
qlora
multimodal
tool-use
Eval Results (legacy)
conversational
Instructions to use ram1234598766/Cesium2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ram1234598766/Cesium2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ram1234598766/Cesium2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ram1234598766/Cesium2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ram1234598766/Cesium2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: llama cli -hf ram1234598766/Cesium2:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: llama cli -hf ram1234598766/Cesium2:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf ram1234598766/Cesium2:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ram1234598766/Cesium2:Q8_0
Use Docker
docker model run hf.co/ram1234598766/Cesium2:Q8_0
- LM Studio
- Jan
- vLLM
How to use ram1234598766/Cesium2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ram1234598766/Cesium2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ram1234598766/Cesium2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ram1234598766/Cesium2:Q8_0
- SGLang
How to use ram1234598766/Cesium2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ram1234598766/Cesium2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ram1234598766/Cesium2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ram1234598766/Cesium2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ram1234598766/Cesium2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use ram1234598766/Cesium2 with Ollama:
ollama run hf.co/ram1234598766/Cesium2:Q8_0
- Unsloth Studio
How to use ram1234598766/Cesium2 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ram1234598766/Cesium2 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ram1234598766/Cesium2 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for ram1234598766/Cesium2 to start chatting
- Pi
How to use ram1234598766/Cesium2 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ram1234598766/Cesium2:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ram1234598766/Cesium2:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ram1234598766/Cesium2 with Docker Model Runner:
docker model run hf.co/ram1234598766/Cesium2:Q8_0
- Lemonade
How to use ram1234598766/Cesium2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ram1234598766/Cesium2:Q8_0
Run and chat with the model
lemonade run user.Cesium2-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use ram1234598766/Cesium2 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ram1234598766/Cesium2:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ram1234598766/Cesium2:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ram1234598766/Cesium2 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ram1234598766/Cesium2:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ram1234598766/Cesium2:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 5,030 Bytes
82f262a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 | """
RegexFeatureExtractor - token feature gating.
Extends build_code_features (4-dim) to a 7-dim per-token feature vector plus
a running SyntaxState (bracket stack, quote open/closed) used as a cheap
pre-model sanity gate.
Dimensions per token:
[0] is_code_like (indent / brackets / operators / newlines)
[1] indent_depth (normalized leading whitespace)
[2] bracket_balance (1.0 open, 0.5 neutral, 0.0 close)
[3] has_newline
[4] keyword_hit (NEW)
[5] quote_state (running open/closed string literal) (NEW)
[6] numeric_literal (NEW)
"""
import re
import torch
from dataclasses import dataclass, field
from typing import List, Optional, Tuple
CODE_CHARS = set("{}()[];=<>!&|+-*/%'\"`#@.,:")
KEYWORD_RE = re.compile(
r"\b(def|class|import|return|if|else|elif|for|while|try|except|finally|"
r"lambda|pass|with|as|yield|from|async|await)\b"
)
NUMERIC_RE = re.compile(r"\b\d+(\.\d+)?\b")
OPEN_BRACKETS = {"{": 1, "[": 1, "(": 1}
CLOSE_BRACKETS = {"}": 1, "]": 1, ")": 1}
@dataclass
class SyntaxState:
bracket_stack: List[str] = field(default_factory=list)
quote_open: Optional[str] = None # '"' or "'" while a string literal is open
depth: int = 0
def is_balanced(self) -> bool:
return not self.bracket_stack and self.quote_open is None
def to_dict(self) -> dict:
return {
"balanced": self.is_balanced(),
"bracket_depth": len(self.bracket_stack),
"quote_open": self.quote_open,
}
class RegexFeatureExtractor:
def __init__(self, num_features: int = 7):
self.num_features = num_features
def extract(self, tokenizer, input_ids: torch.Tensor) -> Tuple[torch.Tensor, List[SyntaxState]]:
"""Per-token features (B, T, F) + one SyntaxState per row."""
feats: List[List[List[float]]] = []
states: List[SyntaxState] = []
for row in input_ids.tolist():
tokens = tokenizer.convert_ids_to_tokens(row)
state = SyntaxState()
row_feats = []
for tok in tokens:
is_code = any(c in CODE_CHARS for c in tok)
indent = 0.0
stripped = tok.lstrip()
if stripped and tok != stripped:
indent = min((len(tok) - len(stripped)) / 8.0, 1.0)
is_code = True
bal = 0.5
for c in tok:
if c in OPEN_BRACKETS:
bal = 1.0
state.bracket_stack.append(c)
elif c in CLOSE_BRACKETS:
bal = 0.0
if state.bracket_stack:
state.bracket_stack.pop()
# running quote state
for c in tok:
if c in ('"', "'"):
if state.quote_open is None:
state.quote_open = c
elif state.quote_open == c:
state.quote_open = None
quote = 1.0 if state.quote_open is not None else 0.0
newline = 1.0 if "\n" in tok else 0.0
kw = 1.0 if KEYWORD_RE.search(tok) else 0.0
num = 1.0 if NUMERIC_RE.search(tok) else 0.0
row_feats.append([1.0 if is_code else 0.0, indent, bal, newline, kw, quote, num])
state.depth = len(state.bracket_stack)
row_feats = row_feats[: input_ids.shape[1]]
feats.append(row_feats)
states.append(state)
max_len = max(len(r) for r in feats)
padded = [
r + [[0.0, 0.0, 0.5, 0.0, 0.0, 0.0, 0.0]] * (max_len - len(r))
for r in feats
]
t = torch.tensor(padded, dtype=torch.float32)
if t.shape[-1] > self.num_features:
t = t[..., : self.num_features]
return t, states
def extract_text(self, text: str) -> SyntaxState:
"""Run a string-only pass for the pre-model gate (no tokenizer)."""
state = SyntaxState()
for c in text:
if c in OPEN_BRACKETS:
state.bracket_stack.append(c)
elif c in CLOSE_BRACKETS and state.bracket_stack:
state.bracket_stack.pop()
elif c in ('"', "'"):
if state.quote_open is None:
state.quote_open = c
elif state.quote_open == c:
state.quote_open = None
state.depth = len(state.bracket_stack)
return state
def gate(self, syntax_state: SyntaxState) -> str:
"""Returns 'pass' | 'warn' | 'block'."""
if syntax_state.is_balanced():
return "pass"
return "block" if syntax_state.depth > 4 else "warn"
# drop-in replacement for architecture.build_code_features with 7-dim output
def build_code_features_v2(tokenizer, input_ids: torch.Tensor) -> torch.Tensor:
return RegexFeatureExtractor(num_features=7).extract(tokenizer, input_ids)[0] |