Text Generation
Transformers
Safetensors
GGUF
English
causal-lm
qwen2.5
reasoning
code-generation
Mixture of Experts
qlora
multimodal
tool-use
Eval Results (legacy)
conversational
Instructions to use ram1234598766/Cesium2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ram1234598766/Cesium2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ram1234598766/Cesium2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ram1234598766/Cesium2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ram1234598766/Cesium2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: llama cli -hf ram1234598766/Cesium2:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: llama cli -hf ram1234598766/Cesium2:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf ram1234598766/Cesium2:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ram1234598766/Cesium2:Q8_0
Use Docker
docker model run hf.co/ram1234598766/Cesium2:Q8_0
- LM Studio
- Jan
- vLLM
How to use ram1234598766/Cesium2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ram1234598766/Cesium2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ram1234598766/Cesium2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ram1234598766/Cesium2:Q8_0
- SGLang
How to use ram1234598766/Cesium2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ram1234598766/Cesium2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ram1234598766/Cesium2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ram1234598766/Cesium2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ram1234598766/Cesium2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use ram1234598766/Cesium2 with Ollama:
ollama run hf.co/ram1234598766/Cesium2:Q8_0
- Unsloth Studio
How to use ram1234598766/Cesium2 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ram1234598766/Cesium2 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ram1234598766/Cesium2 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for ram1234598766/Cesium2 to start chatting
- Pi
How to use ram1234598766/Cesium2 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ram1234598766/Cesium2:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ram1234598766/Cesium2:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ram1234598766/Cesium2 with Docker Model Runner:
docker model run hf.co/ram1234598766/Cesium2:Q8_0
- Lemonade
How to use ram1234598766/Cesium2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ram1234598766/Cesium2:Q8_0
Run and chat with the model
lemonade run user.Cesium2-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use ram1234598766/Cesium2 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ram1234598766/Cesium2:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ram1234598766/Cesium2:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ram1234598766/Cesium2 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ram1234598766/Cesium2:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ram1234598766/Cesium2:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| """ | |
| VisionAnalyzer - multimodal visual processing for MORPH-AI. | |
| Lazy-loads a Vision Transformer (ViT) + object-detection model when | |
| transformers provides them; otherwise falls back to pure pixel statistics so | |
| visual facts (dominant colors, brightness, edge density, saliency regions) | |
| are still produced with zero model downloads. | |
| Output is a structured ImageFacts object that flows into the FSM VISION | |
| state, feeds the RAG context, and is cross-examined by the verifier. | |
| """ | |
| import io | |
| from dataclasses import dataclass, field | |
| from typing import Any, Dict, List, Optional | |
| import torch | |
| class ImageFacts: | |
| width: int = 0 | |
| height: int = 0 | |
| dominant_colors: List[tuple] = field(default_factory=list) | |
| brightness: float = 0.0 | |
| edge_density: float = 0.0 | |
| saliency_regions: List[Dict[str, Any]] = field(default_factory=list) | |
| objects: List[Dict[str, Any]] = field(default_factory=list) | |
| caption: str = "" | |
| embedding: Optional[torch.Tensor] = None # (patches+1, dim) or None | |
| def to_text(self) -> str: | |
| lines = [f"image {self.width}x{self.height}", f"brightness {self.brightness:.2f}"] | |
| if self.dominant_colors: | |
| lines.append("colors: " + ", ".join( | |
| f"#{r:02x}{g:02x}{b:02x}" for r, g, b in self.dominant_colors[:4] | |
| )) | |
| if self.objects: | |
| lines.append("objects: " + ", ".join( | |
| f"{o.get('label', 'obj')} ({o.get('conf', 0):.2f})" for o in self.objects | |
| )) | |
| if self.saliency_regions: | |
| lines.append("regions: " + ", ".join( | |
| f"{r['x']},{r['y']}" for r in self.saliency_regions[:6] | |
| )) | |
| if self.caption: | |
| lines.append(f"caption: {self.caption}") | |
| return " | ".join(lines) | |
| def to_dict(self) -> Dict[str, Any]: | |
| return { | |
| "width": self.width, | |
| "height": self.height, | |
| "dominant_colors": [list(c) for c in self.dominant_colors], | |
| "brightness": self.brightness, | |
| "edge_density": self.edge_density, | |
| "saliency_regions": self.saliency_regions, | |
| "objects": self.objects, | |
| "caption": self.caption, | |
| } | |
| def _load_pil(): | |
| try: | |
| from PIL import Image | |
| return Image | |
| except ImportError: | |
| return None | |
| class VisionAnalyzer: | |
| def __init__(self, device: Optional[str] = None, use_vit: bool = True, use_detector: bool = True): | |
| self.device = device or ("cuda" if torch.cuda.is_available() else "cpu") | |
| self.vit = None | |
| self.processor = None | |
| self.detector = None | |
| self.det_processor = None | |
| self.use_vit = use_vit | |
| self.use_detector = use_detector | |
| self._load_models() | |
| def _load_models(self): | |
| try: | |
| from transformers import ( | |
| AutoImageProcessor, | |
| AutoModelForObjectDetection, | |
| ViTModel, | |
| ) | |
| if self.use_vit: | |
| self.vit = ViTModel.from_pretrained("google/vit-base-patch16-224-in21k") | |
| self.vit = self.vit.to(self.device).eval() | |
| if self.use_detector: | |
| self.det_processor = AutoImageProcessor.from_pretrained( | |
| "hustvl/yolos-small" | |
| ) | |
| self.detector = AutoModelForObjectDetection.from_pretrained( | |
| "hustvl/yolos-small" | |
| ) | |
| self.detector = self.detector.to(self.device).eval() | |
| except Exception: | |
| self.vit = None | |
| self.detector = None | |
| self.processor = None | |
| def load_image(self, source) -> Any: | |
| """Accept a path, file-like, or bytes; returns PIL Image or None.""" | |
| Image = _load_pil() | |
| if Image is None: | |
| return None | |
| try: | |
| if isinstance(source, (str,)): | |
| return Image.open(source).convert("RGB") | |
| if isinstance(source, bytes): | |
| return Image.open(io.BytesIO(source)).convert("RGB") | |
| if hasattr(source, "read"): | |
| return Image.open(source).convert("RGB") | |
| return source | |
| except Exception: | |
| return None | |
| def _pixel_facts(self, img) -> ImageFacts: | |
| Image = _load_pil() | |
| facts = ImageFacts(width=img.width, height=img.height) | |
| small = img.resize((32, 32)) | |
| px = list(small.getdata()) | |
| n = len(px) | |
| r_sum = g_sum = b_sum = 0 | |
| color_hist: Dict[tuple, int] = {} | |
| for r, g, b in px: | |
| r_sum += r | |
| g_sum += g | |
| b_sum += b | |
| key = (r // 32 * 32, g // 32 * 32, b // 32 * 32) | |
| color_hist[key] = color_hist.get(key, 0) + 1 | |
| facts.brightness = (r_sum + g_sum + b_sum) / (3.0 * n) / 255.0 | |
| facts.dominant_colors = [ | |
| (r + 16, g + 16, b + 16) for (r, g, b), _ in | |
| sorted(color_hist.items(), key=lambda kv: -kv[1])[:4] | |
| ] | |
| # saliency regions: brightest / highest-variance 8x8 cells | |
| import statistics | |
| grid = small.resize((16, 16)) | |
| gx = list(grid.getdata()) | |
| variances = [] | |
| for i in range(16): | |
| for j in range(16): | |
| idx = i * 16 + j | |
| r, g, b = gx[idx][:3] | |
| vals = [r, g, b] | |
| variances.append(((i * 16, j * 16), statistics.pstdev(vals))) | |
| variances.sort(key=lambda kv: -kv[1]) | |
| facts.saliency_regions = [ | |
| {"x": x, "y": y, "score": round(v, 3)} | |
| for (x, y), v in variances[:6] | |
| ] | |
| # edge density via PIL edge detection | |
| try: | |
| import ImageFilter | |
| except ImportError: | |
| from PIL import ImageFilter | |
| edges = small.convert("L").filter(ImageFilter.FIND_EDGES) | |
| epx = list(edges.getdata()) | |
| facts.edge_density = sum(1 for v in epx if v > 100) / len(epx) | |
| return facts | |
| def analyze(self, source) -> ImageFacts: | |
| img = self.load_image(source) | |
| if img is None: | |
| raise ValueError("Could not load image") | |
| facts = self._pixel_facts(img) | |
| # optional real ViT embedding | |
| if self.vit is not None: | |
| try: | |
| from transformers import AutoImageProcessor | |
| if self.processor is None: | |
| self.processor = AutoImageProcessor.from_pretrained( | |
| "google/vit-base-patch16-224-in21k" | |
| ) | |
| with torch.no_grad(): | |
| inputs = self.processor(images=img, return_tensors="pt").to(self.device) | |
| out = self.vit(**inputs) | |
| facts.embedding = out.last_hidden_state # (1, patches+1, dim) | |
| except Exception: | |
| facts.embedding = None | |
| # optional object detection | |
| if self.detector is not None: | |
| try: | |
| with torch.no_grad(): | |
| det = self.det_processor(images=img, return_tensors="pt").to(self.device) | |
| outputs = self.detector(**det) | |
| target_sizes = torch.tensor([[img.height, img.width]]) | |
| results = self.det_processor.post_process_object_detection( | |
| outputs, threshold=0.5, target_sizes=target_sizes | |
| )[0] | |
| for score, label, box in zip( | |
| results["scores"].tolist(), | |
| results["labels"].tolist(), | |
| results["boxes"].tolist(), | |
| ): | |
| label_str = self.detector.config.id2label.get(label, "obj") | |
| facts.objects.append({ | |
| "label": label_str, | |
| "conf": score, | |
| "box": [round(b, 1) for b in box], | |
| }) | |
| except Exception: | |
| pass | |
| return facts |