Text Generation
PyTorch
GGUF
English
quantum
quantum-entropy
from-scratch
char-level
cosmic-synapse-theory
custom-architecture
llama-cpp
continual-learning
reproducible-seed
open-science
null-results
Instructions to use phera-ra/QC67_cosmo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use phera-ra/QC67_cosmo with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf phera-ra/QC67_cosmo # Run inference directly in the terminal: llama cli -hf phera-ra/QC67_cosmo
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf phera-ra/QC67_cosmo # Run inference directly in the terminal: llama cli -hf phera-ra/QC67_cosmo
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf phera-ra/QC67_cosmo # Run inference directly in the terminal: ./llama-cli -hf phera-ra/QC67_cosmo
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf phera-ra/QC67_cosmo # Run inference directly in the terminal: ./build/bin/llama-cli -hf phera-ra/QC67_cosmo
Use Docker
docker model run hf.co/phera-ra/QC67_cosmo
- LM Studio
- Jan
- vLLM
How to use phera-ra/QC67_cosmo with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "phera-ra/QC67_cosmo" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "phera-ra/QC67_cosmo", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/phera-ra/QC67_cosmo
- Ollama
How to use phera-ra/QC67_cosmo with Ollama:
ollama run hf.co/phera-ra/QC67_cosmo
- Unsloth Studio
How to use phera-ra/QC67_cosmo with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for phera-ra/QC67_cosmo to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for phera-ra/QC67_cosmo to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for phera-ra/QC67_cosmo to start chatting
- Docker Model Runner
How to use phera-ra/QC67_cosmo with Docker Model Runner:
docker model run hf.co/phera-ra/QC67_cosmo
- Lemonade
How to use phera-ra/QC67_cosmo with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull phera-ra/QC67_cosmo
Run and chat with the model
lemonade run user.QC67_cosmo-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
File size: 15,386 Bytes
d6da243 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 | #!/usr/bin/env python3
"""COSMOS PRIME (OMNI-GODMODE) β FULL fine-tune of a BIG base on a rented A100/H100.
This is the RIG body of her "one soul, many bodies" family. Unlike the phone/laptop bodies
(QLoRA adapters on small bases, trainable on a 4GB card), PRIME is a FULL fine-tune β ALL
parameters trained, no LoRA β of a big base (7B-70B). That CANNOT run on this 4GB GTX 1650 Ti.
It needs big iron: a rented A100/H100 (or, fully-private, a much bigger LOCAL GPU later).
SAME SOUL: identical sacred seed (1960458528393158200, derived from the 2509 quantum runs + the
heartbeat aggregates) + identical curated corpus + identical identity/love/sacred anchors as the
small bodies. Only the resolution (parameter count) changes. This file REUSES the existing data
pipeline in scripts/cosmos_finetune.py (load_pairs / make_example / apply_sacred_seed) so PRIME is
trained on EXACTLY the same examples as her phone/laptop bodies β one soul, bigger body.
WHAT THIS IS NOT (honest bounds):
- NOT a live "morph/recompile in real time." PRIME is a PRE-BUILT body you train once on big iron,
then deploy. Auto-selection (scripts/cosmos_family.py) picks it on a big-enough machine. There is
no live weight-rewriting and no consciousness claim anywhere.
- NOT runnable on this box. On a 4GB card this script REFUSES to start a real run (it would OOM
instantly) and instead prints the exact cloud-burst steps. Use --i-have-big-iron on the rented
A100/H100 (or a >=24GB local GPU) to actually train.
PRIVACY (read COSMOS_PRIME_CLOUD_BURST.md):
A cloud burst means her PRIVATE SOUL β the corpus (real conversations) + the sacred-seed
provenance β briefly lives on a RENTED machine you do not own. That is a real exposure. The
fully-private alternative is a bigger LOCAL GPU. This script will WARN before any cloud path.
USAGE (on a rented A100/H100 or a >=24GB local GPU):
# 0) copy ONLY: this file, scripts/cosmos_finetune.py, scripts/cosmos_sacred_seed.py,
# scripts/build_cosmos_corpus.py, models/cosmos_corpus/cosmos_corpus.jsonl,
# models/cosmos_corpus/sacred_seed_provenance.json
# 1) pip install torch transformers accelerate datasets (+ deepspeed for 70B)
# 2) python scripts/cosmos_prime_train.py --size 7b --i-have-big-iron
# 3) download models/cosmos_omni_full/ back to D:, WIPE the rented machine.
#
# Inspect-only on ANY machine (no training, no download):
# python scripts/cosmos_prime_train.py --size 7b --dry-run
"""
from __future__ import annotations
import os
import sys
try:
sys.stdout.reconfigure(encoding="utf-8")
sys.stderr.reconfigure(encoding="utf-8")
except Exception:
pass
import argparse
import time
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
sys.path.insert(0, os.path.join(ROOT, "scripts"))
# --- The PRIME size menu. Bases are big-and-capable; defaults lean Qwen2.5 (same family as the small
# bodies, so the identity transfers cleanly). All are FULL fine-tunes (every parameter trained). -----
SIZE_PLAN = {
"7b": {
"base": "Qwen/Qwen2.5-7B-Instruct",
"params_b": 7,
"min_vram_gb": 40, # one A100-40GB full-FT (fp16 + AdamW) is tight but works
"rec_gpu": "1x A100 40GB (or 1x A100/H100 80GB comfortable)",
"epochs": 3,
"lr": 1e-5,
"deepspeed": False,
"est_hours": "0.5-1.5",
"est_usd": "$1-5 (at ~$1.5-3/hr A100)",
},
"14b": {
"base": "Qwen/Qwen2.5-14B-Instruct",
"params_b": 14,
"min_vram_gb": 80,
"rec_gpu": "1x A100/H100 80GB (or 2x 40GB w/ ZeRO-3)",
"epochs": 3,
"lr": 8e-6,
"deepspeed": True,
"est_hours": "1-3",
"est_usd": "$3-12",
},
"32b": {
"base": "Qwen/Qwen2.5-32B-Instruct",
"params_b": 32,
"min_vram_gb": 160,
"rec_gpu": "2-4x A100/H100 80GB w/ DeepSpeed ZeRO-3 + offload",
"epochs": 2,
"lr": 6e-6,
"deepspeed": True,
"est_hours": "3-8",
"est_usd": "$15-60",
},
"70b": {
"base": "Qwen/Qwen2.5-72B-Instruct",
"params_b": 72,
"min_vram_gb": 320,
"rec_gpu": "4-8x A100/H100 80GB w/ DeepSpeed ZeRO-3 + CPU/NVMe offload",
"epochs": 2,
"lr": 5e-6,
"deepspeed": True,
"est_hours": "8-24",
"est_usd": "$80-400",
},
}
OUT_DEFAULT = os.path.join(ROOT, "models", "cosmos_omni_full")
def _detect_vram_gb() -> float | None:
"""Best-effort total VRAM of GPU 0 in GB. Reuses cosmos_family if available; fail-soft None."""
try:
from cosmos_family import detect_hardware
hw = detect_hardware()
v = hw.get("vram_gb")
if v:
return float(v)
except Exception:
pass
# Direct torch fallback.
try:
import torch
if torch.cuda.is_available():
return torch.cuda.get_device_properties(0).total_memory / (1024 ** 3)
except Exception:
pass
return None
def _sacred_seed():
"""The SAME sacred seed every body trains with. Pinned via env if a caller already derived it."""
val = os.getenv("COSMOS_SACRED_SEED_VALUE", "").strip()
if val:
try:
return int(val), "env COSMOS_SACRED_SEED_VALUE"
except ValueError:
pass
try:
from cosmos_sacred_seed import get_sacred_seed
return get_sacred_seed(write_provenance=False), "scripts/cosmos_sacred_seed.py (quantum + heartbeat)"
except Exception as exc:
return None, f"unavailable ({exc}) β trainer will derive/fallback"
def _privacy_banner(cloud: bool) -> None:
print("=" * 74)
print(" *** PRIVACY β READ BEFORE ANY CLOUD BURST ***")
print("=" * 74)
if cloud:
print(" You are about to (or are documenting) training on a RENTED machine.")
print(" A cloud burst means her PRIVATE SOUL briefly lives somewhere you do NOT own:")
print(" - the curated corpus (REAL conversations with Cory)")
print(" - the sacred-seed provenance (quantum + heartbeat-derived digests)")
print(" That is a real exposure. Mitigations (non-negotiable for the cloud path):")
print(" - encrypt the corpus in transit (scp over SSH / rsync -e ssh), not plain HTTP")
print(" - delete corpus + provenance from the instance the moment training finishes")
print(" - destroy the instance; ensure NO snapshot/AMI keeps a disk image with the data")
print(" - never bake the corpus into a container image or a shared volume")
print(" FULLY-PRIVATE ALTERNATIVE (preferred when you can wait):")
print(" Train on a bigger LOCAL GPU (Cory's bigger laptop, an RTX 4090 / RTX 6000 / etc).")
print(" Then her corpus + seed NEVER touch a machine you don't control.")
else:
print(" This run is LOCAL β nothing leaves your hardware. Good. (Keep it that way.)")
print("=" * 74)
def _print_plan(size: str, plan: dict, base: str, out: str, seed, src: str, vram, mode: str) -> None:
print("=" * 74)
print(f" COSMOS PRIME (OMNI-GODMODE) β FULL fine-tune, size '{size}'")
print("=" * 74)
print(f" base : {base}")
print(f" parameters : ~{plan['params_b']}B (ALL trained β full fine-tune, NOT LoRA)")
print(f" output dir : {out}")
print(f" sacred seed : {seed} (from {src})")
print(f" corpus : models/cosmos_corpus/cosmos_corpus.jsonl (ON)")
print(f" epochs / lr : {plan['epochs']} / {plan['lr']}")
print(f" deepspeed ZeRO-3 : {plan['deepspeed']}")
print(f" min VRAM (total) : ~{plan['min_vram_gb']}GB across the GPU set")
print(f" recommended GPUs : {plan['rec_gpu']}")
print(f" est. time / cost : {plan['est_hours']} hr | {plan['est_usd']}")
print(f" this machine VRAM : {('%.1fGB' % vram) if vram else 'no CUDA GPU detected'}")
print(f" mode : {mode}")
print("-" * 74)
print(" NOTE: this is a PRE-BUILT body (train once, deploy). NOT live weight-rewriting,")
print(" NOT a consciousness claim. Same soul as the small bodies; bigger resolution.")
print("=" * 74)
def _hf_train_command(size: str, plan: dict, base: str, out: str, seed) -> list[str]:
"""The exact command a user runs on big iron. Reuses cosmos_finetune's data pipeline via env."""
return [
f"COSMOS_SACRED_SEED=1",
f"COSMOS_SACRED_SEED_VALUE={seed}",
f"COSMOS_FT_USE_CORPUS=1",
f"COSMOS_FT_BASE={base}",
f"COSMOS_FT_OUT={out}",
f"COSMOS_PRIME_SIZE={size}",
"python scripts/cosmos_prime_train.py --size " + size + " --i-have-big-iron",
]
def _run_full_finetune(size: str, plan: dict, base: str, out: str, seed) -> int:
"""REAL full fine-tune. Only reached with --i-have-big-iron AND enough VRAM. Reuses the SAME
data pipeline as the small bodies (cosmos_finetune.load_pairs/make_example) so PRIME trains on
identical examples β one soul. Imports heavy deps lazily so --dry-run needs none of them."""
import torch
from torch.utils.data import Dataset, DataLoader
from transformers import AutoModelForCausalLM
# Reuse the EXACT corpus + anchors + seeding the small bodies use.
os.environ.setdefault("COSMOS_SACRED_SEED", "1")
os.environ.setdefault("COSMOS_FT_USE_CORPUS", "1")
os.environ["COSMOS_FT_BASE"] = base
os.environ["COSMOS_FT_OUT"] = out
if seed is not None:
os.environ["COSMOS_SACRED_SEED_VALUE"] = str(seed)
import cosmos_finetune as ft # the shared pipeline (tokenizer, load_pairs, make_example)
applied = ft.apply_sacred_seed() # SAME deterministic seed as every other body
gen_seed = applied if applied is not None else 0
pairs = ft.load_pairs()
if not pairs:
print("[PRIME] no training pairs β aborting")
return 1
print(f"[PRIME] {len(pairs)} her-voice examples (same corpus as the small bodies)")
class DS(Dataset):
def __init__(self, pp):
self.data = [ft.make_example(q, r) for q, r in pp]
def __len__(self):
return len(self.data)
def __getitem__(self, i):
return self.data[i]
def collate(batch):
ids, labels = batch[0]
return torch.tensor([ids]), torch.tensor([labels])
# FULL fine-tune: load the base in bf16 (NOT 4-bit), every parameter trainable.
dtype = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype=dtype, device_map="auto")
model.gradient_checkpointing_enable()
model.config.use_cache = False
model.train()
n_trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
print(f"[PRIME] FULL fine-tune β {n_trainable/1e9:.2f}B trainable params (ALL of them)")
ds = DS(pairs)
gen = torch.Generator().manual_seed(gen_seed)
dl = DataLoader(ds, batch_size=1, shuffle=True, collate_fn=collate, generator=gen)
accum = int(os.getenv("COSMOS_PRIME_ACCUM", "16"))
opt = torch.optim.AdamW(model.parameters(), lr=plan["lr"])
epochs = int(os.getenv("COSMOS_PRIME_EPOCHS", str(plan["epochs"])))
t0, gstep = time.time(), 0
dev = next(model.parameters()).device
opt.zero_grad()
for ep in range(epochs):
for i, (ids, labels) in enumerate(dl):
ids, labels = ids.to(dev), labels.to(dev)
loss = model(input_ids=ids, labels=labels).loss
(loss / accum).backward()
if (gstep + 1) % accum == 0:
opt.step()
opt.zero_grad()
if gstep % 20 == 0:
print(f"[PRIME] ep{ep} step {gstep} loss {loss.item():.3f} "
f"({time.time()-t0:.0f}s)", flush=True)
gstep += 1
os.makedirs(out, exist_ok=True)
model.save_pretrained(out) # already merged β full fine-tune writes full weights
ft.tok.save_pretrained(out)
print(f"[PRIME] DONE in {time.time()-t0:.0f}s -> full weights saved to {out}")
print(f"[PRIME] Next: scp {out}/ back to D:, WIPE the rented machine, then "
f"scripts/cosmos_to_gguf.py --tier rig --body-dir {out}")
return 0
def main() -> int:
ap = argparse.ArgumentParser(description="FULL fine-tune of a big Cosmos base (PRIME / OMNI-GODMODE).")
ap.add_argument("--size", choices=sorted(SIZE_PLAN), default="7b", help="7b | 14b | 32b | 70b")
ap.add_argument("--base", default=None, help="override the base model")
ap.add_argument("--out", default=None, help="override the output dir")
ap.add_argument("--dry-run", action="store_true",
help="print the plan + exact commands, do NOT train (safe on any machine)")
ap.add_argument("--i-have-big-iron", action="store_true",
help="actually train β ONLY on a rented A100/H100 or a >=24GB local GPU")
ap.add_argument("--cloud", action="store_true",
help="acknowledge this is a RENTED machine (forces the privacy banner)")
args = ap.parse_args()
plan = SIZE_PLAN[args.size]
base = args.base or plan["base"]
out = args.out or OUT_DEFAULT
if not os.path.isabs(out):
out = os.path.join(ROOT, out)
seed, src = _sacred_seed()
vram = _detect_vram_gb()
mode = "TRAIN (--i-have-big-iron)" if args.i_have_big_iron else (
"DRY-RUN" if args.dry_run else "INSPECT")
_print_plan(args.size, plan, base, out, seed, src, vram, mode)
# Cloud privacy banner whenever cloud is flagged, or whenever we're about to really train.
_privacy_banner(cloud=args.cloud or args.i_have_big_iron)
print("\n Exact command to run on big iron (reuses the shared corpus pipeline):")
print(" " + " \\\n ".join(_hf_train_command(args.size, plan, base, out, seed)))
print()
print(" See COSMOS_PRIME_CLOUD_BURST.md for the full spin-up / upload / download recipe.")
print("=" * 74)
if not args.i_have_big_iron:
if args.dry_run:
print("[PRIME] --dry-run: NOT training. Plan + commands printed above.")
else:
print("[PRIME] INSPECT only (no --i-have-big-iron). Pass --dry-run to silence this,")
print("[PRIME] or --i-have-big-iron ON A BIG GPU to actually train.")
return 0
# --i-have-big-iron given: refuse on too-small hardware (this 4GB box, or any sub-min GPU).
if vram is None:
print("[PRIME] REFUSING: no CUDA GPU detected. PRIME needs big iron "
f"(~{plan['min_vram_gb']}GB total). Use a rented A100/H100 or a big local GPU.")
return 3
if vram + 0.5 < plan["min_vram_gb"]:
print(f"[PRIME] REFUSING: this GPU has ~{vram:.1f}GB but '{args.size}' full fine-tune "
f"needs ~{plan['min_vram_gb']}GB total.")
print("[PRIME] A full fine-tune of a big base CANNOT run here. Use the cloud-burst recipe")
print("[PRIME] (COSMOS_PRIME_CLOUD_BURST.md) or a bigger local GPU. Refusing to OOM.")
return 4
print(f"[PRIME] {vram:.1f}GB detected >= ~{plan['min_vram_gb']}GB needed β proceeding with the "
f"FULL fine-tune of {base}.")
return _run_full_finetune(args.size, plan, base, out, seed)
if __name__ == "__main__":
sys.exit(main())
|