Instructions to use dkudos/cinimod-devops with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use dkudos/cinimod-devops with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf dkudos/cinimod-devops:Q8_0 # Run inference directly in the terminal: llama cli -hf dkudos/cinimod-devops:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf dkudos/cinimod-devops:Q8_0 # Run inference directly in the terminal: llama cli -hf dkudos/cinimod-devops:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf dkudos/cinimod-devops:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf dkudos/cinimod-devops:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf dkudos/cinimod-devops:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf dkudos/cinimod-devops:Q8_0
Use Docker
docker model run hf.co/dkudos/cinimod-devops:Q8_0
- LM Studio
- Jan
- vLLM
How to use dkudos/cinimod-devops with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dkudos/cinimod-devops" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dkudos/cinimod-devops", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/dkudos/cinimod-devops:Q8_0
- Ollama
How to use dkudos/cinimod-devops with Ollama:
ollama run hf.co/dkudos/cinimod-devops:Q8_0
- Unsloth Desktop
- Docker Model Runner
How to use dkudos/cinimod-devops with Docker Model Runner:
docker model run hf.co/dkudos/cinimod-devops:Q8_0
- Lemonade
How to use dkudos/cinimod-devops with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull dkudos/cinimod-devops:Q8_0
Run and chat with the model
lemonade run user.cinimod-devops-Q8_0
List all available models
lemonade list
- Atomic Chat
File size: 5,820 Bytes
6d2daa2 a7a07a2 5a15826 a7a07a2 5a15826 6d2daa2 cb6abf3 5a15826 6d2daa2 5a15826 6d2daa2 5a15826 ad7c8c2 5a15826 8d7c1d2 1c616ab 8d7c1d2 1c616ab 8d7c1d2 1c616ab 5a15826 a7a07a2 5a15826 a7a07a2 5a15826 a7a07a2 1c616ab a7a07a2 5a15826 a7a07a2 5a15826 a7a07a2 5a15826 a7a07a2 5a15826 a7a07a2 5a15826 a7a07a2 5a15826 a7a07a2 5a15826 a7a07a2 5a15826 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 | ---
language: en
license: apache-2.0
base_model: dkudos/cinimod-devops
tags:
- cinimod
- devops
- llm
- causal-lm
- llama
- llama.cpp
pipeline_tag: text-generation
model-index:
- name: Cinimod DevOps 300M
results: []
---
# The Model is ONLY PRE-TRAINED ATM
# Cinimod DevOps 300M
A 287M-parameter decoder-only causal language model (Llama-3 style architecture), trained from scratch on a DevOps/ops SysAdmin domain corpus. Target usage: devops tooling assistance, ops documentation, and small on-box language modeling.
## Model Details
| Property | Value |
|---|---|
| Parameters | 287,310,848 (~287M) |
| Architecture | Llama-style decoder-only (custom, not stock `transformers` LlamaForCausalLM params) |
| Hidden size | 1024 |
| Layers | 20 |
| Attention heads | 16 |
| KV heads (GQA) | 4 |
| Intermediate size | 2730 |
| Vocab size | 65,536 (BPE) |
| Position embeddings | RoPE, theta = 500000 |
| Trained context | 4096 tokens |
| Max context (served) | up to 256K via linear RoPE scaling |
| Embeddings | tied (no separate lm_head) |
## Training
- **Objective**: from-scratch pretraining on a DevOps/ops corpus.
- **Compute**: 2x RTX 4090 (24 GB each, bf16), DeepSpeed ZeRO-2, FP32 master weights via bf16 autocast.
- **Tokens**: one epoch over ~132,068 sequences at seq_len 4096 (~540M tokens).
- **Steps**: 4000, warmup 40, LR 6e-4 cosine decay (final step LR ~0).
- **Efficient attention**: `torch.nn.functional.scaled_dot_product_attention` (flash path via flash-attn 2).
- **Loss trajectory**: train loss 0.43 (step 2000) -> 0.35 (step 4000).
### Evaluation
- **Full validation** (17,492 bins / 123,656 sequences @ 4096): mean eval loss **2.3163** (perplexity **10.14**). Final log in `full_val_eval.log`.
## Files
| File | Description | Size |
|---|---|---|
| `model.safetensors` | Full bf16 PyTorch weights (HF format with `config.json`, `tokenizer.json`/`tokenizer_config.json`) | 548 MiB |
| `config.json` | Model config (transformers) | - |
| `tokenizer.json` / `tokenizer_config.json` | BPE tokenizer (vocab 65,536) | - |
| `train_log.log` | Full training log (steps, losses, LR) | - |
| `full_val_eval.log` | Held-out full validation eval log | - |
GGUF files are listed in the [GGUF section](#gguf-llamacpp--recommended) above.
## GGUF (llama.cpp) — recommended
Ready-to-serve GGUF quantizations. Both are standalone single files with no dependencies (no Cinimod source code needed). The token embedding tensor is left in BF16/F16 (the Q8_0 quantizer keeps non-32-divisible dims at F16); all other weights are as noted.
| Quantization | File | Size | Notes |
|---|---|---|---|
| **Q8_0** | [`checkpoint-4000-Q8_0.gguf`](https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-Q8_0.gguf) | 344 MiB | Recommended default. ~8-bit, near-lossless, ~2x smaller than F16 |
| **F16** | [`checkpoint-4000-f16.gguf`](https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-f16.gguf) | 550 MiB | Best fidelity for llama.cpp |
## How to run
### HuggingFace transformers (PyTorch)
The `model.safetensors` require the Cinimod architecture classes (`cinimod.model.llama.LlamaForCausalLM`) — a custom Llama variant, **not** the stock `transformers.LlamaForCausalLM`. Load from the repo source only:
```python
import sys
sys.path.insert(0, "/path/to/cinimod-llm/src") # package src/cinimod
from cinimod.model.llama import LlamaForCausalLM
from transformers import PreTrainedTokenizerFast
model = LlamaForCausalLM.from_pretrained("dkudos/cinimod-devops")
tok = PreTrainedTokenizerFast.from_pretrained("dkudos/cinimod-devops")
ids = tok.encode("how do I check nginx status", return_tensors="pt")
out = model.generate(ids, max_new_tokens=64)
print(tok.decode(out[0]))
```
> If you are not in the Cinimod repo, use the GGUFs instead — they are standalone and need no source code. We publish GGUFs precisely because the HF-PyTorch path depends on the custom architecture classes.
### llama.cpp (recommended for serving)
Both GGUFs load directly in llama.cpp / llama-server with no external deps.
```bash
# Q8_0 (default)
wget https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-Q8_0.gguf
llama-server -m checkpoint-4000-Q8_0.gguf --port 8080
# or F16 for best fidelity
wget https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-f16.gguf
llama-server -m checkpoint-4000-f16.gguf --port 8080
```
256K context via linear RoPE scaling (trained at 4096):
```bash
llama-server -m dkudos/cinimod-devops/checkpoint-4000-Q8_0.gguf \
--ctx-size 262144 --rope-scaling linear --rope-scale 64 --port 8080
```
Rope scaling is **serve-time only**; this model ships with `rope_scaling: null`. For aggressive 64x scaling, Yarn (`--rope-scaling yarn --rope-scale 64`) often generalizes better than linear if long-range coherence suffers.
One-line test:
```bash
llama-server -m checkpoint-4000-Q8_0.gguf --ctx-size 262144 --rope-scaling linear --rope-scale 64
curl http://localhost:8080/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"messages":[{"role":"user","content":"List 5 common systemd service commands"}],"max_tokens":128}'
```
## Notes on the tokenizer
Vocabulary is a 65,536-token BPE (custom, `tokenizers` backend). `<pad>`, `<s>`, `</s>`, `<unk>` are at indices 0-3, trained with `pad_token_id=0`. It is a plain causal LM — no chat template is baked in. If GGUF chat-format warnings appear they are just llama.cpp server defaults, not part of the model.
## Limitations
- Pretrained from scratch on a single domain (DevOps) for one epoch at small scale (~287M) — expect domain-limited fluency, not general world knowledge.
- Exact transformers architecture classes are Cinimod-custom; use the GGUFs for maximum portability (no source code needed).
## License
Apache 2.0 |