Instructions to use dkudos/cinimod-devops with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use dkudos/cinimod-devops with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf dkudos/cinimod-devops:Q8_0 # Run inference directly in the terminal: llama cli -hf dkudos/cinimod-devops:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf dkudos/cinimod-devops:Q8_0 # Run inference directly in the terminal: llama cli -hf dkudos/cinimod-devops:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf dkudos/cinimod-devops:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf dkudos/cinimod-devops:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf dkudos/cinimod-devops:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf dkudos/cinimod-devops:Q8_0
Use Docker
docker model run hf.co/dkudos/cinimod-devops:Q8_0
- LM Studio
- Jan
- vLLM
How to use dkudos/cinimod-devops with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dkudos/cinimod-devops" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dkudos/cinimod-devops", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/dkudos/cinimod-devops:Q8_0
- Ollama
How to use dkudos/cinimod-devops with Ollama:
ollama run hf.co/dkudos/cinimod-devops:Q8_0
- Unsloth Desktop
- Docker Model Runner
How to use dkudos/cinimod-devops with Docker Model Runner:
docker model run hf.co/dkudos/cinimod-devops:Q8_0
- Lemonade
How to use dkudos/cinimod-devops with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull dkudos/cinimod-devops:Q8_0
Run and chat with the model
lemonade run user.cinimod-devops-Q8_0
List all available models
lemonade list
- Atomic Chat
|
Download README.md from dkudos/cinimod-devops: direct link, hf CLI and curl.
- Browser
- Download file 5.82 kB
-
https://huggingface.co/dkudos/cinimod-devops/resolve/main/README.md
- Command line
-
hf download hf://dkudos/cinimod-devops/README.md
-
curl -L -o README.md https://huggingface.co/dkudos/cinimod-devops/resolve/main/README.md
5.82 kB
| language: en | |
| license: apache-2.0 | |
| base_model: dkudos/cinimod-devops | |
| tags: | |
| - cinimod | |
| - devops | |
| - llm | |
| - causal-lm | |
| - llama | |
| - llama.cpp | |
| pipeline_tag: text-generation | |
| model-index: | |
| - name: Cinimod DevOps 300M | |
| results: [] | |
| # The Model is ONLY PRE-TRAINED ATM | |
| # Cinimod DevOps 300M | |
| A 287M-parameter decoder-only causal language model (Llama-3 style architecture), trained from scratch on a DevOps/ops SysAdmin domain corpus. Target usage: devops tooling assistance, ops documentation, and small on-box language modeling. | |
| ## Model Details | |
| | Property | Value | | |
| |---|---| | |
| | Parameters | 287,310,848 (~287M) | | |
| | Architecture | Llama-style decoder-only (custom, not stock `transformers` LlamaForCausalLM params) | | |
| | Hidden size | 1024 | | |
| | Layers | 20 | | |
| | Attention heads | 16 | | |
| | KV heads (GQA) | 4 | | |
| | Intermediate size | 2730 | | |
| | Vocab size | 65,536 (BPE) | | |
| | Position embeddings | RoPE, theta = 500000 | | |
| | Trained context | 4096 tokens | | |
| | Max context (served) | up to 256K via linear RoPE scaling | | |
| | Embeddings | tied (no separate lm_head) | | |
| ## Training | |
| - **Objective**: from-scratch pretraining on a DevOps/ops corpus. | |
| - **Compute**: 2x RTX 4090 (24 GB each, bf16), DeepSpeed ZeRO-2, FP32 master weights via bf16 autocast. | |
| - **Tokens**: one epoch over ~132,068 sequences at seq_len 4096 (~540M tokens). | |
| - **Steps**: 4000, warmup 40, LR 6e-4 cosine decay (final step LR ~0). | |
| - **Efficient attention**: `torch.nn.functional.scaled_dot_product_attention` (flash path via flash-attn 2). | |
| - **Loss trajectory**: train loss 0.43 (step 2000) -> 0.35 (step 4000). | |
| ### Evaluation | |
| - **Full validation** (17,492 bins / 123,656 sequences @ 4096): mean eval loss **2.3163** (perplexity **10.14**). Final log in `full_val_eval.log`. | |
| ## Files | |
| | File | Description | Size | | |
| |---|---|---| | |
| | `model.safetensors` | Full bf16 PyTorch weights (HF format with `config.json`, `tokenizer.json`/`tokenizer_config.json`) | 548 MiB | | |
| | `config.json` | Model config (transformers) | - | | |
| | `tokenizer.json` / `tokenizer_config.json` | BPE tokenizer (vocab 65,536) | - | | |
| | `train_log.log` | Full training log (steps, losses, LR) | - | | |
| | `full_val_eval.log` | Held-out full validation eval log | - | | |
| GGUF files are listed in the [GGUF section](#gguf-llamacpp--recommended) above. | |
| ## GGUF (llama.cpp) — recommended | |
| Ready-to-serve GGUF quantizations. Both are standalone single files with no dependencies (no Cinimod source code needed). The token embedding tensor is left in BF16/F16 (the Q8_0 quantizer keeps non-32-divisible dims at F16); all other weights are as noted. | |
| | Quantization | File | Size | Notes | | |
| |---|---|---|---| | |
| | **Q8_0** | [`checkpoint-4000-Q8_0.gguf`](https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-Q8_0.gguf) | 344 MiB | Recommended default. ~8-bit, near-lossless, ~2x smaller than F16 | | |
| | **F16** | [`checkpoint-4000-f16.gguf`](https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-f16.gguf) | 550 MiB | Best fidelity for llama.cpp | | |
| ## How to run | |
| ### HuggingFace transformers (PyTorch) | |
| The `model.safetensors` require the Cinimod architecture classes (`cinimod.model.llama.LlamaForCausalLM`) — a custom Llama variant, **not** the stock `transformers.LlamaForCausalLM`. Load from the repo source only: | |
| ```python | |
| import sys | |
| sys.path.insert(0, "/path/to/cinimod-llm/src") # package src/cinimod | |
| from cinimod.model.llama import LlamaForCausalLM | |
| from transformers import PreTrainedTokenizerFast | |
| model = LlamaForCausalLM.from_pretrained("dkudos/cinimod-devops") | |
| tok = PreTrainedTokenizerFast.from_pretrained("dkudos/cinimod-devops") | |
| ids = tok.encode("how do I check nginx status", return_tensors="pt") | |
| out = model.generate(ids, max_new_tokens=64) | |
| print(tok.decode(out[0])) | |
| ``` | |
| > If you are not in the Cinimod repo, use the GGUFs instead — they are standalone and need no source code. We publish GGUFs precisely because the HF-PyTorch path depends on the custom architecture classes. | |
| ### llama.cpp (recommended for serving) | |
| Both GGUFs load directly in llama.cpp / llama-server with no external deps. | |
| ```bash | |
| # Q8_0 (default) | |
| wget https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-Q8_0.gguf | |
| llama-server -m checkpoint-4000-Q8_0.gguf --port 8080 | |
| # or F16 for best fidelity | |
| wget https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-f16.gguf | |
| llama-server -m checkpoint-4000-f16.gguf --port 8080 | |
| ``` | |
| 256K context via linear RoPE scaling (trained at 4096): | |
| ```bash | |
| llama-server -m dkudos/cinimod-devops/checkpoint-4000-Q8_0.gguf \ | |
| --ctx-size 262144 --rope-scaling linear --rope-scale 64 --port 8080 | |
| ``` | |
| Rope scaling is **serve-time only**; this model ships with `rope_scaling: null`. For aggressive 64x scaling, Yarn (`--rope-scaling yarn --rope-scale 64`) often generalizes better than linear if long-range coherence suffers. | |
| One-line test: | |
| ```bash | |
| llama-server -m checkpoint-4000-Q8_0.gguf --ctx-size 262144 --rope-scaling linear --rope-scale 64 | |
| curl http://localhost:8080/v1/chat/completions -H 'Content-Type: application/json' \ | |
| -d '{"messages":[{"role":"user","content":"List 5 common systemd service commands"}],"max_tokens":128}' | |
| ``` | |
| ## Notes on the tokenizer | |
| Vocabulary is a 65,536-token BPE (custom, `tokenizers` backend). `<pad>`, `<s>`, `</s>`, `<unk>` are at indices 0-3, trained with `pad_token_id=0`. It is a plain causal LM — no chat template is baked in. If GGUF chat-format warnings appear they are just llama.cpp server defaults, not part of the model. | |
| ## Limitations | |
| - Pretrained from scratch on a single domain (DevOps) for one epoch at small scale (~287M) — expect domain-limited fluency, not general world knowledge. | |
| - Exact transformers architecture classes are Cinimod-custom; use the GGUFs for maximum portability (no source code needed). | |
| ## License | |
| Apache 2.0 |