File size: 5,820 Bytes
6d2daa2
a7a07a2
 
5a15826
a7a07a2
5a15826
 
 
 
 
 
 
 
 
 
6d2daa2
cb6abf3
5a15826
6d2daa2
5a15826
6d2daa2
 
 
5a15826
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ad7c8c2
5a15826
 
 
 
 
 
 
 
 
 
 
8d7c1d2
 
1c616ab
 
8d7c1d2
1c616ab
8d7c1d2
 
 
 
1c616ab
5a15826
 
 
 
 
a7a07a2
 
5a15826
 
 
 
 
 
 
 
 
 
a7a07a2
 
5a15826
 
 
 
 
 
a7a07a2
1c616ab
 
 
 
 
 
 
a7a07a2
 
5a15826
a7a07a2
5a15826
 
 
 
a7a07a2
5a15826
a7a07a2
5a15826
a7a07a2
5a15826
 
 
 
 
a7a07a2
5a15826
a7a07a2
5a15826
a7a07a2
5a15826
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
---
language: en
license: apache-2.0
base_model: dkudos/cinimod-devops
tags:
  - cinimod
  - devops
  - llm
  - causal-lm
  - llama
  - llama.cpp
pipeline_tag: text-generation
model-index:
  - name: Cinimod DevOps 300M
    results: []
---
# The Model is ONLY PRE-TRAINED ATM
# Cinimod DevOps 300M

A 287M-parameter decoder-only causal language model (Llama-3 style architecture), trained from scratch on a DevOps/ops SysAdmin domain corpus. Target usage: devops tooling assistance, ops documentation, and small on-box language modeling.

## Model Details

| Property | Value |
|---|---|
| Parameters | 287,310,848 (~287M) |
| Architecture | Llama-style decoder-only (custom, not stock `transformers` LlamaForCausalLM params) |
| Hidden size | 1024 |
| Layers | 20 |
| Attention heads | 16 |
| KV heads (GQA) | 4 |
| Intermediate size | 2730 |
| Vocab size | 65,536 (BPE) |
| Position embeddings | RoPE, theta = 500000 |
| Trained context | 4096 tokens |
| Max context (served) | up to 256K via linear RoPE scaling |
| Embeddings | tied (no separate lm_head) |

## Training

- **Objective**: from-scratch pretraining on a DevOps/ops corpus.
- **Compute**: 2x RTX 4090 (24 GB each, bf16), DeepSpeed ZeRO-2, FP32 master weights via bf16 autocast.
- **Tokens**: one epoch over ~132,068 sequences at seq_len 4096 (~540M tokens).
- **Steps**: 4000, warmup 40, LR 6e-4 cosine decay (final step LR ~0).
- **Efficient attention**: `torch.nn.functional.scaled_dot_product_attention` (flash path via flash-attn 2).
- **Loss trajectory**: train loss 0.43 (step 2000) -> 0.35 (step 4000).

### Evaluation

- **Full validation** (17,492 bins / 123,656 sequences @ 4096): mean eval loss **2.3163** (perplexity **10.14**). Final log in `full_val_eval.log`.

## Files

| File | Description | Size |
|---|---|---|
| `model.safetensors` | Full bf16 PyTorch weights (HF format with `config.json`, `tokenizer.json`/`tokenizer_config.json`) | 548 MiB |
| `config.json` | Model config (transformers) | - |
| `tokenizer.json` / `tokenizer_config.json` | BPE tokenizer (vocab 65,536) | - |
| `train_log.log` | Full training log (steps, losses, LR) | - |
| `full_val_eval.log` | Held-out full validation eval log | - |

GGUF files are listed in the [GGUF section](#gguf-llamacpp--recommended) above.

## GGUF (llama.cpp) — recommended

Ready-to-serve GGUF quantizations. Both are standalone single files with no dependencies (no Cinimod source code needed). The token embedding tensor is left in BF16/F16 (the Q8_0 quantizer keeps non-32-divisible dims at F16); all other weights are as noted.

| Quantization | File | Size | Notes |
|---|---|---|---|
| **Q8_0** | [`checkpoint-4000-Q8_0.gguf`](https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-Q8_0.gguf) | 344 MiB | Recommended default. ~8-bit, near-lossless, ~2x smaller than F16 |
| **F16** | [`checkpoint-4000-f16.gguf`](https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-f16.gguf) | 550 MiB | Best fidelity for llama.cpp |

## How to run

### HuggingFace transformers (PyTorch)

The `model.safetensors` require the Cinimod architecture classes (`cinimod.model.llama.LlamaForCausalLM`) — a custom Llama variant, **not** the stock `transformers.LlamaForCausalLM`. Load from the repo source only:

```python
import sys
sys.path.insert(0, "/path/to/cinimod-llm/src")   # package src/cinimod
from cinimod.model.llama import LlamaForCausalLM
from transformers import PreTrainedTokenizerFast

model = LlamaForCausalLM.from_pretrained("dkudos/cinimod-devops")
tok = PreTrainedTokenizerFast.from_pretrained("dkudos/cinimod-devops")
ids = tok.encode("how do I check nginx status", return_tensors="pt")
out = model.generate(ids, max_new_tokens=64)
print(tok.decode(out[0]))
```

> If you are not in the Cinimod repo, use the GGUFs instead — they are standalone and need no source code. We publish GGUFs precisely because the HF-PyTorch path depends on the custom architecture classes.

### llama.cpp (recommended for serving)

Both GGUFs load directly in llama.cpp / llama-server with no external deps.

```bash
# Q8_0 (default)
wget https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-Q8_0.gguf
llama-server -m checkpoint-4000-Q8_0.gguf --port 8080

# or F16 for best fidelity
wget https://huggingface.co/dkudos/cinimod-devops/resolve/main/checkpoint-4000-f16.gguf
llama-server -m checkpoint-4000-f16.gguf --port 8080
```

256K context via linear RoPE scaling (trained at 4096):

```bash
llama-server -m dkudos/cinimod-devops/checkpoint-4000-Q8_0.gguf \
  --ctx-size 262144 --rope-scaling linear --rope-scale 64 --port 8080
```

Rope scaling is **serve-time only**; this model ships with `rope_scaling: null`. For aggressive 64x scaling, Yarn (`--rope-scaling yarn --rope-scale 64`) often generalizes better than linear if long-range coherence suffers.

One-line test:

```bash
llama-server -m checkpoint-4000-Q8_0.gguf --ctx-size 262144 --rope-scaling linear --rope-scale 64
curl http://localhost:8080/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"messages":[{"role":"user","content":"List 5 common systemd service commands"}],"max_tokens":128}'
```

## Notes on the tokenizer

Vocabulary is a 65,536-token BPE (custom, `tokenizers` backend). `<pad>`, `<s>`, `</s>`, `<unk>` are at indices 0-3, trained with `pad_token_id=0`. It is a plain causal LM — no chat template is baked in. If GGUF chat-format warnings appear they are just llama.cpp server defaults, not part of the model.

## Limitations

- Pretrained from scratch on a single domain (DevOps) for one epoch at small scale (~287M) — expect domain-limited fluency, not general world knowledge.
- Exact transformers architecture classes are Cinimod-custom; use the GGUFs for maximum portability (no source code needed).

## License

Apache 2.0