File size: 5,552 Bytes
ce2fb03
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
---
license: apache-2.0
base_model: shuhulx/Qwopus3.5-4B-Coder-Fable5-v1
datasets:
- Glint-Research/Fable-5-traces
language:
- en
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- llama-cpp
- lm-studio
- qwen3_5
- fable5
- reasoning
- agent
- tool-use
- function-calling
- coder
- coding
- debugging
- local-inference
- quantized
- conversational
---

<div align="center">

# 馃捇 Qwopus3.5-4B-Coder-Fable5-v1 GGUF

### GGUF builds for llama.cpp, LM Studio, and local inference

<p><b>Fable-5 traces</b><b>agentic coding</b><b>tool use</b><b>debugging</b></p>

</div>

---

## Overview

**Qwopus3.5-4B-Coder-Fable5-v1** is a Fable-5 trace continuation of [`Jackrong/Qwopus3.5-4B-Coder`](https://huggingface.co/Jackrong/Qwopus3.5-4B-Coder).

The base model, Qwopus3.5-4B-Coder, is a compact Qwen3.5-based coding model trained for reasoning, tool use, function calling, coding workflows, and agent-style behavior.

This release continues that model on [`Glint-Research/Fable-5-traces`](https://huggingface.co/datasets/Glint-Research/Fable-5-traces), a dataset of Claude Fable 5 local coding-agent traces. The dataset is heavily oriented around tool-use trajectories, repository work, local command context, code editing, debugging loops, and `<think>`-style reasoning completions.

The result is a small local coding-agent model intended for:

| Area | Description |
|---|---|
| Tool-use workflows | Bash, Read, Write, Edit, repo inspection, and action traces. |
| Debugging | Failing tests, stack traces, root-cause analysis, and patch planning. |
| Trace-style reasoning | Long-form planning and `<think>` style reasoning traces. |
| Local agents | Hermes-style, Claude-Code-style, OpenCode-style, and LM Studio workflows. |

## Files

Typical GGUF files:

- `Qwopus3.5-4B-Coder-Fable5-v1-Q4_K_M.gguf`
- `Qwopus3.5-4B-Coder-Fable5-v1-Q5_K_M.gguf`
- `Qwopus3.5-4B-Coder-Fable5-v1-mmproj-BF16.gguf`

## Which file should I use?

| File | Use case |
|---|---|
| `Q4_K_M` | Best default. Small, fast, good quality. |
| `Q5_K_M` | Better quality while still compact. |
| `Q8_0` | Higher quality, larger memory use, if included. |
| `mmproj-BF16` | Multimodal projector for compatible runtimes. |

## llama.cpp

```bash
llama-cli \
  -m Qwopus3.5-4B-Coder-Fable5-v1-Q5_K_M.gguf \
  -p "Write a Bash/Read/Edit style plan for debugging a failing Python repo." \
  -n 768 \
  --temp 0.7 \
  --top-p 0.95
```

## llama.cpp Server

```bash
llama-server \
  -m Qwopus3.5-4B-Coder-Fable5-v1-Q5_K_M.gguf \
  --host 0.0.0.0 \
  --port 8080 \
  --ctx-size 8192
```

Then call it with an OpenAI-compatible client:

```bash
curl -X POST "http://localhost:8080/v1/chat/completions" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "Qwopus3.5-4B-Coder-Fable5-v1-Q5_K_M.gguf",
    "messages": [
      {"role": "user", "content": "Write a tool-use plan for debugging a Python repo."}
    ],
    "temperature": 0.7,
    "top_p": 0.95
  }'
```


## About the Fable-5 Traces

[`Glint-Research/Fable-5-traces`](https://huggingface.co/datasets/Glint-Research/Fable-5-traces) contains Claude Fable 5 coding traces.

The dataset includes fields such as:

```text
uid
source_file
session
model
context
cot
output_type
output
completion
origin
```

The examples are not simple chat pairs. They are multi-step agent trajectories with local development context, reasoning traces, and tool-use outputs.

Common patterns in the dataset include:

- user coding requests
- local-command caveats
- repository inspection
- Bash command usage
- file reads
- file writes
- edits
- debugging passes
- playtesting / validation loops
- `<think>...</think>` reasoning traces
- tool-use completions

A large portion of the dataset is `tool_use` style data, which makes it especially relevant for local coding agents and developer automation.

## Capabilities

### Agentic coding

Designed for coding-agent loops where the model must inspect a repo, plan work, call tools, edit files, and validate changes.

### Tool-use style outputs

Works well with prompts that expose structured tools such as:

```text
Bash
Read
Write
Edit
Search
Grep
```

### Debugging and repair

Useful for:

- finding likely failing files
- explaining stack traces
- planning test commands
- proposing minimal patches
- iterating after errors

### Local-first deployment

The release includes Transformers, GGUF, MLX, and MLX 4-bit formats so it can run in Python, llama.cpp, LM Studio, and Apple Silicon workflows.


## Available Releases

| Release | Repo | Best for |
|---|---|---|
| Transformers / Safetensors | [`shuhulx/Qwopus3.5-4B-Coder-Fable5-v1`](https://huggingface.co/shuhulx/Qwopus3.5-4B-Coder-Fable5-v1) | Python, Transformers, custom inference. |
| GGUF | [`shuhulx/Qwopus3.5-4B-Coder-Fable5-v1-GGUF`](https://huggingface.co/shuhulx/Qwopus3.5-4B-Coder-Fable5-v1-GGUF) | llama.cpp, LM Studio, local CPU/GPU inference. |
| MLX | [`shuhulx/Qwopus3.5-4B-Coder-Fable5-v1-MLX`](https://huggingface.co/shuhulx/Qwopus3.5-4B-Coder-Fable5-v1-MLX) | Apple Silicon full MLX inference. |
| MLX 4-bit | [`shuhulx/Qwopus3.5-4B-Coder-Fable5-v1-MLX-4bit`](https://huggingface.co/shuhulx/Qwopus3.5-4B-Coder-Fable5-v1-MLX-4bit) | Apple Silicon low-memory inference. |

## Credits

Built on:

- [`Jackrong/Qwopus3.5-4B-Coder`](https://huggingface.co/Jackrong/Qwopus3.5-4B-Coder) by Jackrong
- [`Glint-Research/Fable-5-traces`](https://huggingface.co/datasets/Glint-Research/Fable-5-traces) by Glint-Research
- Qwen / Qwen3.5 model family
- Unsloth
- Hugging Face
- llama.cpp
- mlx-lm