File size: 2,805 Bytes
37ebf62
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
---
license: apache-2.0
base_model:
- bravesoftware/Ocelot-1-VL
- Qwen/Qwen3-VL-4B-Instruct
pipeline_tag: image-text-to-text
tags: [gguf, llama.cpp, qwen3-vl, vision, summarization]
---

# Ocelot-1-VL GGUF

Merged and quantized GGUF builds of [bravesoftware/Ocelot-1-VL](https://huggingface.co/bravesoftware/Ocelot-1-VL), a Qwen3-VL-4B LoRA specialized for webpage summarization from page text or screenshots.

The Ocelot LoRA was merged into `Qwen/Qwen3-VL-4B-Instruct` in BF16 before conversion. The vision projector is separate and is unchanged by the language-side LoRA.

## Intended use

This is not a general-purpose assistant. Use it only to summarize webpage text or webpage screenshots with the training-aligned prompt. See the [original model card](https://huggingface.co/bravesoftware/Ocelot-1-VL) for the full prompt contract, limitations, and safety guidance.

```text
The is the text of a webpage: <page>
... page plain text here ...
</page>
Summarise the content between the <page> tags, or if no content is found use the screenshots provided, in the Brave Summary style.
```

For screenshots, begin with `The following is a screenshot of a webpage:` and append the same fixed instruction.

## Files

| Quant | Approx. size | Notes |
|---|---:|---|
| `BF16` | 7.5 GB | Maximum fidelity |
| `Q8_0` | 4.0 GB | Near-BF16 quality |
| `Q6_K` | 3.1 GB | High quality |
| `Q5_K_M` / `Q5_K_S` | 2.7 / 2.6 GB | 5-bit variants |
| `Q4_K_M` / `Q4_K_S` | 2.3 / 2.2 GB | Recommended balance / smaller |
| `Q4_1` / `Q4_0` | 2.4 / 2.2 GB | Legacy compatibility |
| `IQ4_NL` / `IQ4_XS` | 2.2 / 2.1 GB | Non-linear 4-bit variants |
| `Q3_K_L` / `Q3_K_M` / `Q3_K_S` | 2.1 / 1.9 / 1.8 GB | 3-bit variants |
| `Q2_K` | 1.6 GB | Maximum compression, largest quality loss |
| `mmproj-Q8_0` / `mmproj-BF16` | 433 / 813 MB | Vision projectors |

## Run

```bash
llama-cli \
  -m Ocelot-1-VL-Q4_K_M.gguf \
  -p 'The is the text of a webpage: <page>Page text here</page> Summarise the content between the <page> tags, or if no content is found use the screenshots provided, in the Brave Summary style.' \
  -n 512 -c 8192
```

For screenshots, add `--mmproj mmproj-Ocelot-1-VL-Q8_0.gguf --image webpage.png` and use the screenshot prompt described above.

## Conversion and validation

- Adapter revision: `fbc473443eec370c3b551363556ab75d2b485db6`
- Base revision: `ebb281ec70b05090aa6165b016eac8ec08e71b17`
- llama.cpp revision: `0278d8362d78c5de291bc03b76016f7f74b2ab77`
- All 15 language files passed generation smoke tests.
- `Q4_K_M` with the Q8 projector passed screenshot summarization.
- Apple M4 Pro Q4 benchmark: 313 prompt tokens/s and 26 generation tokens/s with Metal.

Quantization can reduce output quality, especially at 2 and 3 bits. Validate against your own webpage corpus before production use.