Image-Text-to-Text
Transformers
Safetensors
English
qwen3_5
liveledger
epistemic-ledger
agent
tool-use
function-calling
deep-research
conversational
Instructions to use dayoon/LiveLedger-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dayoon/LiveLedger-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="dayoon/LiveLedger-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("dayoon/LiveLedger-4B") model = AutoModelForMultimodalLM.from_pretrained("dayoon/LiveLedger-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dayoon/LiveLedger-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dayoon/LiveLedger-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dayoon/LiveLedger-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/dayoon/LiveLedger-4B
- SGLang
How to use dayoon/LiveLedger-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dayoon/LiveLedger-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dayoon/LiveLedger-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dayoon/LiveLedger-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dayoon/LiveLedger-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use dayoon/LiveLedger-4B with Docker Model Runner:
docker model run hf.co/dayoon/LiveLedger-4B
File size: 7,408 Bytes
b99da57 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 | ---
license: apache-2.0
base_model: Qwen/Qwen3.5-4B
base_model_relation: finetune
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- liveledger
- epistemic-ledger
- agent
- tool-use
- function-calling
- deep-research
- qwen3_5
language:
- en
---
# LiveLedger-4B
**LiveLedger-4B** is the *ledger-update* module of **LiveLedger**, an epistemic agent framework for
answering multi-constraint questions with iterative web search and structured evidence tracking.
LiveLedger decomposes a question into atomic constraints and maintains an **Epistemic Ledger** — a
`(candidate × constraint)` table whose cells record whether a candidate satisfies a constraint and
the evidence for that verdict. Keeping this bookkeeping in an explicit structure, rather than in the
agent's context window, is what lets the agent stay coherent over long search trajectories.
Running the ledger update with the frontier search model is expensive: it is called after *every*
search step. LiveLedger-4B is a small, distilled specialist that takes over exactly that step. It is
supervised-fine-tuned on ledger-update trajectories produced by **gpt-oss-120b**, so a large model
drives search while this 4B model does the per-step evidence bookkeeping.
> This is **not** a general-purpose chat or search model. It is trained for one job: read search
> results and emit an `update_ledger` tool call. Used outside that harness, its behaviour is
> undefined.
## Model details
| | |
|---|---|
| Base model | [`Qwen/Qwen3.5-4B`](https://huggingface.co/Qwen/Qwen3.5-4B) |
| Parameters | 4.54 B (incl. the inherited vision tower) |
| Architecture | `Qwen3_5ForConditionalGeneration` |
| Precision | bfloat16 |
| Context length | 262,144 |
| Role in LiveLedger | ledger-update step (`update_ledger` tool call) |
| Teacher | gpt-oss-120b |
The base checkpoint is multimodal and the vision tower is carried over **unmodified** — only the
language-model weights were fine-tuned. LiveLedger is a text-only pipeline; the vision weights are
kept solely so the checkpoint stays loadable as a stock `Qwen3.5-4B`.
## Training
Full-parameter SFT (no LoRA) with DeepSpeed ZeRO-3 + optimizer offload.
| Hyperparameter | Value |
|---|---|
| Objective | supervised fine-tuning on assistant turns |
| Learning rate | 3e-5, cosine schedule, warmup ratio 0.05 |
| Weight decay | 0.1 |
| Gradient clipping | 1.0 |
| Batch | 1 per device × 8 grad-accum × 16 GPUs = **128** effective |
| Max sequence length | 8,192 (longer examples dropped) |
| Precision | bf16 |
| Seed | 42 |
**Data.** Tool-calling trajectories distilled from gpt-oss-120b LiveLedger rollouts on HDS-QA and
MultiConIR-books: 2,885 `update_ledger` examples and 1,393 constraint-extraction examples. After
dropping examples over 8,192 tokens the corpus is ~3.7K, split 90/10 into train/validation.
**Released checkpoint.** The run was scheduled for 10 epochs (260 steps) but validation loss bottoms
out early and rises steadily afterwards, so the released weights are **step 52 (epoch 2)**:
| Step | Epoch | Validation loss | Token accuracy |
|---|---|---|---|
| 26 | 1 | 0.9179 | 0.7405 |
| **52** | **2** | **0.9188** | **0.7417** |
| 78 | 3 | 0.9620 | 0.7370 |
| 104 | 4 | 1.0645 | 0.7291 |
| 130 | 5 | 1.2241 | 0.7194 |
| 156 | 6 | 1.4123 | 0.7115 |
Step 26 has a marginally lower loss, but step 52 has higher token accuracy and is the checkpoint used
for every reported LiveLedger result.
## Usage
### vLLM (how it is actually served)
The ledger model runs as its own endpoint alongside the search model. Tool-call and reasoning parsers
are required — the agent reads the structured `update_ledger` call, not free text.
```bash
vllm serve dayoon/LiveLedger-4B \
--port 8100 \
--tensor-parallel-size 1 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--max-num-seqs 16
```
### Transformers
```python
import torch
from transformers import AutoTokenizer, Qwen3_5ForConditionalGeneration
model_id = "dayoon/LiveLedger-4B"
tok = AutoTokenizer.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
model_id, dtype=torch.bfloat16, device_map="auto"
)
tools = [{
"type": "function",
"function": {
"name": "update_ledger",
"description": (
"Record evidence in the ledger based on search results.\n\n"
"Each entry records evidence for ONE constraint of ONE candidate.\n"
"Only include entries where new evidence was found."
),
"parameters": {
"type": "object",
"properties": {
"entries": {
"type": "array",
"items": {
"type": "object",
"properties": {
"candidate": {"type": "string"},
"constraint": {"type": "string"},
"obj": {"type": "boolean"},
"obj_evidence": {"type": "string"},
},
"required": ["candidate", "constraint", "obj", "obj_evidence"],
},
}
},
"required": ["entries"],
},
},
}]
messages = [
{"role": "system", "content": "You are a careful and thorough ledger update assistant."},
{"role": "user", "content": (
"Constraints:\nC1: Directed by Christopher Nolan\nC2: Released in 2010\n\n"
"Current ledger: {}\n\n"
"Search results: 'Inception is a 2010 science fiction film written and "
"directed by Christopher Nolan.'\n\nUpdate the ledger."
)},
]
text = tok.apply_chat_template(messages, tools=tools, tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=False))
```
The model reasons about the search results, then emits an `update_ledger` call whose `entries` set
`obj: true/false` per `(candidate, constraint)` pair with a supporting `obj_evidence` quote.
## Intended use and limitations
**Intended use.** The ledger-update step of a LiveLedger-style agent: given constraints, the current
ledger, a search query and its retrieval results, decide which `(candidate, constraint)` cells the
new evidence settles.
**Limitations.**
- **Narrow specialist.** Fine-tuned on a single tool-calling format. It is not a chat assistant and
will not behave like the `Qwen3.5-4B` it came from on general tasks.
- **Trained in English** on English web content.
- **Evidence is extracted, not verified.** The model reports what retrieved text asserts. Wrong,
stale or adversarial sources propagate into the ledger as confident verdicts.
- **Distillation ceiling.** Behaviour is bounded by the gpt-oss-120b trajectories it learned from,
including their errors.
- **Small training corpus** (~3.7K examples) drawn from two source datasets; out-of-domain
constraint types are less reliable.
- **Vision inputs are untested.** The vision tower is inherited, never fine-tuned, and never
exercised in LiveLedger.
## License
Apache 2.0, inherited from [`Qwen/Qwen3.5-4B`](https://huggingface.co/Qwen/Qwen3.5-4B). The bundled
`LICENSE` file is the base model's.
|