Cubex11 commited on
Commit
b93ed7d
·
verified ·
1 Parent(s): 6ff9a2c

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +160 -125
README.md CHANGED
@@ -1,199 +1,234 @@
1
  ---
2
  library_name: transformers
3
- tags: []
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  ---
5
 
6
- # Model Card for Model ID
7
-
8
- <!-- Provide a quick summary of what the model is/does. -->
9
-
10
 
 
11
 
12
  ## Model Details
13
 
14
  ### Model Description
15
 
16
- <!-- Provide a longer summary of what this model is. -->
17
-
18
- This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
19
-
20
- - **Developed by:** [More Information Needed]
21
- - **Funded by [optional]:** [More Information Needed]
22
- - **Shared by [optional]:** [More Information Needed]
23
- - **Model type:** [More Information Needed]
24
- - **Language(s) (NLP):** [More Information Needed]
25
- - **License:** [More Information Needed]
26
- - **Finetuned from model [optional]:** [More Information Needed]
27
 
28
- ### Model Sources [optional]
 
 
 
 
29
 
30
- <!-- Provide the basic links for the model. -->
31
 
32
- - **Repository:** [More Information Needed]
33
- - **Paper [optional]:** [More Information Needed]
34
- - **Demo [optional]:** [More Information Needed]
35
 
36
  ## Uses
37
 
38
- <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
39
-
40
  ### Direct Use
41
 
42
- <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
43
-
44
- [More Information Needed]
45
-
46
- ### Downstream Use [optional]
47
 
48
- <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
49
-
50
- [More Information Needed]
51
 
52
  ### Out-of-Scope Use
53
 
54
- <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
55
-
56
- [More Information Needed]
57
-
58
- ## Bias, Risks, and Limitations
59
-
60
- <!-- This section is meant to convey both technical and sociotechnical limitations. -->
61
-
62
- [More Information Needed]
63
-
64
- ### Recommendations
65
-
66
- <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
67
-
68
- Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
69
 
70
  ## How to Get Started with the Model
71
 
72
- Use the code below to get started with the model.
73
-
74
- [More Information Needed]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
75
 
76
  ## Training Details
77
 
78
  ### Training Data
79
 
80
- <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
81
-
82
- [More Information Needed]
83
 
84
  ### Training Procedure
85
 
86
- <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
87
-
88
- #### Preprocessing [optional]
89
-
90
- [More Information Needed]
91
 
 
92
 
93
  #### Training Hyperparameters
94
 
95
- - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
96
-
97
- #### Speeds, Sizes, Times [optional]
98
-
99
- <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
100
-
101
- [More Information Needed]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
102
 
103
  ## Evaluation
104
 
105
- <!-- This section describes the evaluation protocols and provides the results. -->
106
-
107
  ### Testing Data, Factors & Metrics
108
 
109
- #### Testing Data
110
-
111
- <!-- This should link to a Dataset Card if possible. -->
112
-
113
- [More Information Needed]
114
-
115
- #### Factors
116
-
117
- <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
118
-
119
- [More Information Needed]
120
 
121
  #### Metrics
122
 
123
- <!-- These are the evaluation metrics being used, ideally with a description of why. -->
124
-
125
- [More Information Needed]
 
 
 
 
 
126
 
127
  ### Results
128
 
129
- [More Information Needed]
 
 
 
 
 
 
 
 
 
 
 
 
130
 
131
  #### Summary
132
 
 
133
 
 
 
 
 
134
 
135
- ## Model Examination [optional]
136
-
137
- <!-- Relevant interpretability work for the model goes here -->
138
 
139
- [More Information Needed]
 
 
 
140
 
141
- ## Environmental Impact
142
 
143
- <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
 
 
144
 
145
- Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
146
 
147
- - **Hardware Type:** [More Information Needed]
148
- - **Hours used:** [More Information Needed]
149
- - **Cloud Provider:** [More Information Needed]
150
- - **Compute Region:** [More Information Needed]
151
- - **Carbon Emitted:** [More Information Needed]
152
 
153
- ## Technical Specifications [optional]
154
 
155
  ### Model Architecture and Objective
156
 
157
- [More Information Needed]
 
 
158
 
159
  ### Compute Infrastructure
160
 
161
- [More Information Needed]
162
-
163
  #### Hardware
164
 
165
- [More Information Needed]
166
 
167
  #### Software
168
 
169
- [More Information Needed]
 
 
 
170
 
171
- ## Citation [optional]
172
-
173
- <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
174
 
175
  **BibTeX:**
176
 
177
- [More Information Needed]
178
-
179
- **APA:**
180
-
181
- [More Information Needed]
182
-
183
- ## Glossary [optional]
184
-
185
- <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
186
-
187
- [More Information Needed]
188
-
189
- ## More Information [optional]
190
-
191
- [More Information Needed]
192
-
193
- ## Model Card Authors [optional]
194
-
195
- [More Information Needed]
196
-
197
- ## Model Card Contact
198
-
199
- [More Information Needed]
 
1
  ---
2
  library_name: transformers
3
+ tags:
4
+ - smolvlm
5
+ - vlm
6
+ - dpo
7
+ - hallucination-reduction
8
+ - accessibility
9
+ - qlora
10
+ - rlaif
11
+ license: apache-2.0
12
+ language:
13
+ - en
14
+ pipeline_tag: image-text-to-text
15
+ base_model: HuggingFaceTB/SmolVLM2-500M-Video-Instruct
16
+ datasets:
17
+ - HuggingFaceH4/rlaif-v_formatted
18
  ---
19
 
20
+ # Solari: Hallucination-Reduced Vision Language Model
 
 
 
21
 
22
+ Solari is a 500M parameter vision-language model fine-tuned for **reduced hallucination** on real-world images. Built on [SmolVLM2-500M-Video-Instruct](https://huggingface.co/HuggingFaceTB/SmolVLM2-500M-Video-Instruct), Solari uses **QLoRA + Direct Preference Optimization (DPO)** on the [RLAIF-V](https://huggingface.co/datasets/HuggingFaceH4/rlaif-v_formatted) dataset to align the model toward more faithful visual descriptions.
23
 
24
  ## Model Details
25
 
26
  ### Model Description
27
 
28
+ Solari targets **hallucination reduction** in vision-language tasks, with a focus on improving reliability for **accessibility applications** (e.g., assisting visually impaired users). The model was trained using parameter-efficient fine-tuning (QLoRA) with DPO to learn preferences between accurate and hallucinated image descriptions, achieving improved hallucination benchmarks while preserving general VLM capabilities.
 
 
 
 
 
 
 
 
 
 
29
 
30
+ - **Developed by:** Cubex11
31
+ - **Model type:** Vision-Language Model (Image-Text-to-Text)
32
+ - **Language(s):** English
33
+ - **License:** Apache-2.0
34
+ - **Finetuned from:** [HuggingFaceTB/SmolVLM2-500M-Video-Instruct](https://huggingface.co/HuggingFaceTB/SmolVLM2-500M-Video-Instruct)
35
 
36
+ ### Model Sources
37
 
38
+ - **Base Model:** [SmolVLM2-500M-Video-Instruct](https://huggingface.co/HuggingFaceTB/SmolVLM2-500M-Video-Instruct)
39
+ - **Training Dataset:** [RLAIF-V (Formatted)](https://huggingface.co/datasets/HuggingFaceH4/rlaif-v_formatted) — 72K AI-generated preference pairs for hallucination reduction
 
40
 
41
  ## Uses
42
 
 
 
43
  ### Direct Use
44
 
45
+ Solari can be used for image understanding tasks where **factual accuracy** is critical:
 
 
 
 
46
 
47
+ - Describing real-world scenes for visually impaired users
48
+ - Visual question answering with reduced hallucination
49
+ - Image captioning with improved object recognition reliability
50
 
51
  ### Out-of-Scope Use
52
 
53
+ - Tasks requiring strong mathematical reasoning or code understanding (degraded from base model)
54
+ - Non-English language tasks
55
+ - Medical or safety-critical applications without additional validation
 
 
 
 
 
 
 
 
 
 
 
 
56
 
57
  ## How to Get Started with the Model
58
 
59
+ ```python
60
+ import torch
61
+ from transformers import AutoModelForImageTextToText, AutoProcessor
62
+ from PIL import Image
63
+ import requests
64
+
65
+ model_id = "Cubex11/Solari"
66
+ model = AutoModelForImageTextToText.from_pretrained(
67
+ model_id,
68
+ torch_dtype=torch.bfloat16,
69
+ device_map="auto",
70
+ )
71
+ processor = AutoProcessor.from_pretrained(model_id)
72
+
73
+ # Load an image
74
+ url = "https://upload.wikimedia.org/wikipedia/commons/thumb/4/47/PNG_transparency_demonstration_1.png/280px-PNG_transparency_demonstration_1.png"
75
+ image = Image.open(requests.get(url, stream=True).raw).convert("RGB")
76
+
77
+ # Create prompt
78
+ messages = [
79
+ {
80
+ "role": "user",
81
+ "content": [
82
+ {"type": "image"},
83
+ {"type": "text", "text": "Describe this image in detail."}
84
+ ]
85
+ }
86
+ ]
87
+
88
+ text = processor.apply_chat_template(messages, add_generation_prompt=True)
89
+ inputs = processor(text=text, images=[[image]], return_tensors="pt").to(model.device)
90
+ output = model.generate(**inputs, max_new_tokens=256)
91
+ trimmed = output[0][len(inputs.input_ids[0]):]
92
+ print(processor.decode(trimmed, skip_special_tokens=True))
93
+ ```
94
 
95
  ## Training Details
96
 
97
  ### Training Data
98
 
99
+ [RLAIF-V (Formatted)](https://huggingface.co/datasets/HuggingFaceH4/rlaif-v_formatted) a large-scale multimodal preference dataset containing ~72K preference pairs. Each sample includes an image, a prompt, a **chosen** response (more accurate), and a **rejected** response (more hallucinated). Preferences are generated by open-source AI models following the RLAIF-V methodology.
 
 
100
 
101
  ### Training Procedure
102
 
103
+ **Method:** QLoRA + Direct Preference Optimization (DPO)
 
 
 
 
104
 
105
+ The base model was quantized to 4-bit (NF4) and fine-tuned using Low-Rank Adaptation (LoRA) with DPO to learn preferences between accurate and hallucinated responses.
106
 
107
  #### Training Hyperparameters
108
 
109
+ | Parameter | Value |
110
+ |-----------|-------|
111
+ | **Training regime** | bf16 mixed precision |
112
+ | **Quantization** | 4-bit NF4 (double quantization) |
113
+ | **LoRA rank (r)** | 16 |
114
+ | **LoRA alpha** | 16 |
115
+ | **LoRA dropout** | 0.1 |
116
+ | **DoRA** | Enabled |
117
+ | **Target modules** | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
118
+ | **Trainable params** | ~1.9% of total |
119
+ | **Learning rate** | 5e-5 |
120
+ | **DPO beta** | 0.1 |
121
+ | **Batch size** | 8 (per device) |
122
+ | **Gradient accumulation** | 4 (effective batch = 32) |
123
+ | **Epochs** | 2 (best checkpoint at ~1 epoch / step 2500) |
124
+ | **Warmup ratio** | 0.1 |
125
+ | **Optimizer** | AdamW |
126
+
127
+ #### Speeds, Sizes, Times
128
+
129
+ - **Training time:** ~9 hours on NVIDIA L4 (24GB)
130
+ - **Best checkpoint:** Step 2500 (selected by lowest validation loss)
131
+ - **Model size:** ~1 GB (bf16 safetensors)
132
 
133
  ## Evaluation
134
 
 
 
135
  ### Testing Data, Factors & Metrics
136
 
137
+ Evaluated using [VLMEvalKit](https://github.com/open-compass/VLMEvalKit) on 8 standard benchmarks covering hallucination, general VLM capability, and real-world understanding.
 
 
 
 
 
 
 
 
 
 
138
 
139
  #### Metrics
140
 
141
+ - **POPE:** F1 score across random/popular/adversarial splits (object hallucination)
142
+ - **AMBER:** Attribute, Existence, Relation accuracy (multi-dimensional hallucination)
143
+ - **HallusionBench:** aAcc, fAcc, qAcc (hallucination detection)
144
+ - **A-OKVQA:** Accuracy on outside-knowledge VQA
145
+ - **MME:** Perception and Reasoning scores
146
+ - **MMStar:** Multi-modal reasoning accuracy
147
+ - **MMBench:** General multi-modal understanding
148
+ - **RealWorldQA:** Real-world image understanding accuracy
149
 
150
  ### Results
151
 
152
+ | Benchmark | Metric | Base Model | **Solari** | Change |
153
+ |-----------|--------|------------|------------|--------|
154
+ | **POPE** | Overall | 82.67 | **85.08** | **+2.41** |
155
+ | **POPE** | Recall | 76.73 | **85.33** | **+8.60** |
156
+ | **AMBER** | Avg ACC | 79.38 | **79.77** | **+0.39** |
157
+ | **AMBER** | Relation | 72.36 | **75.42** | **+3.06** |
158
+ | **HallusionBench** | fAcc | 18.21 | **19.36** | **+1.16** |
159
+ | **A-OKVQA** | Overall | 68.12% | **69.00%** | **+0.88%** |
160
+ | **MMStar** | Overall | 38.33% | **39.60%** | **+1.27%** |
161
+ | **MMBench** | Test | 53.14% | **53.42%** | **+0.28%** |
162
+ | **RealWorldQA** | Overall | 49.80% | **50.59%** | **+0.78%** |
163
+ | **MME** | Perception | **1216.19** | 1118.51 | -97.68 |
164
+ | **MME** | Reasoning | **237.50** | 211.79 | -25.71 |
165
 
166
  #### Summary
167
 
168
+ Solari improves on **7 out of 8 benchmarks** compared to the base model:
169
 
170
+ - **POPE recall +8.60%** — dramatically better at recognizing objects actually present in images
171
+ - **All hallucination benchmarks improved** — POPE, AMBER, and HallusionBench
172
+ - **General capabilities preserved or improved** — A-OKVQA, MMStar, MMBench, RealWorldQA all show gains
173
+ - **Trade-off on MME** — perception score dropped ~98 points, primarily on counting (-26.7), position (-26.7), and code reasoning (-27.5) subtasks due to the model becoming more conservative
174
 
175
+ ## Bias, Risks, and Limitations
 
 
176
 
177
+ - **Counting and spatial reasoning degraded:** The DPO alignment made the model more conservative, reducing performance on fine-grained counting and positional reasoning tasks (reflected in MME scores).
178
+ - **Small model capacity:** At 500M parameters, the model has inherent limitations on complex reasoning tasks.
179
+ - **English only:** The model was trained and evaluated only on English-language tasks.
180
+ - **Training data bias:** RLAIF-V preferences are AI-generated, which may introduce systematic biases.
181
 
182
+ ### Recommendations
183
 
184
+ - Best suited for binary object recognition tasks ("Is there a X?") and general scene description
185
+ - For tasks requiring precise counting or spatial reasoning, consider using the base model or a larger VLM
186
+ - Always validate outputs in safety-critical applications
187
 
188
+ ## Environmental Impact
189
 
190
+ - **Hardware Type:** NVIDIA L4 (24GB)
191
+ - **Hours used:** ~9 hours
192
+ - **Cloud Provider:** Lightning AI
193
+ - **Compute Region:** US
 
194
 
195
+ ## Technical Specifications
196
 
197
  ### Model Architecture and Objective
198
 
199
+ - **Architecture:** SmolVLM2 (ViT vision encoder + LLM decoder with multi-modal projector)
200
+ - **Parameters:** ~500M total
201
+ - **Objective:** Direct Preference Optimization (DPO) — learns to prefer accurate descriptions over hallucinated ones
202
 
203
  ### Compute Infrastructure
204
 
 
 
205
  #### Hardware
206
 
207
+ NVIDIA L4 GPU (24GB VRAM) on Lightning AI
208
 
209
  #### Software
210
 
211
+ - Transformers
212
+ - TRL (DPO Trainer)
213
+ - PEFT (QLoRA)
214
+ - BitsAndBytes (4-bit quantization)
215
 
216
+ ## Citation
 
 
217
 
218
  **BibTeX:**
219
 
220
+ ```bibtex
221
+ @misc{solari2026,
222
+ title={Solari: Hallucination-Reduced Vision Language Model via QLoRA DPO on RLAIF-V},
223
+ author={Cubex11},
224
+ year={2026},
225
+ url={https://huggingface.co/Cubex11/Solari}
226
+ }
227
+ ```
228
+
229
+ ## Acknowledgments
230
+
231
+ - [HuggingFace](https://huggingface.co/) for SmolVLM2 and the RLAIF-V formatted dataset
232
+ - [OpenBMB](https://github.com/OpenBMB) for the RLAIF-V and RLHF-V research
233
+ - [Lightning AI](https://lightning.ai/) for compute resources
234
+ - [OpenCompass](https://github.com/open-compass/VLMEvalKit) for the VLMEvalKit evaluation toolkit