Neemias's picture
Update README.md
4ec4bd6 verified
|
Raw
History Blame Contribute Delete
5.74 kB
---
license: cc-by-4.0
arxiv: 2508.16873
library_name: pytorch
pipeline_tag: text-classification
tags:
- sentiment-analysis
- multimodal
- image-sentiment
- perceptsent
- mllm
language:
- en
base_model:
- answerdotai/ModernBERT-large
- facebook/bart-large-mnli
---
# MLLMsent — sentiment classifiers over multimodal-LLM image descriptions
Fine-tuned text classifiers from **"Multimodal LLMs See Sentiment"**
([arXiv:2508.16873](https://arxiv.org/abs/2508.16873)). Each checkpoint scores the sentiment of an image
*description* produced by a multimodal LLM, which is the second stage of the MLLMsent
pipeline.
- **Paper:** [arXiv:2508.16873](https://arxiv.org/abs/2508.16873)
- **Code, training and inference:** https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment
- **Datasets (inputs, captions and every result CSV):** https://huggingface.co/datasets/Neemias/multimodal-LLMs-See-Sentiment
## Pipeline
```
image ──▶ multimodal LLM ──▶ description ──▶ text classifier ──▶ sentiment
(GPT-4o mini, Gemini, (these checkpoints:
DeepSeek-VL2, Phi-4, ModernBERT-large,
Gemma-4, MiniGPT-4) BART-large-MNLI)
```
The paper's best configuration is **GPT-4o mini captions + fine-tuned ModernBERT**.
## Layout
```
{caption_mllm}/{backbone}/{problem}/sigma{n}/{finetuned|not_finetuned}/
model.safetensors
config.json
MANIFEST.json
```
- **problem** — label granularity: `p5` (5 classes), `p3` (3), `p2plus`/`p2neg` (2).
- **sigma** — annotator-agreement threshold used to filter the training set (3 or 5).
- **finetuned** — whole backbone trained. **not_finetuned** — backbone frozen, head only.
Every `config.json` carries the base model id, `id2label`/`label2id`, the source
checkpoint's SHA-256 and the 5-fold scores that checkpoint achieved.
## Coverage
| caption MLLM | base model | checkpoints |
|---|---|---|
| `deepseek` | `answerdotai/ModernBERT-large` | 8 |
| `deepseek` | `facebook/bart-large-mnli` | 6 |
| `gemini` | `answerdotai/ModernBERT-large` | 8 |
| `gemma4` | `answerdotai/ModernBERT-large` | 8 |
| `minigpt4` | `answerdotai/ModernBERT-large` | 8 |
| `minigpt4` | `facebook/bart-large-mnli` | 6 |
| `openai` | `answerdotai/ModernBERT-large` | 12 |
| `openai` | `facebook/bart-large-mnli` | 8 |
| `phi4` | `answerdotai/ModernBERT-large` | 8 |
Weights are **fp16 safetensors** converted from the original fp32 training checkpoints.
## Best checkpoints
| checkpoint | track | mean 5-fold F1 | classes |
|---|---|---|---|
| openai-modernbert-p3-sigma5 | finetuning | 0.9581 | 3 |
| openai-bart-p3-sigma5 | finetuning | 0.9532 | 3 |
| gemini-modernbert-p3-sigma5 | finetuning | 0.9451 | 3 |
| gemma4-modernbert-p3-sigma5 | finetuning | 0.9417 | 3 |
| phi4-modernbert-p3-sigma5 | finetuning | 0.9332 | 3 |
| openai-modernbert-p3-sigma5 | not-finetuning | 0.9063 | 3 |
| minigpt4-modernbert-p3-sigma5 | finetuning | 0.9039 | 3 |
| deepseek-modernbert-p3-sigma5 | finetuning | 0.8948 | 3 |
| openai-bart-p3-sigma5 | not-finetuning | 0.8543 | 3 |
| openai-modernbert-p5-sigma5 | finetuning | 0.8445 | 5 |
## Usage
```python
import json, torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from transformers import AutoModel, AutoTokenizer
repo = "Neemias/multimodal-LLMs-See-Sentiment"
folder = "gpt4-openai-classify/modernbert/p3/sigma5/finetuned"
weights = load_file(hf_hub_download(repo, f"{folder}/model.safetensors"))
config = json.load(open(hf_hub_download(repo, f"{folder}/config.json")))
for alias, owner in config["tied_weights"].items():
weights[alias] = weights[owner]
class SentimentClassifier(torch.nn.Module):
def __init__(self, base_model, num_classes):
super().__init__()
self.model = AutoModel.from_pretrained(base_model)
self.classifier = torch.nn.Sequential(
torch.nn.Linear(self.model.config.hidden_size, 1024),
torch.nn.ReLU(),
torch.nn.Linear(1024, num_classes),
)
def forward(self, ids, mask):
return self.classifier(self.model(ids, attention_mask=mask).last_hidden_state[:, 0])
model = SentimentClassifier(config["base_model"], config["num_classes"])
model.load_state_dict({k: v.float() for k, v in weights.items()})
model.eval()
tokenizer = AutoTokenizer.from_pretrained(config["base_model"])
batch = tokenizer(["A bright park full of children playing."], return_tensors="pt",
padding="max_length", truncation=True, max_length=config["max_len"])
prediction = model(batch["input_ids"], batch["attention_mask"]).argmax(-1).item()
print(config["id2label"][str(prediction)])
```
Or through the project CLI:
```bash
mllmsent hub pull-checkpoint openai-modernbert-p3-sigma5
mllmsent predict --spec openai-modernbert-p3-sigma5 --input captions.csv --output predictions.csv
```
## Not published here
- **LLaMA-3 qLoRA adapters** — the adapter weights were never retained; only the
training logs and `adapter_config.json` survive.
- **Swin Transformer baseline** — its checkpoint-saving path was broken, so no
weights were ever written. Results for it are in the dataset repo.
- A few BART sigma-5 fine-tuned cells, for the same reason.
## Citation
```bibtex
@misc{dasilva2026multimodalllmssentiment,
title={Multimodal LLMs See Sentiment},
author={Neemias B. da Silva and John Harrison and Rodrigo Minetto and Myriam R. Delgado and Bogdan T. Nassu and Thiago H. Silva},
year={2026},
eprint={2508.16873},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2508.16873},
}
```
## License
CC-BY-4.0. The base models keep their own licenses.