metadata
license: cc-by-4.0
arxiv: 2508.16873
library_name: pytorch
pipeline_tag: text-classification
tags:
- sentiment-analysis
- multimodal
- image-sentiment
- perceptsent
- mllm
language:
- en
base_model:
- answerdotai/ModernBERT-large
- facebook/bart-large-mnli
MLLMsent — sentiment classifiers over multimodal-LLM image descriptions
Fine-tuned text classifiers from "Multimodal LLMs See Sentiment" (arXiv:2508.16873). Each checkpoint scores the sentiment of an image description produced by a multimodal LLM, which is the second stage of the MLLMsent pipeline.
- Paper: arXiv:2508.16873
- Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment
- Datasets (inputs, captions and every result CSV): https://huggingface.co/datasets/Neemias/multimodal-LLMs-See-Sentiment
Pipeline
image ──▶ multimodal LLM ──▶ description ──▶ text classifier ──▶ sentiment
(GPT-4o mini, Gemini, (these checkpoints:
DeepSeek-VL2, Phi-4, ModernBERT-large,
Gemma-4, MiniGPT-4) BART-large-MNLI)
The paper's best configuration is GPT-4o mini captions + fine-tuned ModernBERT.
Layout
{caption_mllm}/{backbone}/{problem}/sigma{n}/{finetuned|not_finetuned}/
model.safetensors
config.json
MANIFEST.json
- problem — label granularity:
p5(5 classes),p3(3),p2plus/p2neg(2). - sigma — annotator-agreement threshold used to filter the training set (3 or 5).
- finetuned — whole backbone trained. not_finetuned — backbone frozen, head only.
Every config.json carries the base model id, id2label/label2id, the source
checkpoint's SHA-256 and the 5-fold scores that checkpoint achieved.
Coverage
| caption MLLM | base model | checkpoints |
|---|---|---|
deepseek |
answerdotai/ModernBERT-large |
8 |
deepseek |
facebook/bart-large-mnli |
6 |
gemini |
answerdotai/ModernBERT-large |
8 |
gemma4 |
answerdotai/ModernBERT-large |
8 |
minigpt4 |
answerdotai/ModernBERT-large |
8 |
minigpt4 |
facebook/bart-large-mnli |
6 |
openai |
answerdotai/ModernBERT-large |
12 |
openai |
facebook/bart-large-mnli |
8 |
phi4 |
answerdotai/ModernBERT-large |
8 |
Weights are fp16 safetensors converted from the original fp32 training checkpoints.
Best checkpoints
| checkpoint | track | mean 5-fold F1 | classes |
|---|---|---|---|
| openai-modernbert-p3-sigma5 | finetuning | 0.9581 | 3 |
| openai-bart-p3-sigma5 | finetuning | 0.9532 | 3 |
| gemini-modernbert-p3-sigma5 | finetuning | 0.9451 | 3 |
| gemma4-modernbert-p3-sigma5 | finetuning | 0.9417 | 3 |
| phi4-modernbert-p3-sigma5 | finetuning | 0.9332 | 3 |
| openai-modernbert-p3-sigma5 | not-finetuning | 0.9063 | 3 |
| minigpt4-modernbert-p3-sigma5 | finetuning | 0.9039 | 3 |
| deepseek-modernbert-p3-sigma5 | finetuning | 0.8948 | 3 |
| openai-bart-p3-sigma5 | not-finetuning | 0.8543 | 3 |
| openai-modernbert-p5-sigma5 | finetuning | 0.8445 | 5 |
Usage
import json, torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from transformers import AutoModel, AutoTokenizer
repo = "Neemias/multimodal-LLMs-See-Sentiment"
folder = "gpt4-openai-classify/modernbert/p3/sigma5/finetuned"
weights = load_file(hf_hub_download(repo, f"{folder}/model.safetensors"))
config = json.load(open(hf_hub_download(repo, f"{folder}/config.json")))
for alias, owner in config["tied_weights"].items():
weights[alias] = weights[owner]
class SentimentClassifier(torch.nn.Module):
def __init__(self, base_model, num_classes):
super().__init__()
self.model = AutoModel.from_pretrained(base_model)
self.classifier = torch.nn.Sequential(
torch.nn.Linear(self.model.config.hidden_size, 1024),
torch.nn.ReLU(),
torch.nn.Linear(1024, num_classes),
)
def forward(self, ids, mask):
return self.classifier(self.model(ids, attention_mask=mask).last_hidden_state[:, 0])
model = SentimentClassifier(config["base_model"], config["num_classes"])
model.load_state_dict({k: v.float() for k, v in weights.items()})
model.eval()
tokenizer = AutoTokenizer.from_pretrained(config["base_model"])
batch = tokenizer(["A bright park full of children playing."], return_tensors="pt",
padding="max_length", truncation=True, max_length=config["max_len"])
prediction = model(batch["input_ids"], batch["attention_mask"]).argmax(-1).item()
print(config["id2label"][str(prediction)])
Or through the project CLI:
mllmsent hub pull-checkpoint openai-modernbert-p3-sigma5
mllmsent predict --spec openai-modernbert-p3-sigma5 --input captions.csv --output predictions.csv
Not published here
- LLaMA-3 qLoRA adapters — the adapter weights were never retained; only the
training logs and
adapter_config.jsonsurvive. - Swin Transformer baseline — its checkpoint-saving path was broken, so no weights were ever written. Results for it are in the dataset repo.
- A few BART sigma-5 fine-tuned cells, for the same reason.
Citation
@misc{dasilva2026multimodalllmssentiment,
title={Multimodal LLMs See Sentiment},
author={Neemias B. da Silva and John Harrison and Rodrigo Minetto and Myriam R. Delgado and Bogdan T. Nassu and Thiago H. Silva},
year={2026},
eprint={2508.16873},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2508.16873},
}
License
CC-BY-4.0. The base models keep their own licenses.