Multimodal LLMs See Sentiment
Paper • 2508.16873 • Published
Fine-tuned text classifiers from "Multimodal LLMs See Sentiment" (arXiv:2508.16873). Each checkpoint scores the sentiment of an image description produced by a multimodal LLM, which is the second stage of the MLLMsent pipeline.
image ──▶ multimodal LLM ──▶ description ──▶ text classifier ──▶ sentiment
(GPT-4o mini, Gemini, (these checkpoints:
DeepSeek-VL2, Phi-4, ModernBERT-large,
Gemma-4, MiniGPT-4) BART-large-MNLI)
The paper's best configuration is GPT-4o mini captions + fine-tuned ModernBERT.
{caption_mllm}/{backbone}/{problem}/sigma{n}/{finetuned|not_finetuned}/
model.safetensors
config.json
MANIFEST.json
p5 (5 classes), p3 (3), p2plus/p2neg (2).Every config.json carries the base model id, id2label/label2id, the source
checkpoint's SHA-256 and the 5-fold scores that checkpoint achieved.
| caption MLLM | base model | checkpoints |
|---|---|---|
deepseek |
answerdotai/ModernBERT-large |
8 |
deepseek |
facebook/bart-large-mnli |
6 |
gemini |
answerdotai/ModernBERT-large |
8 |
gemma4 |
answerdotai/ModernBERT-large |
8 |
minigpt4 |
answerdotai/ModernBERT-large |
8 |
minigpt4 |
facebook/bart-large-mnli |
6 |
openai |
answerdotai/ModernBERT-large |
12 |
openai |
facebook/bart-large-mnli |
8 |
phi4 |
answerdotai/ModernBERT-large |
8 |
Weights are fp16 safetensors converted from the original fp32 training checkpoints.
| checkpoint | track | mean 5-fold F1 | classes |
|---|---|---|---|
| openai-modernbert-p3-sigma5 | finetuning | 0.9581 | 3 |
| openai-bart-p3-sigma5 | finetuning | 0.9532 | 3 |
| gemini-modernbert-p3-sigma5 | finetuning | 0.9451 | 3 |
| gemma4-modernbert-p3-sigma5 | finetuning | 0.9417 | 3 |
| phi4-modernbert-p3-sigma5 | finetuning | 0.9332 | 3 |
| openai-modernbert-p3-sigma5 | not-finetuning | 0.9063 | 3 |
| minigpt4-modernbert-p3-sigma5 | finetuning | 0.9039 | 3 |
| deepseek-modernbert-p3-sigma5 | finetuning | 0.8948 | 3 |
| openai-bart-p3-sigma5 | not-finetuning | 0.8543 | 3 |
| openai-modernbert-p5-sigma5 | finetuning | 0.8445 | 5 |
import json, torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from transformers import AutoModel, AutoTokenizer
repo = "Neemias/multimodal-LLMs-See-Sentiment"
folder = "gpt4-openai-classify/modernbert/p3/sigma5/finetuned"
weights = load_file(hf_hub_download(repo, f"{folder}/model.safetensors"))
config = json.load(open(hf_hub_download(repo, f"{folder}/config.json")))
for alias, owner in config["tied_weights"].items():
weights[alias] = weights[owner]
class SentimentClassifier(torch.nn.Module):
def __init__(self, base_model, num_classes):
super().__init__()
self.model = AutoModel.from_pretrained(base_model)
self.classifier = torch.nn.Sequential(
torch.nn.Linear(self.model.config.hidden_size, 1024),
torch.nn.ReLU(),
torch.nn.Linear(1024, num_classes),
)
def forward(self, ids, mask):
return self.classifier(self.model(ids, attention_mask=mask).last_hidden_state[:, 0])
model = SentimentClassifier(config["base_model"], config["num_classes"])
model.load_state_dict({k: v.float() for k, v in weights.items()})
model.eval()
tokenizer = AutoTokenizer.from_pretrained(config["base_model"])
batch = tokenizer(["A bright park full of children playing."], return_tensors="pt",
padding="max_length", truncation=True, max_length=config["max_len"])
prediction = model(batch["input_ids"], batch["attention_mask"]).argmax(-1).item()
print(config["id2label"][str(prediction)])
Or through the project CLI:
mllmsent hub pull-checkpoint openai-modernbert-p3-sigma5
mllmsent predict --spec openai-modernbert-p3-sigma5 --input captions.csv --output predictions.csv
adapter_config.json survive.@misc{dasilva2026multimodalllmssentiment,
title={Multimodal LLMs See Sentiment},
author={Neemias B. da Silva and John Harrison and Rodrigo Minetto and Myriam R. Delgado and Bogdan T. Nassu and Thiago H. Silva},
year={2026},
eprint={2508.16873},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2508.16873},
}
CC-BY-4.0. The base models keep their own licenses.
Base model
answerdotai/ModernBERT-large