| --- |
| license: cc-by-4.0 |
| arxiv: 2508.16873 |
| library_name: pytorch |
| pipeline_tag: text-classification |
| tags: |
| - sentiment-analysis |
| - multimodal |
| - image-sentiment |
| - perceptsent |
| - mllm |
| language: |
| - en |
| base_model: |
| - answerdotai/ModernBERT-large |
| - facebook/bart-large-mnli |
| --- |
| |
| # MLLMsent — sentiment classifiers over multimodal-LLM image descriptions |
|
|
| Fine-tuned text classifiers from **"Multimodal LLMs See Sentiment"** |
| ([arXiv:2508.16873](https://arxiv.org/abs/2508.16873)). Each checkpoint scores the sentiment of an image |
| *description* produced by a multimodal LLM, which is the second stage of the MLLMsent |
| pipeline. |
|
|
| - **Paper:** [arXiv:2508.16873](https://arxiv.org/abs/2508.16873) |
| - **Code, training and inference:** https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment |
| - **Datasets (inputs, captions and every result CSV):** https://huggingface.co/datasets/Neemias/multimodal-LLMs-See-Sentiment |
|
|
| ## Pipeline |
|
|
| ``` |
| image ──▶ multimodal LLM ──▶ description ──▶ text classifier ──▶ sentiment |
| (GPT-4o mini, Gemini, (these checkpoints: |
| DeepSeek-VL2, Phi-4, ModernBERT-large, |
| Gemma-4, MiniGPT-4) BART-large-MNLI) |
| ``` |
|
|
| The paper's best configuration is **GPT-4o mini captions + fine-tuned ModernBERT**. |
|
|
| ## Layout |
|
|
| ``` |
| {caption_mllm}/{backbone}/{problem}/sigma{n}/{finetuned|not_finetuned}/ |
| model.safetensors |
| config.json |
| MANIFEST.json |
| ``` |
|
|
| - **problem** — label granularity: `p5` (5 classes), `p3` (3), `p2plus`/`p2neg` (2). |
| - **sigma** — annotator-agreement threshold used to filter the training set (3 or 5). |
| - **finetuned** — whole backbone trained. **not_finetuned** — backbone frozen, head only. |
| |
| Every `config.json` carries the base model id, `id2label`/`label2id`, the source |
| checkpoint's SHA-256 and the 5-fold scores that checkpoint achieved. |
| |
| ## Coverage |
| |
| | caption MLLM | base model | checkpoints | |
| |---|---|---| |
| | `deepseek` | `answerdotai/ModernBERT-large` | 8 | |
| | `deepseek` | `facebook/bart-large-mnli` | 6 | |
| | `gemini` | `answerdotai/ModernBERT-large` | 8 | |
| | `gemma4` | `answerdotai/ModernBERT-large` | 8 | |
| | `minigpt4` | `answerdotai/ModernBERT-large` | 8 | |
| | `minigpt4` | `facebook/bart-large-mnli` | 6 | |
| | `openai` | `answerdotai/ModernBERT-large` | 12 | |
| | `openai` | `facebook/bart-large-mnli` | 8 | |
| | `phi4` | `answerdotai/ModernBERT-large` | 8 | |
| |
| Weights are **fp16 safetensors** converted from the original fp32 training checkpoints. |
| |
| ## Best checkpoints |
| |
| | checkpoint | track | mean 5-fold F1 | classes | |
| |---|---|---|---| |
| | openai-modernbert-p3-sigma5 | finetuning | 0.9581 | 3 | |
| | openai-bart-p3-sigma5 | finetuning | 0.9532 | 3 | |
| | gemini-modernbert-p3-sigma5 | finetuning | 0.9451 | 3 | |
| | gemma4-modernbert-p3-sigma5 | finetuning | 0.9417 | 3 | |
| | phi4-modernbert-p3-sigma5 | finetuning | 0.9332 | 3 | |
| | openai-modernbert-p3-sigma5 | not-finetuning | 0.9063 | 3 | |
| | minigpt4-modernbert-p3-sigma5 | finetuning | 0.9039 | 3 | |
| | deepseek-modernbert-p3-sigma5 | finetuning | 0.8948 | 3 | |
| | openai-bart-p3-sigma5 | not-finetuning | 0.8543 | 3 | |
| | openai-modernbert-p5-sigma5 | finetuning | 0.8445 | 5 | |
| |
| ## Usage |
| |
| ```python |
| import json, torch |
| from huggingface_hub import hf_hub_download |
| from safetensors.torch import load_file |
| from transformers import AutoModel, AutoTokenizer |
| |
| repo = "Neemias/multimodal-LLMs-See-Sentiment" |
| folder = "gpt4-openai-classify/modernbert/p3/sigma5/finetuned" |
| |
| weights = load_file(hf_hub_download(repo, f"{folder}/model.safetensors")) |
| config = json.load(open(hf_hub_download(repo, f"{folder}/config.json"))) |
| |
| for alias, owner in config["tied_weights"].items(): |
| weights[alias] = weights[owner] |
| |
| |
| class SentimentClassifier(torch.nn.Module): |
| def __init__(self, base_model, num_classes): |
| super().__init__() |
| self.model = AutoModel.from_pretrained(base_model) |
| self.classifier = torch.nn.Sequential( |
| torch.nn.Linear(self.model.config.hidden_size, 1024), |
| torch.nn.ReLU(), |
| torch.nn.Linear(1024, num_classes), |
| ) |
| |
| def forward(self, ids, mask): |
| return self.classifier(self.model(ids, attention_mask=mask).last_hidden_state[:, 0]) |
| |
| |
| model = SentimentClassifier(config["base_model"], config["num_classes"]) |
| model.load_state_dict({k: v.float() for k, v in weights.items()}) |
| model.eval() |
| |
| tokenizer = AutoTokenizer.from_pretrained(config["base_model"]) |
| batch = tokenizer(["A bright park full of children playing."], return_tensors="pt", |
| padding="max_length", truncation=True, max_length=config["max_len"]) |
| prediction = model(batch["input_ids"], batch["attention_mask"]).argmax(-1).item() |
| print(config["id2label"][str(prediction)]) |
| ``` |
| |
| Or through the project CLI: |
| |
| ```bash |
| mllmsent hub pull-checkpoint openai-modernbert-p3-sigma5 |
| mllmsent predict --spec openai-modernbert-p3-sigma5 --input captions.csv --output predictions.csv |
| ``` |
| |
| ## Not published here |
| |
| - **LLaMA-3 qLoRA adapters** — the adapter weights were never retained; only the |
| training logs and `adapter_config.json` survive. |
| - **Swin Transformer baseline** — its checkpoint-saving path was broken, so no |
| weights were ever written. Results for it are in the dataset repo. |
| - A few BART sigma-5 fine-tuned cells, for the same reason. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{dasilva2026multimodalllmssentiment, |
| title={Multimodal LLMs See Sentiment}, |
| author={Neemias B. da Silva and John Harrison and Rodrigo Minetto and Myriam R. Delgado and Bogdan T. Nassu and Thiago H. Silva}, |
| year={2026}, |
| eprint={2508.16873}, |
| archivePrefix={arXiv}, |
| primaryClass={cs.CV}, |
| url={https://arxiv.org/abs/2508.16873}, |
| } |
| ``` |
|
|
| ## License |
|
|
| CC-BY-4.0. The base models keep their own licenses. |
|
|