SciGram-7B
SciGram-7B is a 7B-class vision-language model specialized in understanding scientific diagrams. It is based on a LLaVA-style architecture with a pretrained CLIP vision encoder and a Qwen2-Instruct 7B language model, and is trained on SciGram, a large-scale dataset of scientific diagrams and synthetic visual instructions covering life, earth, and physical sciences.
The model is designed for multimodal understanding tasks involving scientific diagrams, including diagram question answering, visual grounding, diagram description, and science-related visual reasoning.
SciGram-7B corresponds to the LLaVA-SciGram 7B model described in:
Raul Ortega and José Manuel Gómez-Pérez. From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding. Proceedings of the Conference on Language Modeling (COLM 2026).
Model Details
Model Description
SciGram-7B is a vision-language model trained to improve scientific diagram understanding. It uses a pretrained CLIP vision encoder together with a Qwen2-Instruct 7B language model. The model is trained using three SciGram training stages: SciGram-Align, SciGram-VIT, and SciGram-M³.
The training data contains scientific diagrams paired with synthetic captions and multiple-choice visual instructions covering life sciences, earth sciences, and physical sciences.
The model is intended primarily for research on scientific diagram understanding and multimodal scientific question answering.
- Developed by: Expert.ai Language Technology Research Laboratory
- Authors: Raul Ortega and José Manuel Gómez-Pérez
- Shared by: Expert.ai Language Technology Research Laboratory
- Model type: Vision-language model (VLM)
- Architecture: LLaVA-Qwen
- Language(s): English
- License: CC BY 4.0
- Vision encoder: Pretrained CLIP vision encoder
- Language model: Qwen2-Instruct 7B
- Model format: Safetensors, FP16
- Maximum context length: 32,768 tokens
The released checkpoint contains approximately 8B parameters and is approximately 16.3 GB in size.
Model Sources
- Repository: https://github.com/expertailab/scigram
- Dataset: https://huggingface.co/datasets/expertailab/scigram_dataset
- Paper: https://openreview.net/pdf?id=Xdi2q5XKgB
How to Get Started with the Model
The model is released using the LLaVA-Qwen architecture. It can be loaded using a compatible LLaVA implementation and the Hugging Face checkpoint.
A typical workflow is:
from transformers import AutoProcessor, AutoModel
from PIL import Image
import torch
model_id = "expertailab/scigram-7b"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModel.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto",
)
image = Image.open("scientific_diagram.png").convert("RGB")
prompt = "What does this scientific diagram show? Explain the main components and their relationships."
inputs = processor(
images=image,
text=prompt,
return_tensors="pt"
).to(model.device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=512,
)
answer = processor.batch_decode(
output,
skip_special_tokens=True
)[0]
print(answer)
The model repository contains a LlavaProcessor configuration and uses the Qwen2 tokenizer with the special <image> token. For reproducible benchmark evaluation,
users should follow the prompts and evaluation procedure described in the paper rather than relying on the generation defaults.
Training Details
Training Data
SciGram-7B was trained using the three SciGram subsets:
- SciGram-Align: 582,213 image-caption instruction pairs used for visual-language alignment.
- SciGram-VIT: 737,887 diagram-based multiple-choice instruction examples.
- SciGram-M³: 47,506 examples curated from TQA, ScienceQA, OpenBookQA, and ARC-Easy/Challenge.
Together, these subsets contain approximately 1.37 million training instructions.
The complete SciGram dataset contains 194,071 scientific diagrams and more than 1.4 million synthetic visual instructions. The diagrams cover life sciences, earth sciences, and physical sciences.
The dataset construction pipeline extracts scientific terminology from middle-school science curricula, generates atomic scientific facts, retrieves candidate diagrams from the web, filters and deduplicates the retrieved images, and generates captions and multiple-choice questions using a vision-language model.
The instruction-generation process used Qwen2-VL-7B to generate captions and diagram-grounded multiple-choice questions.
Training Procedure
SciGram-7B was trained using a LLaVA-style three-stage training pipeline:
- Alignment: The multimodal projection/alignment components are trained using SciGram-Align while the relevant pretrained components are kept frozen.
- Visual Instruction Tuning: LoRA fine-tuning is performed on SciGram-VIT.
- Further Fine-tuning: The resulting model is further fine-tuned with LoRA on SciGram-M³.
The model was trained on two NVIDIA A100 GPUs for approximately 450 GPU-hours.
Preprocessing
The model follows the LLaVA image-processing and multimodal training pipeline.
For the visual instruction tuning and further fine-tuning stages, the training configuration uses an any-resolution image processing strategy and the following image grid pinpoints:
(384, 768)
(768, 384)
(768, 768)
(1152, 384)
(384, 1152)
The multimodal projector uses an MLP with two GELU layers (mlp2x_gelu), and the selected vision feature layer is -2.
Training Hyperparameters
Alignment stage — SciGram-Align
- Epochs: 1
- Batch size: 4
- Learning rate: 1e-3
- Weight decay: 0
- Multimodal projector learning rate: 2e-5
- Maximum sequence length: 32,768
Visual Instruction Tuning — SciGram-VIT
- Epochs: 1
- LoRA rank: 128
- LoRA alpha: 256
- Batch size: 1
- Gradient accumulation steps: 16
- Learning rate: 1e-5
- Weight decay: 0
- Warmup ratio: 0.03
- Scheduler: cosine
- Maximum sequence length: 32,768
- Multimodal projector learning rate: 2e-5
Further Fine-tuning — SciGram-M³
- Epochs: 3
- LoRA rank: 128
- LoRA alpha: 256
- Batch size: 1
- Gradient accumulation steps: 16
- Learning rate: 1e-5
- Weight decay: 0
- Warmup ratio: 0.03
- Scheduler: cosine
- Maximum sequence length: 32,768
- Multimodal projector learning rate: 2e-5
Speeds, Sizes, Times
- Training hardware: 2 × NVIDIA A100 GPUs
- Training compute: approximately 450 GPU-hours
- Parameter count: approximately 8B
- Maximum sequence length: 32,768 tokens
Evaluation
Testing Data, Factors & Metrics
Testing Data
SciGram-7B was evaluated on three scientific multimodal question-answering benchmarks:
- TQA (Textbook Question Answering): includes text-only, true/false, and diagram-grounded questions across physical, life, and earth sciences.
- ScienceQA (SQA): contains multimodal science questions covering elementary and high-school science curricula.
- AI2D: contains grade-school science diagrams paired with multiple-choice questions, including evaluations with opaque and transparent diagram labels.
Metrics
The primary evaluation metric is accuracy (%), computed as the proportion of correctly answered multiple-choice questions.
Results
The reported results for LLaVA-SciGram 7B are:
| Benchmark | Subset | Accuracy |
|---|---|---|
| TQA | Text MC | 90.87 |
| TQA | True/False | 92.00 |
| TQA | Diagram MC | 76.68 |
| TQA | Overall | 83.66 |
| ScienceQA | Natural Sciences | 96.27 |
| ScienceQA | Social Sciences | 97.53 |
| ScienceQA | Language | 91.64 |
| ScienceQA | Text | 99.11 |
| ScienceQA | Visual/Image Support | 95.24 |
| ScienceQA | No Support | 93.38 |
| ScienceQA | Grade 1–6 | 95.49 |
| ScienceQA | Grade 7–12 | 95.06 |
| ScienceQA | Overall | 95.33 |
| AI2D | Opaque Labels | 80.21 |
| AI2D | Transparent Labels | 89.93 |
| AI2D | Overall | 85.07 |
Technical Specifications
Model Architecture and Objective
SciGram-7B uses a LLaVA-Qwen multimodal architecture.
The model consists of:
- Architecture:
LlavaQwenForCausalLM - Language backbone:
Qwen/Qwen2-7B-Instruct - Vision encoder:
openai/clip-vit-large-patch14-336 - Multimodal projector: 2-layer MLP with GELU activation
- Image processing: Any-resolution (
anyres) - Maximum sequence length: 32,768
- Data type: FP16
- Attention implementation: SDPA
- Model objective: Autoregressive multimodal language modeling
Compute Infrastructure
Training was performed using two NVIDIA A100 GPUs and required approximately 450 GPU-hours.
Hardware
- 2 × NVIDIA A100 GPUs
Software
The released checkpoint is compatible with the LLaVA-Qwen architecture.
The associated environment includes:
- srsly~=2.5.1
- nltk~=3.9.1
- flair~=0.15.1
- transformers~=4.56.2
- tqdm~=4.67.1
- numpy~=2.3.3
- torch~=2.8.0
- textdistance~=4.6.3
- datasets~=4.1.1
- pillow~=11.3.0
Bias, Risks, and Limitations
Licensing and Data Usage.
All datasets and pretrained models are subject to their respective licenses, and future users are responsible for complying with their terms. Improper use of copyrighted datasets or proprietary models may result in legal or ethical violations. As noted in our GitHub repository on the license and copyright of content linked from SciGram:
- Images linked from SciGram are copyrighted by their respective owners; the SciGram authors do not host or redistribute them.
- Image URLs are publicly available on the internet and were not scraped from private sources.
- We respected robots.txt rules and site Terms of Service (ToS) during URL collection.
- SciGram is intended for educational and research purposes only; its creators do not claim ownership of linked content.
- SciGram released under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0).
Bias and Fairness.
Pretrained models may reflect biases in their training data. Although our study focuses on diagram reasoning, such biases may affect downstream outputs, potentially disadvantaging certain groups or misrepresenting information. Users should consider these risks when deploying similar models. Environmental Impact. Training and fine-tuning large models are computationally expensive and contribute to carbon emissions. We encourage efficient training strategies and consideration of environmental costs when developing similar systems. Misuse Potential. Although intended for research and educational purposes, our approach could be misused for automated content generation or misinformation. Appropriate safeguards and ethical guidelines should be followed to minimize potential harm.
Citation
If you use SciGram-7B or the SciGram dataset, please cite the accompanying paper:
BibTeX:
@inproceedings{
ortega2026from,
title={From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding},
author={Raúl Ortega and José Maunel Gómez-Pérez,
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=Xdi2q5XKgB}
}
APA:
Ortega, R., & Gómez-Pérez, J. M. (2026). From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding. Proceedings of the Conference on Language Modeling (COLM 2026).
Model Card Contact
For questions concerning the model or dataset, please contact rortega@expert.ai or jmgomez@expert.ai.
- Downloads last month
- 38