Safetensors
llava_qwen

SciGram-7B

SciGram-7B is a 7B-class vision-language model specialized in understanding scientific diagrams. It is based on a LLaVA-style architecture with a pretrained CLIP vision encoder and a Qwen2-Instruct 7B language model, and is trained on SciGram, a large-scale dataset of scientific diagrams and synthetic visual instructions covering life, earth, and physical sciences.

The model is designed for multimodal understanding tasks involving scientific diagrams, including diagram question answering, visual grounding, diagram description, and science-related visual reasoning.

SciGram-7B corresponds to the LLaVA-SciGram 7B model described in:

Raul Ortega and José Manuel Gómez-Pérez. From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding. Proceedings of the Conference on Language Modeling (COLM 2026).

Model Details

Model Description

SciGram-7B is a vision-language model trained to improve scientific diagram understanding. It uses a pretrained CLIP vision encoder together with a Qwen2-Instruct 7B language model. The model is trained using three SciGram training stages: SciGram-Align, SciGram-VIT, and SciGram-M³.

The training data contains scientific diagrams paired with synthetic captions and multiple-choice visual instructions covering life sciences, earth sciences, and physical sciences.

The model is intended primarily for research on scientific diagram understanding and multimodal scientific question answering.

  • Developed by: Expert.ai Language Technology Research Laboratory
  • Authors: Raul Ortega and José Manuel Gómez-Pérez
  • Shared by: Expert.ai Language Technology Research Laboratory
  • Model type: Vision-language model (VLM)
  • Architecture: LLaVA-Qwen
  • Language(s): English
  • License: CC BY 4.0
  • Vision encoder: Pretrained CLIP vision encoder
  • Language model: Qwen2-Instruct 7B
  • Model format: Safetensors, FP16
  • Maximum context length: 32,768 tokens

The released checkpoint contains approximately 8B parameters and is approximately 16.3 GB in size.

Model Sources

Github repo   Dataset on Hugging Face COLM Paper

How to Get Started with the Model

The model is released using the LLaVA-Qwen architecture. It can be loaded using a compatible LLaVA implementation and the Hugging Face checkpoint.

A typical workflow is:

    from transformers import AutoProcessor, AutoModel
    from PIL import Image
    import torch

    model_id = "expertailab/scigram-7b"

    processor = AutoProcessor.from_pretrained(model_id)
    model = AutoModel.from_pretrained(
        model_id,
        torch_dtype=torch.float16,
        device_map="auto",
    )

    image = Image.open("scientific_diagram.png").convert("RGB")

    prompt = "What does this scientific diagram show? Explain the main components and their relationships."

    inputs = processor(
        images=image,
        text=prompt,
        return_tensors="pt"
    ).to(model.device)

    with torch.no_grad():
        output = model.generate(
            **inputs,
            max_new_tokens=512,
        )

    answer = processor.batch_decode(
        output,
        skip_special_tokens=True
    )[0]

    print(answer)

The model repository contains a LlavaProcessor configuration and uses the Qwen2 tokenizer with the special <image> token. For reproducible benchmark evaluation, users should follow the prompts and evaluation procedure described in the paper rather than relying on the generation defaults.

Training Details

Training Data

SciGram-7B was trained using the three SciGram subsets:

  1. SciGram-Align: 582,213 image-caption instruction pairs used for visual-language alignment.
  2. SciGram-VIT: 737,887 diagram-based multiple-choice instruction examples.
  3. SciGram-M³: 47,506 examples curated from TQA, ScienceQA, OpenBookQA, and ARC-Easy/Challenge.

Together, these subsets contain approximately 1.37 million training instructions.

The complete SciGram dataset contains 194,071 scientific diagrams and more than 1.4 million synthetic visual instructions. The diagrams cover life sciences, earth sciences, and physical sciences.

The dataset construction pipeline extracts scientific terminology from middle-school science curricula, generates atomic scientific facts, retrieves candidate diagrams from the web, filters and deduplicates the retrieved images, and generates captions and multiple-choice questions using a vision-language model.

The instruction-generation process used Qwen2-VL-7B to generate captions and diagram-grounded multiple-choice questions.

Training Procedure

SciGram-7B was trained using a LLaVA-style three-stage training pipeline:

  1. Alignment: The multimodal projection/alignment components are trained using SciGram-Align while the relevant pretrained components are kept frozen.
  2. Visual Instruction Tuning: LoRA fine-tuning is performed on SciGram-VIT.
  3. Further Fine-tuning: The resulting model is further fine-tuned with LoRA on SciGram-M³.

The model was trained on two NVIDIA A100 GPUs for approximately 450 GPU-hours.

Preprocessing

The model follows the LLaVA image-processing and multimodal training pipeline.

For the visual instruction tuning and further fine-tuning stages, the training configuration uses an any-resolution image processing strategy and the following image grid pinpoints:

    (384, 768)
    (768, 384)
    (768, 768)
    (1152, 384)
    (384, 1152)

The multimodal projector uses an MLP with two GELU layers (mlp2x_gelu), and the selected vision feature layer is -2.

Training Hyperparameters

Alignment stage — SciGram-Align

  • Epochs: 1
  • Batch size: 4
  • Learning rate: 1e-3
  • Weight decay: 0
  • Multimodal projector learning rate: 2e-5
  • Maximum sequence length: 32,768

Visual Instruction Tuning — SciGram-VIT

  • Epochs: 1
  • LoRA rank: 128
  • LoRA alpha: 256
  • Batch size: 1
  • Gradient accumulation steps: 16
  • Learning rate: 1e-5
  • Weight decay: 0
  • Warmup ratio: 0.03
  • Scheduler: cosine
  • Maximum sequence length: 32,768
  • Multimodal projector learning rate: 2e-5

Further Fine-tuning — SciGram-M³

  • Epochs: 3
  • LoRA rank: 128
  • LoRA alpha: 256
  • Batch size: 1
  • Gradient accumulation steps: 16
  • Learning rate: 1e-5
  • Weight decay: 0
  • Warmup ratio: 0.03
  • Scheduler: cosine
  • Maximum sequence length: 32,768
  • Multimodal projector learning rate: 2e-5

Speeds, Sizes, Times

  • Training hardware: 2 × NVIDIA A100 GPUs
  • Training compute: approximately 450 GPU-hours
  • Parameter count: approximately 8B
  • Maximum sequence length: 32,768 tokens

Evaluation

Testing Data, Factors & Metrics

Testing Data

SciGram-7B was evaluated on three scientific multimodal question-answering benchmarks:

  • TQA (Textbook Question Answering): includes text-only, true/false, and diagram-grounded questions across physical, life, and earth sciences.
  • ScienceQA (SQA): contains multimodal science questions covering elementary and high-school science curricula.
  • AI2D: contains grade-school science diagrams paired with multiple-choice questions, including evaluations with opaque and transparent diagram labels.

Metrics

The primary evaluation metric is accuracy (%), computed as the proportion of correctly answered multiple-choice questions.

Results

The reported results for LLaVA-SciGram 7B are:

Benchmark Subset Accuracy
TQA Text MC 90.87
TQA True/False 92.00
TQA Diagram MC 76.68
TQA Overall 83.66
ScienceQA Natural Sciences 96.27
ScienceQA Social Sciences 97.53
ScienceQA Language 91.64
ScienceQA Text 99.11
ScienceQA Visual/Image Support 95.24
ScienceQA No Support 93.38
ScienceQA Grade 1–6 95.49
ScienceQA Grade 7–12 95.06
ScienceQA Overall 95.33
AI2D Opaque Labels 80.21
AI2D Transparent Labels 89.93
AI2D Overall 85.07

Technical Specifications

Model Architecture and Objective

SciGram-7B uses a LLaVA-Qwen multimodal architecture.

The model consists of:

  • Architecture: LlavaQwenForCausalLM
  • Language backbone: Qwen/Qwen2-7B-Instruct
  • Vision encoder: openai/clip-vit-large-patch14-336
  • Multimodal projector: 2-layer MLP with GELU activation
  • Image processing: Any-resolution (anyres)
  • Maximum sequence length: 32,768
  • Data type: FP16
  • Attention implementation: SDPA
  • Model objective: Autoregressive multimodal language modeling

Compute Infrastructure

Training was performed using two NVIDIA A100 GPUs and required approximately 450 GPU-hours.

Hardware

  • 2 × NVIDIA A100 GPUs

Software

The released checkpoint is compatible with the LLaVA-Qwen architecture.

The associated environment includes:

  • srsly~=2.5.1
  • nltk~=3.9.1
  • flair~=0.15.1
  • transformers~=4.56.2
  • tqdm~=4.67.1
  • numpy~=2.3.3
  • torch~=2.8.0
  • textdistance~=4.6.3
  • datasets~=4.1.1
  • pillow~=11.3.0

Bias, Risks, and Limitations

Licensing and Data Usage.

All datasets and pretrained models are subject to their respective licenses, and future users are responsible for complying with their terms. Improper use of copyrighted datasets or proprietary models may result in legal or ethical violations. As noted in our GitHub repository on the license and copyright of content linked from SciGram:

  • Images linked from SciGram are copyrighted by their respective owners; the SciGram authors do not host or redistribute them.
  • Image URLs are publicly available on the internet and were not scraped from private sources.
  • We respected robots.txt rules and site Terms of Service (ToS) during URL collection.
  • SciGram is intended for educational and research purposes only; its creators do not claim ownership of linked content.
  • SciGram released under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0).

Bias and Fairness.

Pretrained models may reflect biases in their training data. Although our study focuses on diagram reasoning, such biases may affect downstream outputs, potentially disadvantaging certain groups or misrepresenting information. Users should consider these risks when deploying similar models. Environmental Impact. Training and fine-tuning large models are computationally expensive and contribute to carbon emissions. We encourage efficient training strategies and consideration of environmental costs when developing similar systems. Misuse Potential. Although intended for research and educational purposes, our approach could be misused for automated content generation or misinformation. Appropriate safeguards and ethical guidelines should be followed to minimize potential harm.

Citation

If you use SciGram-7B or the SciGram dataset, please cite the accompanying paper:

BibTeX:

@inproceedings{
  ortega2026from,
  title={From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding},
  author={Raúl Ortega and José Maunel Gómez-Pérez,
  booktitle={Third Conference on Language Modeling},
  year={2026},
  url={https://openreview.net/forum?id=Xdi2q5XKgB}
}

APA:

Ortega, R., & Gómez-Pérez, J. M. (2026). From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding. Proceedings of the Conference on Language Modeling (COLM 2026).

Model Card Contact

For questions concerning the model or dataset, please contact rortega@expert.ai or jmgomez@expert.ai.

Downloads last month
38
Safetensors
Model size
8B params
Tensor type
F32
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train expertailab/scigram-7b

Collection including expertailab/scigram-7b