SnapVL: An Experimental Non-Autoregressive Vision-Language Decision Model

Status: Exploratory Research Checkpoint
Evaluation Stage: Pre-Convergence (Exploratory Training Run)
Paradigm: System-1-Inspired Visual Decision Modeling (Non-Autoregressive)


Important Disclaimers & Evaluation Scope

  1. Pre-Convergence Checkpoint: The metrics reported in this document were evaluated on an intermediate experimental checkpoint where training was paused early to analyze representation dynamics. These numbers do not represent fully converged final weights.
  2. Closed-Schema Evaluation Protocol: Unlike generative vision-language models that decode answers token-by-token from an open vocabulary, SnapVL operates as a closed-schema decision classifier. Questions are framed either as Boolean verification (True/False) or multiple-choice candidate selection (N-choice). Direct numerical comparison with generative LLM decoders should take this protocol difference into account.
  3. Scope: SnapVL is designed strictly for discrete visual decision-making, not for open-ended conversation or free-form text generation.

1. Research Motivation

Standard Vision-Language Models (VLMs) typically couple a vision encoder with an autoregressive Large Language Model (LLM). While this architecture provides conversational flexibility, it introduces non-trivial inference latency due to autoregressive decoding loops, high memory footprints, and potential susceptibility to language-prior affirmation biases.

SnapVL investigates an alternative question:

How effectively can a compact (~450M parameter) model perform structured visual decision-making using a single forward pass without an autoregressive text generator?

This project explores System-1-inspired visual decision modeling—focusing on fast, single-pass perceptual decisions over structured outputs (Boolean truth values and candidate choice selections).


2. Model Architecture

The architecture is streamlined for single forward-pass execution without token decoding loops:

[Image: 224x224] ──► CLIP ViT-L/14 (Top-4 Layers Unfrozen) ──► 256 Patch Tokens (1024-d)
                                                                       │
[Question Text]  ──► CLIP Text Encoder (Frozen)              ──► Text Query (768-d)
                                                                       │
                                                                       ▼
                                                     Cross-Attention Fusion (2 Layers, 512-d)
                                                                       │
                                                                       ▼
                                                     Schema Head (Task Output)
                                                        ├─► Boolean: Logit -> Sigmoid
                                                        └─► Choice: Cosine similarity with candidate choices

Architectural Breakdown:

  • Vision Backbone: OpenAI CLIP ViT-L/14 (24 Transformer layers; top 4 layers unfrozen during stage 3 fine-tuning).
  • Text Backbone: OpenAI CLIP ViT-L/14 Text Encoder (Frozen).
  • Fusion Module: 2-layer Cross-Attention (512 hidden dimension).
  • Schema Head: Shared projection mapping fused multimodal states into task-specific decision outputs (boolean and choice).
  • Parameter Counts:
    • Total Parameters: ~450 Million
    • Trainable Parameters: ~57.2 Million (Top 4 ViT blocks + Cross-Attention Fusion + Schema Head)
  • Measured Batched Throughput: 1.71s (±0.02s) per batch of 256 samples (128 * 2 across 2x NVIDIA T4 GPUs), corresponding to an amortized ~6.68 ms per sample in batched execution (~150 samples/second). Note: This represents amortized multi-GPU batched throughput, not single-sample serial latency.

3. Experimental Results (Pre-Convergence)

The checkpoint was evaluated across multiple benchmark datasets under its structured decision protocol:

A. Benchmark Summary

Dataset Task Type Evaluation Protocol Loss Accuracy Checkpoint Status
GQA Spatial & Relational Reasoning Closed-Schema (Boolean & 4-Choice) 0.3481 82.81% Pre-convergence intermediate checkpoint
VQA-v2 Factual Recognition Closed-Schema (Boolean & 4-Choice) 0.4901 78.77% Pre-convergence intermediate checkpoint
POPE Object Probing (Hallucination) Standard Binary Verification (9,000 queries) — 78.73% Pre-convergence intermediate checkpoint
SNLI-VE Visual Entailment Boolean (Entailment vs Contradiction) 0.6321 67.00% Zero-Shot: Model was never trained on SNLI-VE (17 percentage points above 50% random baseline)

Protocol Distinction: These results are not directly comparable to standard open-ended VQA/GQA leaderboard scores because SnapVL evaluates a transformed closed-schema task rather than unrestricted answer generation.


B. POPE Benchmark Breakdown (9,000 Samples)

The POPE (Polling-based Object Probing Evaluation) benchmark evaluates visual faithfulness by asking whether specific objects exist in an image across three sampling splits:

Split Query Count Accuracy Precision Recall F1-Score Yes-Ratio
Random 3,000 81.90% 85.74% 76.53% 80.87% 44.63%
Popular 3,000 79.70% 81.71% 76.53% 79.04% 46.83%
Adversarial 3,000 74.60% 73.90% 76.07% 74.97% 51.47%
Overall Average 9,000 78.73% 80.15% 76.38% 78.22% 47.64%

Observation: While generative models frequently display an unbalanced Yes-ratio (often 60–70%) due to language co-occurrence priors, SnapVL exhibited a 47.64% Yes-ratio output distribution across the 9,000 queries, showing a relatively balanced positive/negative distribution across splits without exhibiting extreme affirmation skew.


4. Empirical Observations & Limitations

Positive Findings:

  1. Batched Throughput: At ~6.68 ms per sample amortized across a 256-sample batch on 2x T4 GPUs (1.71s per batch), non-autoregressive decision architectures provide efficient throughput for parallel processing pipelines (e.g., frame filtering, offline verification).
  2. Schema-Constrained Outputs: Formatting outputs as structured primitives (boolean or choice) prevents textual parsing failures and unbounded response lengths.
  3. Zero-Shot Transfer on Entailment: The model was not trained on any SNLI-VE data. On unseen SNLI-VE validation data, it achieved **67.00% accuracy (Loss: 0.6321)**—17 percentage points above the 50% random baseline—demonstrating that cross-attention representations trained on VQA/GQA transferred partially to binary visual entailment without task-specific fine-tuning.

Technical Limitations & Trade-offs:

  1. Linguistic Complexity & Text Representation: Because the CLIP text encoder remains frozen, fine-grained compositional and syntactic structures (such as nested sentence negations) face representation limitations inherent to a frozen contrastive encoder.
  2. Protocol Specificity: Performance on GQA (82.81%) and VQA-v2 (78.77%) relies on pre-defined choice candidates and Boolean formats. The model cannot answer unconstrained, open-ended questions outside these schemas.
  3. Lack of Generative Capability: The model cannot explain its reasoning, produce multi-sentence descriptions, or hold conversational dialog.

5. Quick Usage & Verification

import torch
from huggingface_hub import hf_hub_download
from models.vision_encoder import CLIPViTEncoder
from models.question_encoder import CLIPTextEncoder
from models.fusion import CrossAttentionFusion
from models.schema_head import SchemaHead, ArchitectureBModel

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

# 1. Instantiate architecture
vision_enc = CLIPViTEncoder(unfreeze_last_n=4).to(device)
text_enc   = CLIPTextEncoder().to(device)
fusion     = CrossAttentionFusion(vision_dim=1024, text_dim=768, hidden_dim=512).to(device)
head       = SchemaHead(input_dim=512, text_dim=768, hidden_dim=256).to(device)
model      = ArchitectureBModel(vision_enc, text_enc, fusion, head).to(device)

# 2. Load weights
ckpt_path = hf_hub_download(repo_id="akeshkumar/SnapVL", filename="best_model.pth")
checkpoint = torch.load(ckpt_path, map_location="cpu", weights_only=False)
model.load_state_dict(checkpoint.get("model_state_dict", checkpoint), strict=False)
model.eval()

# 3. Model supports structured evaluation:
# - Boolean schema: question + image -> True / False logit
# - Choice schema: question + image + candidate list -> Argmax over candidate embeddings

6. Citation

@misc{snapvl_experiment_2026,
  author = {Kumar, Akesh},
  title = {SnapVL: An Experimental Non-Autoregressive Vision-Language Decision Model},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/akeshkumar/SnapVL}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results