Apollo-VL-Edge-3B — GGUF Quantizations

Intelligence isn't about scale. It's about precision.

Apollo-VL-Edge-3B is an efficient Vision-Language Model (VLM) developed by Pluto AI Labs, designed for practical multimodal inference on consumer hardware.

This repository provides the official GGUF quantizations of Apollo-VL-Edge-3B for local inference with llama.cpp and compatible runtimes.


Overview

Apollo-VL-Edge-3B is built upon the foundation of Qwen2.5-VL-3B-Instruct, bringing multimodal understanding, visual reasoning, document understanding, and OCR capabilities to hardware with significantly lower memory requirements.

The GGUF release is optimized for:

  • Local and offline inference
  • CPU inference
  • Low-VRAM GPUs
  • Consumer laptops and desktops
  • Apple Silicon devices
  • Edge AI deployments
  • llama.cpp
  • LM Studio
  • Ollama-compatible workflows

For the native FP16 Transformers checkpoint:

Apollo-VL-Edge-3B — Hugging Face


Model Highlights

  • 3B parameter-class VLM
  • Multimodal image + text understanding
  • Visual question answering
  • Chart understanding
  • Science diagram reasoning
  • Document understanding
  • OCR and text extraction
  • Visual reasoning
  • Local inference
  • GGUF quantization support
  • Designed for edge and consumer hardware

Note: Apollo-VL-Edge-3B GGUF models require the corresponding multimodal projector (mmproj) file for image understanding.


Provided Quantizations

File Format Approx. Size Recommended VRAM Target Hardware Precision Loss
Apollo-VL-Edge-3B-F16.gguf FP16 ~6.0 GB 6 GB RTX 3060/4060, Apple M1/M2/M3 Baseline
Apollo-VL-Edge-3B-Q8_0.gguf Q8_0 ~3.3 GB 4 GB RTX 3050, Apple M-Series < 0.3%
Apollo-VL-Edge-3B-Q6_K.gguf Q6_K ~2.6 GB 3.5 GB Consumer GPUs < 0.8%
Apollo-VL-Edge-3B-Q5_K_M.gguf Q5_K_M ~2.2 GB 3.0 GB Laptops, high-RAM mobile < 1.5%
Apollo-VL-Edge-3B-Q4_K_M.gguf Q4_K_M ~1.9 GB 2.5 GB Edge devices, CPU-only < 2.1%

Multimodal Projector

To use the vision capabilities, you must load:

mmproj-Apollo-VL-Edge-3B-f16.gguf

alongside your selected GGUF model.

The mmproj file handles the connection between the visual encoder and language model.


Benchmark Performance

Apollo-VL-Edge-3B was evaluated across visual reasoning, chart understanding, and document OCR benchmarks.

Model Organization Params AI2D ChartQA OCRBench
Apollo-VL-Edge-3B Pluto AI Labs 3.0B 77.98% 78.60% 786
Qwen2.5-VL-3B-Instruct Alibaba Qwen 3.0B 78.00% 78.50% 785
InternVL2-4B OpenGVLab 4.2B 76.20% 78.40% 768
Phi-3.5-Vision-Instruct Microsoft 4.2B 75.40% 76.20% 695

Evaluation Results

Benchmark Score
AI2D 77.98%
ChartQA 78.60%
OCRBench 786

Quantization note: Lower-bit quantizations such as Q4_K_M may introduce some degradation on complex visual reasoning tasks. For maximum fidelity, use F16 or Q8_0.


Hardware Recommendations

Quantization Recommended Hardware
F16 GPU with 6+ GB VRAM / Apple Silicon
Q8_0 GPU with 4+ GB VRAM
Q6_K GPU with 3.5+ GB VRAM
Q5_K_M GPU with 3+ GB VRAM
Q4_K_M 2.5+ GB VRAM / CPU inference

Actual memory requirements may vary depending on context length, runtime configuration, GPU offloading, and system RAM.


Usage

llama.cpp

Apollo-VL-Edge-3B can be run locally using llama.cpp with both the main GGUF model and multimodal projector.

CLI Inference

./llama-cli \
  -m Apollo-VL-Edge-3B-Q4_K_M.gguf \
  --mmproj mmproj-Apollo-VL-Edge-3B-f16.gguf \
  -p "Describe this image in detail." \
  --image "path/to/your/image.png"

Replace the model filename with your preferred quantization.


llama-server

Apollo-VL-Edge-3B can also be served through the OpenAI-compatible llama-server API.

./llama-server \
  -m Apollo-VL-Edge-3B-Q4_K_M.gguf \
  --mmproj mmproj-Apollo-VL-Edge-3B-f16.gguf \
  --host 0.0.0.0 \
  --port 8080

The server can then be accessed through compatible OpenAI API clients.


LM Studio

Apollo-VL-Edge-3B can be used locally through LM Studio.

Steps

  1. Open LM Studio.
  2. Search for Apollo-VL-Edge-3B-GGUF.
  3. Download your preferred quantization.
  4. Download the corresponding mmproj file.
  5. Load the GGUF model.
  6. Attach an image.
  7. Start the conversation.

For the best balance between performance and quality, Q4_K_M or Q5_K_M is recommended for typical consumer hardware.


Ollama

If your Ollama workflow supports the required multimodal GGUF configuration, Apollo-VL-Edge-3B can be integrated into a local Ollama deployment.

The recommended configuration is to pair:

Apollo-VL-Edge-3B-Q4_K_M.gguf

with:

mmproj-Apollo-VL-Edge-3B-f16.gguf

Refer to the runtime-specific documentation for the exact multimodal configuration required by your Ollama version.


Multimodal Projector

The multimodal projector is required for visual inference.

mmproj-Apollo-VL-Edge-3B-f16.gguf

Required Files

A typical deployment should contain:

Apollo-VL-Edge-3B-GGUF/
├── Apollo-VL-Edge-3B-Q4_K_M.gguf
├── mmproj-Apollo-VL-Edge-3B-f16.gguf
└── README.md

You can replace Q4_K_M with another available quantization.


Example Use Cases

Apollo-VL-Edge-3B is designed for practical multimodal workloads such as:

  • Image understanding
  • Visual question answering
  • Chart interpretation
  • Diagram reasoning
  • Document understanding
  • OCR
  • Screenshot analysis
  • Visual classification
  • Local AI assistants
  • Edge AI applications
  • Offline multimodal workflows
  • Low-VRAM VLM deployment

Performance vs. Size

The different quantizations provide a range of quality, memory usage, and deployment flexibility.

F16

Best for:

  • Maximum model fidelity
  • High-memory GPUs
  • Benchmarking
  • Research

Q8_0

Best for:

  • Near-FP16 quality
  • Consumer GPUs
  • High-quality local inference

Q6_K

Best for:

  • Strong quality-to-memory ratio
  • General-purpose local VLM inference

Q5_K_M

Best for:

  • Balanced performance
  • Laptops
  • Consumer hardware

Q4_K_M

Best for:

  • Low-VRAM GPUs
  • CPU inference
  • Edge deployments
  • Maximum memory efficiency

Model Architecture

Apollo-VL-Edge-3B is based on the Qwen2.5-VL-3B-Instruct architecture and is designed for multimodal visual-language workloads.

The system combines:

Image
  ↓
Vision Encoder
  ↓
Multimodal Projector
  ↓
Language Model
  ↓
Text Response

The GGUF distribution separates the language-model weights from the multimodal projector so that compatible inference engines can efficiently load and execute the complete VLM.


Limitations

Apollo-VL-Edge-3B is intended for research, experimentation, and local AI applications.

Potential limitations include:

  • Visual hallucinations
  • OCR errors on low-quality images
  • Reduced accuracy on highly complex diagrams
  • Sensitivity to image resolution
  • Reduced reasoning performance under aggressive quantization
  • Potential inaccuracies in numerical chart interpretation
  • Performance differences across inference runtimes

Model outputs should be independently verified for high-stakes applications.


Recommended Quantization

For most users:

Q4_K_M — Best overall edge deployment choice

For higher-quality inference:

Q5_K_M / Q6_K

For maximum fidelity:

Q8_0 / F16


Acknowledgements

Apollo-VL-Edge-3B builds upon the foundational work of the Qwen Team at Alibaba and the open-source multimodal ecosystem.

We would like to acknowledge:

  • Qwen Team / Alibaba for Qwen2.5-VL
  • llama.cpp contributors for the GGUF inference ecosystem
  • Hugging Face for model hosting and open-source ML infrastructure
  • The broader open-source VLM and multimodal research community

About Pluto AI Labs

Pluto AI Labs is an independent open-source AI research lab focused on efficient intelligence, multimodal reasoning, model distillation, and edge deployment.

We explore how capable AI systems can be made smaller, faster, and more accessible without requiring hyperscale infrastructure.

Links


Citation

If you use Apollo-VL-Edge-3B in your research or projects, please cite:

@misc{apollo_vl_edge_3b,
  title        = {Apollo-VL-Edge-3B: Elite Visual Reasoning on Edge Hardware},
  author       = {Siddharth N.R. and Pluto AI Labs},
  year         = {2026},
  howpublished = {Hugging Face},
  url          = {https://huggingface.co/Pluto-AI-Labs/Apollo-VL-Edge-3B}
}

Apollo-VL-Edge-3B — GGUF

Precision over scale. Reasoning at the edge.

Built by Pluto AI Labs

Downloads last month
2
GGUF
Model size
3B params
Architecture
qwen2vl
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Pluto-AI-Labs/Apollo-VL-Edge-3B-GGUF

Quantized
(3)
this model

Dataset used to train Pluto-AI-Labs/Apollo-VL-Edge-3B-GGUF

Collection including Pluto-AI-Labs/Apollo-VL-Edge-3B-GGUF

Evaluation results