How to use from
Docker Model Runner
docker model run hf.co/Kimokcheon/Fundus-R1-7B
Quick Links

Fundus-R1

Fundus-R1 is a fundus-reading multimodal large language model introduced in the paper:

Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data
Yuchuan Deng, Qijie Wei, Kaiheng Qian, Jiazhen Liu, Zijie Xin, Bangxiang Lan, Jingyu Liu, Jianfeng Dong, Xirong Li
Paper: https://arxiv.org/abs/2604.08322

Fundus-R1 is designed for fundus image understanding, including color fundus photography (CFP), optical coherence tomography (OCT), and ultra-widefield fundus imaging (UWF). The model is trained using publicly available data and aims to improve knowledge-aware reasoning for retinal image analysis.

Model Variants

Model Repository
Fundus-R1-3B Kimokcheon/Fundus-R1-3B
Fundus-R1-7B Kimokcheon/Fundus-R1-7B

The examples below use Fundus-R1-7B, based on Qwen2.5-VL-7B-Instruct. The 3B checkpoint is listed above as a related model.

Method Overview

Fundus-R1 addresses the difficulty of training fundus-reading MLLMs without private clinical-report data. According to the paper, the model is trained exclusively on public datasets, where most samples contain only image-level labels rather than detailed diagnostic reports.

The training pipeline contains two key components:

  1. Knowledge-aware reasoning trace construction. A retrieval-augmented generation (RAG) procedure is used to compose image-specific reasoning traces that connect visual findings to image labels through ophthalmic knowledge.
  2. Reasoning-enhanced RLVR. Reinforcement learning with verifiable rewards (RLVR) is enhanced with a process reward that encourages self-consistency in the generated reasoning trace.

The paper reports evaluation on three fundus-reading benchmarks: FunBench, Omni-Fundus, and GMAI-Fundus.

Intended Use

Fundus-R1 is intended for research on fundus-image understanding, medical multimodal reasoning, ophthalmic MLLMs, and public-data-based post-training of medical vision-language models.

Possible research uses include:

  • fundus image question answering;
  • retinal disease recognition experiments;
  • reasoning-trace analysis for medical MLLMs;
  • comparison with general-purpose MLLMs and ophthalmology-specific MLLMs;
  • studies on RAG-generated medical reasoning traces and RLVR training.

Important Medical Disclaimer

This model is released for research use. It is not a certified medical device and should not be used as the sole basis for clinical diagnosis, treatment planning, triage, or patient management. Outputs should be reviewed by qualified medical professionals before any clinical interpretation or downstream use.

Example Usage

The exact loading code may depend on the checkpoint configuration and your installed transformers version. A typical Qwen2.5-VL-style loading pattern is:

import torch
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
from qwen_vl_utils import process_vision_info

model_id = "Kimokcheon/Fundus-R1-7B"

model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
processor = AutoProcessor.from_pretrained(model_id)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "path/to/fundus_image.jpg"},
            {"type": "text", "text": "Describe the retinal findings in this fundus image."},
        ],
    }
]

text = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
    text=[text],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
).to(model.device)

with torch.no_grad():
    generated_ids = model.generate(**inputs, max_new_tokens=512)

output_ids = generated_ids[:, inputs.input_ids.shape[1]:]
response = processor.batch_decode(
    output_ids,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False,
)[0]
print(response)

Install common dependencies:

pip install -U transformers accelerate safetensors qwen-vl-utils

vLLM Usage

Fundus-R1-7B uses the Qwen2_5_VLForConditionalGeneration architecture and can be served directly from its Hugging Face checkpoint with vLLM. No GGUF conversion is needed for this route.

The following example targets a Linux machine with an NVIDIA GPU, a compatible driver, and sufficient GPU memory. Use a fresh Python environment to avoid conflicts with existing PyTorch installations.

Start the server

python3 -m venv .venv-vllm
source .venv-vllm/bin/activate
python -m pip install --upgrade pip
python -m pip install --upgrade vllm openai

CUDA_VISIBLE_DEVICES=0 vllm serve Kimokcheon/Fundus-R1-7B \
  --served-model-name fundus-r1-7b \
  --host 127.0.0.1 \
  --port 8000 \
  --dtype bfloat16 \
  --max-model-len 4096 \
  --max-num-seqs 1 \
  --gpu-memory-utilization 0.90 \
  --limit-mm-per-prompt '{"image":1,"video":0}' \
  --mm-processor-kwargs '{"min_pixels":200704,"max_pixels":1003520}' \
  --enforce-eager

Keep this terminal running. Once the server is ready, check it from another terminal:

curl http://127.0.0.1:8000/v1/models

This configuration allows one image per request, limits concurrency to one sequence, and disables CUDA graphs to reduce startup memory pressure.

Query a local image

In a second terminal, activate the same environment and replace ./fundus_image.jpg with your image path. The image is sent to the local vLLM server as a base64 data URL; no publicly hosted image is required.

source .venv-vllm/bin/activate
python - ./fundus_image.jpg <<'PYTHON'
import base64
import mimetypes
import sys
from pathlib import Path
from openai import OpenAI

image_path = Path(sys.argv[1]).expanduser()
if not image_path.is_file():
    raise FileNotFoundError(f"Image not found: {image_path}")
mime_type = mimetypes.guess_type(image_path.name)[0]
if mime_type not in {"image/jpeg", "image/png"}:
    raise ValueError("Use a JPEG or PNG image for this example.")

image_b64 = base64.b64encode(image_path.read_bytes()).decode("ascii")
client = OpenAI(
    base_url="http://127.0.0.1:8000/v1",
    api_key="EMPTY",  # Local vLLM server; not an OpenAI API key.
    timeout=180.0,
)
response = client.chat.completions.create(
    model="fundus-r1-7b",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "image_url", "image_url": {
                    "url": f"data:{mime_type};base64,{image_b64}"
                }},
                {"type": "text", "text":
                    "Describe the retinal findings in this fundus image."},
            ],
        },
    ],
    temperature=0.0,
    max_tokens=1024,
)
print(response.choices[0].message.content)
PYTHON

See the vLLM supported-model list and multimodal input guide for additional deployment options.

Ollama Usage

Use the Fundus-R1 fine-tuned weights, not the standard qwen2.5vl:7b model from the Ollama library. The latter is the base model, not Fundus-R1.

The example below uses a local GGUF conversion and targets recent Ollama releases with support for importing a language-model GGUF together with its multimodal projector. Install or update Ollama first. The commands below use a Bash-compatible shell; keep the Ollama app or service running. If no service is running, launch ollama serve in a separate terminal.

Download and convert the checkpoint

Use a separate environment for conversion so its dependencies do not interfere with vLLM.

python3 -m venv .venv-ollama
source .venv-ollama/bin/activate
python -m pip install --upgrade pip
python -m pip install --upgrade huggingface_hub ollama

hf download Kimokcheon/Fundus-R1-7B --local-dir ./Fundus-R1-7B

git clone --depth 1 https://github.com/ggml-org/llama.cpp.git
python -m pip install -r llama.cpp/requirements/requirements-convert_hf_to_gguf.txt
mkdir -p ./Fundus-R1-7B-GGUF

# Export the fine-tuned language-model weights with Q8_0 quantization.
python llama.cpp/convert_hf_to_gguf.py ./Fundus-R1-7B \
  --outtype q8_0 \
  --outfile ./Fundus-R1-7B-GGUF/Fundus-R1-7B-Q8_0.gguf

# Export the matching vision encoder/projector from the SAME checkpoint.
python llama.cpp/convert_hf_to_gguf.py ./Fundus-R1-7B \
  --mmproj \
  --outtype f16 \
  --outfile ./Fundus-R1-7B-GGUF/mmproj-Fundus-R1-7B-f16.gguf

Both files are required for image understanding. Do not substitute a projector from another checkpoint: the fine-tuned visual weights must be preserved. Keep the original Hugging Face checkpoint for full-precision inference and evaluation. Q8_0 conversion changes the weight precision and can change model outputs.

Create the Ollama model

Run the following from the directory containing Fundus-R1-7B-GGUF:

cat > Modelfile <<'MODELFILE'
FROM ./Fundus-R1-7B-GGUF/Fundus-R1-7B-Q8_0.gguf
FROM ./Fundus-R1-7B-GGUF/mmproj-Fundus-R1-7B-f16.gguf

SYSTEM """You are a helpful assistant."""
TEMPLATE """{{- range .Messages }}
{{- printf "<|im_start|>%s\n%s<|im_end|>\n" .Role .Content }}
{{- end }}
{{- printf "<|im_start|>assistant\n" }}"""

PARAMETER num_ctx 4096
PARAMETER num_predict 1024
PARAMETER temperature 0
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
MODELFILE

ollama create fundus-r1-7b -f ./Modelfile
ollama show fundus-r1-7b

Query a local image

source .venv-ollama/bin/activate
python - ./fundus_image.jpg <<'PYTHON'
import sys
from pathlib import Path
from ollama import chat

image_path = Path(sys.argv[1]).expanduser().resolve()
if not image_path.is_file():
    raise FileNotFoundError(f"Image not found: {image_path}")

response = chat(
    model="fundus-r1-7b",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": "Describe the retinal findings in this fundus image.",
            "images": [str(image_path)],
        },
    ],
    options={"temperature": 0.0, "num_predict": 1024},
)
print(response.message.content)
PYTHON

Alternatively, provide an image path directly in the Ollama CLI:

ollama run fundus-r1-7b ./fundus_image.jpg "Describe the retinal findings in this fundus image."

See the Ollama Modelfile reference, Ollama vision guide, and llama.cpp multimodal conversion guide.

Download Through HF Mirror

For users in regions where the official Hugging Face endpoint is slow, the checkpoints can be downloaded through the Hugging Face mirror endpoint. Install the CLI with python -m pip install --upgrade huggingface_hub if it is not already available:

export HF_ENDPOINT=https://hf-mirror.com

hf download Kimokcheon/Fundus-R1-3B --local-dir ./Fundus-R1-3B
hf download Kimokcheon/Fundus-R1-7B --local-dir ./Fundus-R1-7B

To verify mirror availability with a lightweight file:

export HF_ENDPOINT=https://hf-mirror.com

hf download Kimokcheon/Fundus-R1-3B config.json --local-dir /tmp/fundus-r1-3b-check
hf download Kimokcheon/Fundus-R1-7B config.json --local-dir /tmp/fundus-r1-7b-check

Citation

If you use Fundus-R1, please cite the paper:

@article{deng2026fundusr1,
  title={Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data},
  author={Deng, Yuchuan and Wei, Qijie and Qian, Kaiheng and Liu, Jiazhen and Xin, Zijie and Lan, Bangxiang and Liu, Jingyu and Dong, Jianfeng and Li, Xirong},
  journal={arXiv preprint arXiv:2604.08322},
  year={2026}
}

Links

Downloads last month
53
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kimokcheon/Fundus-R1-7B

Finetuned
(1250)
this model
Quantizations
1 model

Paper for Kimokcheon/Fundus-R1-7B