Instructions to use google/gemma-3-27b-it with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use google/gemma-3-27b-it with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="google/gemma-3-27b-it")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
pipe(text=messages)

# Load model directly
from transformers import AutoProcessor, AutoModelForImageTextToText

processor = AutoProcessor.from_pretrained("google/gemma-3-27b-it")
model = AutoModelForImageTextToText.from_pretrained("google/gemma-3-27b-it")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
inputs = processor.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:]))

Inference
HuggingChat
Notebooks
Google Colab
Kaggle
Local Apps

vLLM

How to use google/gemma-3-27b-it with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "google/gemma-3-27b-it"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "google/gemma-3-27b-it",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'

Use Docker

docker model run hf.co/google/gemma-3-27b-it

SGLang

How to use google/gemma-3-27b-it with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "google/gemma-3-27b-it" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "google/gemma-3-27b-it",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "google/gemma-3-27b-it" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "google/gemma-3-27b-it",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'

Docker Model Runner
How to use google/gemma-3-27b-it with Docker Model Runner:
```
docker model run hf.co/google/gemma-3-27b-it
```

Why MLP so tiny but vision part of the model works quite well

#52

by CCRss - opened Apr 10, 2025

Discussion

CCRss

Apr 10, 2025

•

edited Apr 10, 2025

Why in gemma model they have such a tiny MLP between VisionTower and LLM. There is only 1 .matmul that is trainable how is that works, when you train only LLM and MLP?
Just curious maybe someone know the answer

And in paper they said they only trained LLM without touching siglip so is that mean such projection layer is enough to transfer vision features to llm?

projected_vision_outputs = torch.matmul(normed_vision_outputs, self.mm_input_projection_weight)

For 27B it's about 6 million parameters that will be trainable.

Or I understand it wrong and there is something that is also trainable except this one?

CCRss changed discussion title from Why MLP so slow but vision part of the model works quite well to Why MLP so tiny but vision part of the model works quite well Apr 10, 2025

GopiUppari

Google org Apr 15, 2025

Hi @CCRss ,

The Gemma model connects vision and language using a lightweight projection layer. This layer acts as a translator, converting visual features into a form that the language model can understand. Despite its small size around 6 million parameters in the 27B version. It works effectively because the vision encoder is already highly capable and doesn’t require further training. Its strong visual representations make this minimal projection sufficient to align visual inputs with text processing.

Please take a look at this blog. It provides detailed insights into how Vision-Language Models work.

Thank you.

CCRss

Apr 15, 2025

•

edited Apr 15, 2025

@GopiUppari
Thank you so much for the detailed answer. 🍀. I will certainly check the blogpost to deeper my understanding of how it's working.

May I ask a question about Gemma vision fine-tuning. How to do it properly, when MLP is small we need to fine-tune our LLM weights as well, when we train on vision tasks.
Before I was working usually with MLP only fine-tuning with larger sizes like 70-200 millions and during MLP training model was able to properly understand how to connect vision and LLM and improve results on task specific cases.

But for Gemma we will tune LLM so I wonder how to do it properly, such that it will not degrade the text generation performance.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment