Instructions to use cyboghostginx/gemma-4-31B-it-Adetayo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cyboghostginx/gemma-4-31B-it-Adetayo with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="cyboghostginx/gemma-4-31B-it-Adetayo") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("cyboghostginx/gemma-4-31B-it-Adetayo") model = AutoModelForMultimodalLM.from_pretrained("cyboghostginx/gemma-4-31B-it-Adetayo", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use cyboghostginx/gemma-4-31B-it-Adetayo with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "cyboghostginx/gemma-4-31B-it-Adetayo" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cyboghostginx/gemma-4-31B-it-Adetayo", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/cyboghostginx/gemma-4-31B-it-Adetayo
- SGLang
How to use cyboghostginx/gemma-4-31B-it-Adetayo with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "cyboghostginx/gemma-4-31B-it-Adetayo" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cyboghostginx/gemma-4-31B-it-Adetayo", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "cyboghostginx/gemma-4-31B-it-Adetayo" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cyboghostginx/gemma-4-31B-it-Adetayo", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use cyboghostginx/gemma-4-31B-it-Adetayo with Docker Model Runner:
docker model run hf.co/cyboghostginx/gemma-4-31B-it-Adetayo
gemma-4-31B-it-Adetayo
Icelandic fine-tune of Google Gemma 4 31B instruction-tuned, targeting Icelandic morphology and grammar.
Scores (Miðeind Icelandic LLM leaderboard, official run)
| score | |
|---|---|
| this model, 6-task average | 70.09 |
| base gemma-4-31b-it, 6-task average | 71.24 |
| this model, 5-task local average (thinking off) | 83.59 |
Read those two top numbers carefully. The base is a reasoning model and its leaderboard run uses its thinking path; this fine-tune is submitted and evaluated with thinking disabled. On the 5-task subset, thinking-on is worth roughly 7 points on this base (85.0 with, 77.92 without), so the fine-tune and the base entry are not measured the same way. Where they are comparable, morphology is the gain.
A reasoning-preserving variant was also trained on rejection-sampled inflection traces. It held Wino, GED, Belebele and ARC exactly but moved inflection not at all, because 293 of 354 traces covered cases the base already solved. The method is sound, the data was too easy, so that variant was not published. This model is at its supervised fine-tuning ceiling.
Use
Standard text generation. Apply the Gemma chat template (tokenizer.apply_chat_template). Load the text path on GPU (Gemma4ForConditionalGeneration, .to("cuda")). Evaluate with thinking disabled.
License
Gemma derivative. Use is governed by the Gemma Terms of Use and the Gemma Prohibited Use Policy. "Gemma" is retained in the model name as required.
Training data and methodology are proprietary and are not distributed with the model.
- Downloads last month
- 7