Instructions to use VolkanSimsir/LLaMA-3-8B-GRPO-math-tr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VolkanSimsir/LLaMA-3-8B-GRPO-math-tr with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="VolkanSimsir/LLaMA-3-8B-GRPO-math-tr") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("VolkanSimsir/LLaMA-3-8B-GRPO-math-tr") model = AutoModelForCausalLM.from_pretrained("VolkanSimsir/LLaMA-3-8B-GRPO-math-tr", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use VolkanSimsir/LLaMA-3-8B-GRPO-math-tr with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VolkanSimsir/LLaMA-3-8B-GRPO-math-tr" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VolkanSimsir/LLaMA-3-8B-GRPO-math-tr", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VolkanSimsir/LLaMA-3-8B-GRPO-math-tr
- SGLang
How to use VolkanSimsir/LLaMA-3-8B-GRPO-math-tr with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "VolkanSimsir/LLaMA-3-8B-GRPO-math-tr" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VolkanSimsir/LLaMA-3-8B-GRPO-math-tr", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "VolkanSimsir/LLaMA-3-8B-GRPO-math-tr" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VolkanSimsir/LLaMA-3-8B-GRPO-math-tr", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Studio
How to use VolkanSimsir/LLaMA-3-8B-GRPO-math-tr with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for VolkanSimsir/LLaMA-3-8B-GRPO-math-tr to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for VolkanSimsir/LLaMA-3-8B-GRPO-math-tr to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for VolkanSimsir/LLaMA-3-8B-GRPO-math-tr to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="VolkanSimsir/LLaMA-3-8B-GRPO-math-tr", max_seq_length=2048, ) - Docker Model Runner
How to use VolkanSimsir/LLaMA-3-8B-GRPO-math-tr with Docker Model Runner:
docker model run hf.co/VolkanSimsir/LLaMA-3-8B-GRPO-math-tr
Model Details
Base Model: ytu-ce-cosmos/Turkish-Llama-8b-DPO-v0.1
Training Time: 3 hours with A40
Lora config:
- lora_r: 32
- lora_alpha:32
Model Description
This model has been trained using the GRPO (Guided Reward Preference Optimization) method with the Turkish GSM8K dataset to enhance its mathematical reasoning capabilities. However, as with similar large language models, its responses may contain errors or biases. Therefore, outputs should be carefully evaluated, especially in applications where accuracy is critical, and additional verification steps are recommended.
Example Usage
Install
!pip install -U transformers bitsandbytes
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
model_id = "VolkanSimsir/LLaMA-3-8B-GRPO-math-tr"
tokenizer = AutoTokenizer.from_pretrained(model_id)
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype="float16",
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=quantization_config,
device_map="auto"
)
question = """ Randy'nin çiftliğinde 60 mango ağacı var.
Ayrıca mango ağaçlarının yarısından 5 tane daha az Hindistan cevizi ağacı var.
Randy'nin çiftliğinde toplam kaç ağaç var?
"""
inputs = tokenizer(question, return_tensors="pt").to(0)
outputs = model.generate(**inputs)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Evaluation
Citation Information
@software{VolkanSimsir,
author = {VolkanSimsir},
title = {VolkanSimsir/LLaMA-3-8B-GRPO-math-tr},
year = 2025,
url = {https://huggingface.co/VolkanSimsir/LLaMA-3-8B-GRPO-math-tr}
}
Contact
- Downloads last month
- 8
Model tree for VolkanSimsir/LLaMA-3-8B-GRPO-math-tr
Base model
meta-llama/Meta-Llama-3-8B