isaiahbjork/cot-logic-reasoning
Viewer • Updated • 10.5k • 110 • 18
How to use alibidaran/GRPO_LLAMA3-instructive_reasoning1 with Transformers:
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("alibidaran/GRPO_LLAMA3-instructive_reasoning1", device_map="auto")This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
We are using MMLU dataset in different tasks. Here are the results of using 100 random samples of MMLU dataset.