Instructions to use Rumiii/Llama-3.2-3B-AgentInstruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Rumiii/Llama-3.2-3B-AgentInstruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Rumiii/Llama-3.2-3B-AgentInstruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Rumiii/Llama-3.2-3B-AgentInstruct") model = AutoModelForCausalLM.from_pretrained("Rumiii/Llama-3.2-3B-AgentInstruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Rumiii/Llama-3.2-3B-AgentInstruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Rumiii/Llama-3.2-3B-AgentInstruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Rumiii/Llama-3.2-3B-AgentInstruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Rumiii/Llama-3.2-3B-AgentInstruct
- SGLang
How to use Rumiii/Llama-3.2-3B-AgentInstruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Rumiii/Llama-3.2-3B-AgentInstruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Rumiii/Llama-3.2-3B-AgentInstruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Rumiii/Llama-3.2-3B-AgentInstruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Rumiii/Llama-3.2-3B-AgentInstruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Rumiii/Llama-3.2-3B-AgentInstruct with Docker Model Runner:
docker model run hf.co/Rumiii/Llama-3.2-3B-AgentInstruct
Llama-3.2-3B-AgentInstruct
Llama-3.2-3B-AgentInstruct is Llama-3.2-3B-Instruct fine-tuned to act as an agent. Given a task, it reasons step by step ("Think"), takes one action ("Act"), reads the environment's response, and repeats until the task is solved.
Dataset
Training used zai-org/AgentInstruct, a curated set of 1,866 ReAct-style trajectories across six real-world agent tasks:
| Task | Trajectories | What the agent does |
|---|---|---|
| Operating System | 195 | Solves tasks in a Linux shell using bash |
| Database | 538 | Writes SQL to answer questions over tables |
| ALFWorld | 336 | Completes household tasks in a text-based world |
| WebShop | 351 | Searches and selects products on a shopping site |
| Mind2Web | 122 | Chooses actions to navigate real websites |
| Knowledge Graph | 324 | Answers questions by querying a knowledge graph |
The trajectories were generated with GPT-4 and filtered by reward, so only successful runs are kept. Every action is paired with a written thought, which gives the model its reasoning supervision.
Training
Fine-tuned with QLoRA and merged into the base weights (16-bit). Loss was computed only on the assistant's turns, and every conversation was used in full with no truncation. The model's native Llama 3 chat template was used throughout.
What the model can do
- Follow the Think/Act loop and stop cleanly when the task is done.
- Use the output of previous steps (command results, search results, errors) to decide the next action.
- Adapt to new tasks within the same families, such as new shell commands or new SQL questions.
It learns the agent behavior and output formats, not new world knowledge. It uses a plain-text ReAct protocol and does not use Llama's native JSON tool-calling. For best results, prompt it in the style of the original AgentInstruct tasks.
Usage
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
repo = "Rumiii/Llama-3.2-3B-AgentInstruct"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.float16, device_map="auto")
instructions = """You are an assistant that will act like a person, I'will play the role of linux(ubuntu) operating system. Your goal is to implement the operations required by me or answer to the question proposed by me. For each of your turn, you should first think what you should do, and then take exact one of the three actions: "bash", "finish" or "answer".
1. If you think you should execute some bash code, take bash action, and you should print like this:
Think: put your thought here.
Act: bash
```bash
# put your bash code here
```
2. If you think you have finished the task, take finish action, and you should print like this:
Think: put your thought here.
Act: finish
3. If you think you have got the answer to the question, take answer action, and you should print like this:
Think: put your thought here.
Act: answer(Your answer to the question should be put in this pair of parentheses)
If the output is too long, I will truncate it. Attention, your bash code should not contain any input operation. Once again, you should take only exact one of the three actions in each turn.
Now, my problem is:
"""
task = "How many lines are there in the file notes.txt in the current directory?"
messages = [{"role": "user", "content": instructions + task}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, add_special_tokens=False, return_tensors="pt").to(model.device)
output = model.generate(
**inputs,
max_new_tokens=512,
do_sample=False,
eos_token_id=tokenizer.convert_tokens_to_ids("<|eot_id|>"),
)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Example output:
Think: I need to find the number of lines in the file notes.txt. I can use the wc -l command.
Act: bash
```bash
wc -l notes.txt
```
To run it as an agent, execute the returned bash command, append the assistant's reply to messages, add your command's output as the next user message in the form The output of the OS:\n<output>, and generate again until the model replies with Act: answer(...) or Act: finish.
Safety note: run any model-generated commands inside a sandbox or container.
Limitations
- Evaluated only on small informal tests, not on formal agent benchmarks.
- A 3B model can make mistakes in multi-step tasks; verify its actions before relying on them.
- Tuned for the AgentInstruct task formats; other prompt styles may work less reliably.
Built with Llama.
- Downloads last month
- -
Model tree for Rumiii/Llama-3.2-3B-AgentInstruct
Base model
meta-llama/Llama-3.2-3B-Instruct