image

Llama-3.2-3B-AgentInstruct

Llama-3.2-3B-AgentInstruct is Llama-3.2-3B-Instruct fine-tuned to act as an agent. Given a task, it reasons step by step ("Think"), takes one action ("Act"), reads the environment's response, and repeats until the task is solved.

Dataset

Training used zai-org/AgentInstruct, a curated set of 1,866 ReAct-style trajectories across six real-world agent tasks:

Task Trajectories What the agent does
Operating System 195 Solves tasks in a Linux shell using bash
Database 538 Writes SQL to answer questions over tables
ALFWorld 336 Completes household tasks in a text-based world
WebShop 351 Searches and selects products on a shopping site
Mind2Web 122 Chooses actions to navigate real websites
Knowledge Graph 324 Answers questions by querying a knowledge graph

The trajectories were generated with GPT-4 and filtered by reward, so only successful runs are kept. Every action is paired with a written thought, which gives the model its reasoning supervision.

Training

Fine-tuned with QLoRA and merged into the base weights (16-bit). Loss was computed only on the assistant's turns, and every conversation was used in full with no truncation. The model's native Llama 3 chat template was used throughout.

What the model can do

  • Follow the Think/Act loop and stop cleanly when the task is done.
  • Use the output of previous steps (command results, search results, errors) to decide the next action.
  • Adapt to new tasks within the same families, such as new shell commands or new SQL questions.

It learns the agent behavior and output formats, not new world knowledge. It uses a plain-text ReAct protocol and does not use Llama's native JSON tool-calling. For best results, prompt it in the style of the original AgentInstruct tasks.

Usage

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

repo = "Rumiii/Llama-3.2-3B-AgentInstruct"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.float16, device_map="auto")

instructions = """You are an assistant that will act like a person, I'will play the role of linux(ubuntu) operating system. Your goal is to implement the operations required by me or answer to the question proposed by me. For each of your turn, you should first think what you should do, and then take exact one of the three actions: "bash", "finish" or "answer".

1. If you think you should execute some bash code, take bash action, and you should print like this:

Think: put your thought here.

Act: bash

```bash
# put your bash code here
```

2. If you think you have finished the task, take finish action, and you should print like this:

Think: put your thought here.

Act: finish

3. If you think you have got the answer to the question, take answer action, and you should print like this:

Think: put your thought here.

Act: answer(Your answer to the question should be put in this pair of parentheses)

If the output is too long, I will truncate it. Attention, your bash code should not contain any input operation. Once again, you should take only exact one of the three actions in each turn.

Now, my problem is:

"""

task = "How many lines are there in the file notes.txt in the current directory?"
messages = [{"role": "user", "content": instructions + task}]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, add_special_tokens=False, return_tensors="pt").to(model.device)

output = model.generate(
    **inputs,
    max_new_tokens=512,
    do_sample=False,
    eos_token_id=tokenizer.convert_tokens_to_ids("<|eot_id|>"),
)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Example output:

Think: I need to find the number of lines in the file notes.txt. I can use the wc -l command.

Act: bash

```bash
wc -l notes.txt
```

To run it as an agent, execute the returned bash command, append the assistant's reply to messages, add your command's output as the next user message in the form The output of the OS:\n<output>, and generate again until the model replies with Act: answer(...) or Act: finish.

Safety note: run any model-generated commands inside a sandbox or container.

Limitations

  • Evaluated only on small informal tests, not on formal agent benchmarks.
  • A 3B model can make mistakes in multi-step tasks; verify its actions before relying on them.
  • Tuned for the AgentInstruct task formats; other prompt styles may work less reliably.

Built with Llama.

Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Rumiii/Llama-3.2-3B-AgentInstruct

Finetuned
(2017)
this model

Dataset used to train Rumiii/Llama-3.2-3B-AgentInstruct