granite-4.2-8b-builder-lora-peft

PEFT LoRA adapter for the builder persona, trained on ibm-granite/granite-4.2-8b. A llama.cpp version of the same weights is at webAI-Official/granite-4.2-8b-builder-lora-GGUF.

Usage (Transformers + PEFT)

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained("ibm-granite/granite-4.2-8b", torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, "webAI-Official/granite-4.2-8b-builder-lora-peft")
tokenizer = AutoTokenizer.from_pretrained("webAI-Official/granite-4.2-8b-builder-lora-peft")

messages = [{"role": "user", "content": "Hello!"}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
print(tokenizer.decode(model.generate(inputs, max_new_tokens=256)[0][inputs.shape[-1]:], skip_special_tokens=True))

Call model.merge_and_unload() to fold the adapter into the base weights. The adapter was saved with PEFT 0.21.0; the tokenizer config uses the Transformers v5 format.

Files

  • adapter_model.safetensors, adapter_config.json: LoRA weights (fp32) and config
  • tokenizer.json, tokenizer_config.json, chat_template.jinja: tokenizer and chat template used in training
  • persona.json, run_config.json, train_metrics.json: persona settings, training config, and training metrics

Adapter details

  • Type: LoRA, rank 16, alpha 32, dropout 0.05
  • Target modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • Thinking mode: thinking
  • Training: 2776 examples, 3 epochs, lr 0.0001, cosine schedule, max seq len 8192, packing
  • Final step 522; train loss 0.3201; eval loss 0.3395
Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webAI-Official/granite-4.2-8b-builder-lora-peft

Adapter
(5)
this model