agentic-rag-lora

LoRA adapter only (not full base weights) for HuggingFaceTB/SmolLM-135M, lightly trained on CPU with short instruction/response texts about agentic systems and RAG.

Comparison note: This is a tiny PEFT adapter (~SmolLM-135M base), not a 4-bit 7B chat model and not a paid Gradio Space. Load the base model, then attach this adapter.

Important

  • Repository contains adapter weights + tokenizer config, not a full 135M rehost of new authorship.
  • You must load the base SmolLM-135M and then attach this adapter.
  • Intended as a small demo of domain-adapted instruction style for agentic/RAG docs — not a production LLM.

Usage

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "HuggingFaceTB/SmolLM-135M"
tok = AutoTokenizer.from_pretrained("hharsha/agentic-rag-lora")
model = AutoModelForCausalLM.from_pretrained(base)
model = PeftModel.from_pretrained(model, "hharsha/agentic-rag-lora")

prompt = "Explain hybrid search in a RAG pipeline:"
inputs = tok(prompt, return_tensors="pt")
print(tok.decode(model.generate(**inputs, max_new_tokens=64)[0], skip_special_tokens=True))

Training summary

Base HuggingFaceTB/SmolLM-135M
Method PEFT LoRA (r=8, alpha=16, q_proj/v_proj)
Epochs 1 (CPU, float32)
Trainable params ~461k (0.34%)
Data Showcase project summaries + synthetic agentic/RAG instructions

Related

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hharsha/agentic-rag-lora

Adapter
(24)
this model

Dataset used to train hharsha/agentic-rag-lora