Text Generation
PEFT
Safetensors
English
Hindi
llama
lora
fine-tuned
coding
reasoning
hindi
conversational
Instructions to use darshit-0391/Velora-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use darshit-0391/Velora-v1 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit") model = PeftModel.from_pretrained(base_model, "darshit-0391/Velora-v1") - Notebooks
- Google Colab
- Kaggle
Velora v1 β 8B Coding & Reasoning Model
Velora v1 is a fine-tuned version of Meta-Llama-3.1-8B-Instruct, trained using LoRA adapters on high-quality instruction datasets focused on coding, mathematics, and logical reasoning.
Model Details
| Field | Details |
|---|---|
| Developed by | Patel Darshit |
| Base Model | meta-llama/Meta-Llama-3.1-8B-Instruct |
| Model Type | Causal Language Model (QLoRA Fine-tuned) |
| Languages | English, Hindi |
| License | Apache 2.0 |
| Fine-tuning Method | LoRA (r=16), 4-bit Quantization |
Training Data
Velora v1 was trained on a curated mix of:
- OpenHermes-2.5 β 1M high-quality reasoning, coding, and logic examples
- Alpaca Cleaned β 52K general instruction-following examples
Training Hyperparameters
| Parameter | Value |
|---|---|
| LoRA Rank (r) | 16 |
| LoRA Alpha | 16 |
| Optimizer | AdamW 8-bit |
| Learning Rate | 2e-4 |
| Effective Batch Size | 8 |
| Training Steps | 5,000 |
| Precision | BFloat16 |
| GPU | NVIDIA A100 SXM4 40GB |
How to Use
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
from peft import PeftModel
base_model_id = "unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit"
adapter_id = "YOUR_USERNAME/Velora-v1-8B"
tokenizer = AutoTokenizer.from_pretrained(adapter_id)
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16
)
base_model = AutoModelForCausalLM.from_pretrained(
base_model_id,
quantization_config=bnb_config,
device_map="auto"
)
model = PeftModel.from_pretrained(base_model, adapter_id)
model.eval()
prompt = """<|begin_of_text|><|start_header_id|>user<|end_header_id|>
Write a Python function to check if a number is prime.<|eot_id|><|start_header_id|>assistant<|end_header_id|>
"""
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=512,
temperature=0.3,
do_sample=True,
pad_token_id=tokenizer.eos_token_id
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Capabilities
- β Multi-language coding (Python, JavaScript, C++, Java, Rust, Go, and more)
- β Step-by-step mathematical reasoning
- β Complex logical and algorithmic problem solving
- β Code explanation and debugging
- β Multi-step instruction following
Sample Output
Prompt: Write a Python asyncio script to fetch 3 URLs concurrently and explain how the event loop works.
Velora v1 Response: Correctly generated a production-ready asyncio + aiohttp implementation with a five-point structured explanation of the Python event loop model.
Limitations
- Maximum context: 2048 tokens
- May hallucinate on very niche topics
- Not suitable for medical or legal advice
Environmental Impact
| Field | Details |
|---|---|
| Hardware | NVIDIA A100 SXM4 40GB |
| Training Time | ~6 hours |
| Platform | AIKosh |
| Cloud Region | India |
- Downloads last month
- -
Model tree for darshit-0391/Velora-v1
Base model
meta-llama/Llama-3.1-8B Finetuned
meta-llama/Llama-3.1-8B-Instruct