How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="auryn-macmillan/boostedv1")
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("auryn-macmillan/boostedv1")
model = AutoModelForCausalLM.from_pretrained("auryn-macmillan/boostedv1", device_map="auto")
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

BoostedV1

Continued LoRA training of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B for improved reasoning and code generation. This model is the merged output of the boostedv1train pipeline: 400 MLX LoRA steps (Apple Silicon) + 550 Phase-1 continuation steps (dual RTX 3090, bf16).

Architecture

  • Base: DeepSeek-R1-Distill-Qwen-1.5B (Qwen2ForCausalLM, 1.5B params)
  • LoRA rank 8, alpha 160, on layers 20-27 (q/k/v/o + gate/up/down)
  • Merged into a standalone model (no LoRA needed at inference)

Training

Stage Platform Steps Batch Seq len Data
MLX run 1 Apple Silicon 150 4 512 OpenCodeInstruct
MLX run 2 Apple Silicon 250 2 1024 OpenCodeInstruct
Phase 1 2x RTX 3090 550 16 (eff) 2048 OpenThoughts + OpenR1-Math + OpenCodeInstruct

Evaluation

Benchmark Score
GSM8K 46.0%
HumanEval (pass@1) 7.3%

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('auryn-macmillan/boostedv1') tok = AutoTokenizer.from_pretrained('auryn-macmillan/boostedv1') inputs = tok('What is 2+2?', return_tensors='pt') out = model.generate(**inputs, max_new_tokens=128) print(tok.decode(out[0]))

Notes

  • Standard Qwen2 architecture, no custom code, no trust_remote_code needed.
  • Trained in an isolated container; repo contains no training code or data.
Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support