boostedv1 / README.md
auryn-macmillan's picture
BoostedV1: merged 550-step phase-1 LoRA on DeepSeek-R1-Distill-Qwen-1.5B
296969f verified
|
Raw
History Blame Contribute Delete
1.59 kB
metadata
license: mit
language:
  - en
library_name: transformers
pipeline_tag: text-generation
tags:
  - reasoning
  - deepseek
  - lora

BoostedV1

Continued LoRA training of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B for improved reasoning and code generation. This model is the merged output of the boostedv1train pipeline: 400 MLX LoRA steps (Apple Silicon) + 550 Phase-1 continuation steps (dual RTX 3090, bf16).

Architecture

  • Base: DeepSeek-R1-Distill-Qwen-1.5B (Qwen2ForCausalLM, 1.5B params)
  • LoRA rank 8, alpha 160, on layers 20-27 (q/k/v/o + gate/up/down)
  • Merged into a standalone model (no LoRA needed at inference)

Training

Stage Platform Steps Batch Seq len Data
MLX run 1 Apple Silicon 150 4 512 OpenCodeInstruct
MLX run 2 Apple Silicon 250 2 1024 OpenCodeInstruct
Phase 1 2x RTX 3090 550 16 (eff) 2048 OpenThoughts + OpenR1-Math + OpenCodeInstruct

Evaluation

Benchmark Score
GSM8K 46.0%
HumanEval (pass@1) 7.3%

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('auryn-macmillan/boostedv1') tok = AutoTokenizer.from_pretrained('auryn-macmillan/boostedv1') inputs = tok('What is 2+2?', return_tensors='pt') out = model.generate(**inputs, max_new_tokens=128) print(tok.decode(out[0]))

Notes

  • Standard Qwen2 architecture, no custom code, no trust_remote_code needed.
  • Trained in an isolated container; repo contains no training code or data.