File size: 5,665 Bytes
88aaa02
 
b1dbc2a
 
 
 
 
 
 
 
 
 
 
88aaa02
b1dbc2a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
---
license: apache-2.0
base_model: Qwen/Qwen3-1.7B
tags:
- qiskit
- quantum-computing
- code-generation
- lora
- qlora
- unsloth
language:
- en
pipeline_tag: text-generation
---

<div align="center">

![Python 3.11](https://img.shields.io/badge/python-3.11-blue?logo=python&logoColor=white) ![Qiskit](https://img.shields.io/badge/Qiskit-2.5-6929C4?logo=qiskit&logoColor=white) ![Qwen3-1.7B](https://img.shields.io/badge/base%20model-Qwen3--1.7B-orange) ![QLoRA](https://img.shields.io/badge/finetune-QLoRA-9cf) ![Status](https://img.shields.io/badge/status-beta-blue)

</div>

# qrious-code-1.0

A Qiskit coding assistant: [`Qwen/Qwen3-1.7B`](https://huggingface.co/Qwen/Qwen3-1.7B) fine-tuned with QLoRA on quantum-computing code/QA data, designed to be paired with a retrieval-augmented generation (RAG) layer over the official Qiskit documentation.

This repository hosts the **merged** model. The LoRA adapter has been merged into the base weights and saved in standard fp16, so it loads directly with `AutoModelForCausalLM.from_pretrained`, no separate adapter step needed.

## Model Details

- **Base model:** [Qwen/Qwen3-1.7B](https://huggingface.co/Qwen/Qwen3-1.7B)
- **Fine-tuning method:** QLoRA (4-bit NF4 base, LoRA r=16/alpha=16 on `q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj`), merged into fp16 for this release
- **Trainable params during fine-tuning:** 17.4M / 1.74B (1.00%)
- **Training framework:** [Unsloth](https://github.com/unslothai/unsloth) + `trl.SFTTrainer`
- **Hardware:** single RTX 3050 Laptop GPU (6GB VRAM)
- **License:** Apache 2.0 (inherited from the base model)

## Training Data

Fine-tuned on a cleaned subset of [`samuellimabraz/quantum-assistant`](https://huggingface.co/datasets/samuellimabraz/quantum-assistant):

| split | raw rows | cleaned rows |
|---|---|---|
| train | 5837 | 4168 |
| val | 1239 | 887 |
| test | 1290 | 925 |

Cleaning removed image-dependent QA rows (hand-verified that the large majority genuinely depend on the image), de-duplicated near-identical questions, and filtered rows exceeding the 1024-token training context.

## Training Configuration

- 3 epochs, batch size 2 × gradient accumulation 8, packed 1024-token sequences (195 steps total)
- Optimizer `paged_adamw_8bit`, cosine LR schedule, peak LR 2e-4
- Train loss 2.70 → 0.77 (avg 0.989); eval loss 0.896 → 0.819 → 0.809 across 3 epochs (still improving, no overfitting divergence)
- Wall time: 48 minutes

## Evaluation

Scored on held-out test rows (unseen during training), tuned adapter vs. the un-tuned base model:

| Metric | Base | Tuned | Δ |
|---|---|---|---|
| Code pass@1 (744 code rows, executed in a sandbox against ground-truth tests) | 10.2% | 25.0% | **+14.8 pts** |
| QA judge score /5 (136 QA rows, judged by base Qwen3-1.7B) | 2.21 | 2.81 | **+0.60** |
| Instruction-following compliance | n/a | 1.0 | fixed 5-prompt probe set |

Code pass@1 improved in every one of 4 evaluation shards. The QA judge score improved in 3/4 shards, reported as-is rather than smoothed, since self-judging by a same-size base model is a known weak signal.

## Retrieval-Augmented Generation (companion, not baked into these weights)

This model is designed to run behind a RAG layer over the official [Qiskit/documentation](https://github.com/Qiskit/documentation) corpus (chunked, embedded with `BAAI/bge-small-en-v1.5`, reranked with `cross-encoder/ms-marco-MiniLM-L-6-v2`, indexed in Qdrant). In side-by-side testing, RAG grounding measurably reduces hallucination of deprecated/removed Qiskit APIs (`qiskit.opflow`, `qiskit_aqua`, `IBMQ.load_account()`, etc.), though it isn't perfect: the model can still recite deprecated syntax when retrieval doesn't surface the right page for a given query. The RAG pipeline itself is not part of this repository's weights; see [github.com/sarvan-2187/qrious-code-1.0](https://github.com/sarvan-2187/qrious-code-1.0) for the full serving code.

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "sarvan-2187/qrious-code-1.0"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="auto", device_map="auto")

messages = [
    {"role": "system", "content": "You are a Qiskit coding assistant."},
    {"role": "user", "content": "How do I create a Bell state in Qiskit?"},
]
text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```

> Pass `enable_thinking=False` in the chat template: Qwen3's default reasoning mode otherwise spends the generation budget on unrequested `<think>` output before answering.

## Limitations

- QA-answer quality improved less clearly than code-generation correctness (self-judged by a same-size base model, a known weak evaluation signal).
- Can still recite deprecated Qiskit APIs when used without RAG grounding, or when RAG retrieval doesn't surface the right context for a query.
- Trained/evaluated at a 1024-token context; serving at a 4096-token context works (verified) but wasn't part of the original training distribution.

## Citation

Not a formal research release, a personal project. If referencing it: base model [Qwen3](https://huggingface.co/Qwen/Qwen3-1.7B), dataset [`samuellimabraz/quantum-assistant`](https://huggingface.co/datasets/samuellimabraz/quantum-assistant), fine-tuned with [Unsloth](https://github.com/unslothai/unsloth).