File size: 7,958 Bytes
1607c63
 
 
 
 
 
 
 
 
 
9e88762
 
1607c63
 
 
38d66e6
1607c63
9e88762
 
 
1607c63
38d66e6
1607c63
38d66e6
1607c63
 
 
 
 
38d66e6
 
 
9e88762
 
38d66e6
9e88762
 
1607c63
 
38d66e6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9e88762
 
 
38d66e6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9e88762
 
 
38d66e6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
65c7fab
38d66e6
 
 
 
 
 
fe99dc8
 
9e88762
 
38d66e6
1607c63
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38d66e6
1607c63
38d66e6
1607c63
 
 
 
 
 
38d66e6
 
 
1607c63
38d66e6
 
 
 
 
 
 
1607c63
38d66e6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
65c7fab
38d66e6
 
 
 
 
 
 
 
 
 
 
 
1607c63
 
38d66e6
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
# Training β€” Enterprise Contract Guardian

Re-runnable training pipeline for the API Contract Validator environment, using GRPO from TRL with LoRA adapters via Unsloth.

## Files

| File | Purpose |
|---|---|
| `baseline.py` | Run the untrained model on every task; write `baseline_scores.json` |
| `train.py` | GRPO training loop β€” connects to env, rolls out, trains LoRA |
| `run_in_hf_jobs.py` | Self-bootstrapping HF Jobs launcher used for the submitted run |
| `run_trained_inference.py` | HF Jobs inference script for the trained LoRA adapter |
| `plot.py` | Build `reward_curve.png` and `before_after.png` for the README |
| `grpo_colab.ipynb` | One-click Colab notebook (open in Colab β†’ Runtime β†’ Run all) |

## Recommended path β€” HF Jobs (best for the finale)

HF Jobs runs in the cloud, doesn't disconnect, and bills against your hackathon credit.

Important: `hf jobs uv run` uploads exactly one Python file. The submitted run therefore used [`run_in_hf_jobs.py`](run_in_hf_jobs.py), which clones this repo inside the job and then calls `training.train.main()`. This avoids import errors from sibling modules such as `inference.py`, `client.py`, and `server/*`.

### Run 1 β€” Smoke test (~$0.30, 5 min)

Verifies the pipeline works end-to-end before committing to a long run.

```bash
hf jobs uv run \
    --flavor t4-small \
    -s HF_TOKEN -s WANDB_API_KEY \
    -e BASE_MODEL=unsloth/Qwen2.5-1.5B-Instruct-bnb-4bit \
    -e ENV_URL=https://pushpam14-api-contract-validator.hf.space \
    -e MAX_STEPS=10 \
    -e NUM_GENERATIONS=2 \
    -e GIT_REF=main \
    -e WANDB_RUN=smoke-test \
    -d \
    api_contract_validator/training/run_in_hf_jobs.py
```

If this errors, **don't proceed**. Fix the error, re-run smoke test until it returns clean.

### Run 2 β€” Main training, **Qwen2.5-7B on L4** (~$2.40, ~2 hours)

Best balance of model size, speed, and cost for our $60 budget. L4 has 24 GB which fits Qwen2.5-7B with 4-bit quantisation + LoRA r=16.

```bash
hf jobs uv run \
    --flavor l4x1 \
    -s HF_TOKEN -s WANDB_API_KEY \
    -e BASE_MODEL=unsloth/Qwen2.5-7B-Instruct-bnb-4bit \
    -e ENV_URL=https://pushpam14-api-contract-validator.hf.space \
    -e MAX_STEPS=300 \
    -e NUM_GENERATIONS=4 \
    -e LORA_R=16 \
    -e LORA_ALPHA=32 \
    -e WANDB_PROJECT=openenv-contract-guardian \
    -e WANDB_RUN=grpo-7b-l4-300steps \
    -e PUSH_TO_HUB=pushpam14/api-contract-validator-grpo-7b \
    -e GIT_REF=main \
    -d \
    api_contract_validator/training/run_in_hf_jobs.py
```

### Run 3 β€” Insurance run on second account (~$0.40, ~45 min)

Use your second HF account in parallel as a safety net. Smaller model = faster, more dramatic improvement curve. If Run 2 produces a beautiful curve we ship that; if Run 2 has issues we ship this one.

```bash
# Use your SECOND HF account's token here
hf jobs uv run \
    --flavor t4-small \
    -s HF_TOKEN -s WANDB_API_KEY \
    -e BASE_MODEL=unsloth/Qwen2.5-1.5B-Instruct-bnb-4bit \
    -e ENV_URL=https://pushpam14-api-contract-validator.hf.space \
    -e MAX_STEPS=200 \
    -e WANDB_RUN=grpo-1.5b-t4-200steps \
    -e PUSH_TO_HUB=YOUR_SECOND_ACCOUNT/api-contract-validator-grpo-1.5b \
    -e GIT_REF=main \
    -d \
    api_contract_validator/training/run_in_hf_jobs.py
```

### Hardware ↔ model size cheatsheet

| Flavor | VRAM | $/hr | Fits (4-bit + LoRA) |
|---|---|---|---|
| `t4-small` | 16 GB | $0.40 | Up to 3B |
| `l4x1` | 24 GB | $0.80 | Up to 8B comfortably |
| `a10g-large` | 24 GB | $1.50 | Up to 8B, faster than L4 |
| `a100-large` | 80 GB | $3.50 | 14B fp16 or 70B 4-bit |
| `h100x1` | 80 GB | $4.50 | Same as A100 but ~2Γ— faster |

For our env, **`l4x1` + Qwen2.5-7B is the sweet spot.**

### Monitoring the run

```bash
# Watch the job logs in real time
hf jobs logs <job-id> --follow

# List recent jobs
hf jobs list

# WandB run will be auto-linked in the job logs; for public review, link the WandB report URL in the README
```

## Alternative: Colab notebook (if HF Jobs is unavailable)

Open `grpo_colab.ipynb` in Colab. Set `HF_TOKEN` and `WANDB_API_KEY` in the secrets pane. Hit **Runtime β†’ Run all**. Free T4, but disconnects after 3 hours and only fits the 1.5B model.

If Hugging Face's notebook viewer shows a blank/white page for the `.ipynb`, open it in Google Colab or download it and open with Jupyter. The submitted metrics are not stored as heavy notebook outputs; they are available through the public WandB report and the committed `results/reward_curve.png`, `results/training_state.json`, and `results/training_full_log.txt` files.

The notebook is intentionally committed without outputs so reviewers can run it cleanly. The completed submitted run outputs are committed under `api_contract_validator/results/` and linked from the main README.

## Alternative: Local GPU

```bash
pip install trl unsloth wandb matplotlib datasets
export HF_TOKEN="hf_..."
export WANDB_API_KEY="..."
export ENV_URL="http://localhost:7860"   # or your HF Space URL

# 1. Start the env server in another terminal
uvicorn server.app:app --host 0.0.0.0 --port 7860

# 2. Baseline
python training/baseline.py

# 3. Train
python training/train.py

# 4. Inference with trained adapter
export MODEL_NAME="<your-username>/api-contract-validator-grpo"
export SCORES_OUT_PATH="trained_scores.json"
python inference.py

# 5. Plots
python training/plot.py
```

## Key environment variables

| Variable | Default | Notes |
|---|---|---|
| `HF_TOKEN` | β€” | Required. Used for both inference (router) and Hub push |
| `WANDB_API_KEY` | β€” | Optional. If set, training logs go to WandB |
| `BASE_MODEL` | `unsloth/Qwen2.5-7B-Instruct-bnb-4bit` | 7B fits on L4 (24 GB) with 4-bit |
| `ENV_URL` | `http://localhost:7860` | Local server or deployed HF Space |
| `MAX_STEPS` | `300` | GRPO steps. ~2 hours on L4 |
| `NUM_GENERATIONS` | `4` | Completions per prompt for relative ranking |
| `LORA_R` | `16` | LoRA rank |
| `PUSH_TO_HUB` | β€” | `<username>/<repo>` β€” push trained adapter |

## What the judges look at

The hackathon's "Improvement in Rewards" 20% criterion explicitly asks for a **before vs after comparison**. Per the official Q&A:

> "You're expected to show before vs after behavior. Run inference using both models and include the comparison (metrics, rewards, or outputs) in the README."

So after training, you must commit ALL FOUR of these:

```
baseline_scores.json                                     # already committed (the "before")
trained_scores.json                                      # post-training inference output
api_contract_validator/results/reward_curve.png          # GRPO training curve
api_contract_validator/results/before_after.png          # bar chart comparison
```

## Post-training commit checklist

Once the Colab notebook finishes, run this on your laptop:

```bash
cd ~/work/hackathon/hackathon-api-contract-validator

# 1. Pull the four artifacts down from Colab into local repo
#    (Colab File pane β†’ right-click β†’ Download for each)
#
#    Place them at:
#      ./trained_scores.json
#      ./api_contract_validator/results/reward_curve.png
#      ./api_contract_validator/results/before_after.png

# 2. Edit api_contract_validator/README.md
#    - Replace each "_(after training)_" placeholder in the Before vs After
#      table with the score from trained_scores.json
#    - Uncomment the two `<!-- ![](results/...) -->` lines so the plots render
#    - Add the WandB report URL in the Links section (and to the WandB row in
#      the Training Results table)

# 3. Run validation
PYTHONPATH=api_contract_validator python3 -m pytest \
    api_contract_validator/tests/test_environment.py -q

# 4. Commit + push
git add trained_scores.json \
        api_contract_validator/results/*.png \
        api_contract_validator/README.md
git commit -m "Add post-training results: trained scores + reward plots"
git push
```

That single push is the "after" half of the before-vs-after evidence judges grade.