File size: 3,767 Bytes
b6def5b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
# RunPod setup for kyrael + sorelith Krea2 LoRA training

## What's already done
- Both datasets are uploaded as private HF dataset repos:
  - `JBARU/kyrael-dataset`
  - `JBARU/sorelith-dataset`
- No captions yet -- `setup_pod.sh` captions them on the pod (Qwen2.5-VL-7B, fast on a real GPU).

## Files in this folder
- `setup_pod.sh` -- full bootstrap: installs musubi-tuner, downloads models, downloads datasets, captions, caches, trains both LoRAs
- `caption_dataset.py` -- auto-captioning script (called by setup_pod.sh)
- `dataset_kyrael.toml` / `dataset_sorelith.toml` -- musubi-tuner dataset configs

## Steps you need to do yourself (I can't do these for you)

### 1. Accept gated model access
Before anything else, visit https://huggingface.co/krea/Krea-2-Raw while logged in
and accept/request access (it auto-approves). Skip this and the download in step 4
of setup_pod.sh will fail.

### 2. Create a RunPod account + pod
1. Sign up at https://runpod.io (this needs your own payment method -- I can't do this part)
2. Go to **Pods** -> **Deploy**
3. Pick a GPU: **RTX 4090 (24GB)** is the recommended sweet spot for this job (~$0.34-0.69/hr)
4. Pick a template with PyTorch + CUDA pre-installed (e.g. "RunPod PyTorch 2.x")
5. Deploy the pod, wait for it to start
6. Open its **web terminal** (or connect via SSH if you set up a key)

### 3. Get the scripts onto the pod
Once you have the pod's terminal open:
```bash
mkdir -p /workspace
cd /workspace
```
Then either:
- **Easiest**: use the RunPod web UI's file upload to drop `caption_dataset.py`,
  `dataset_kyrael.toml`, and `dataset_sorelith.toml` into `/workspace/`
- Or paste their contents directly using `cat > filename.py << 'EOF' ... EOF` in the terminal

### 4. Run the bootstrap
Upload `setup_pod.sh` the same way, then:
```bash
chmod +x setup_pod.sh
./setup_pod.sh
```
It'll pause at `hf auth login` for you to paste your token interactively --
same rule as before, paste it only when prompted, never on the command line.

### 5. When it's done
Trained LoRAs land in `/workspace/output/kyrael/kyrael_lora.safetensors` and
`/workspace/output/sorelith/sorelith_lora.safetensors`. Push them back to HF
(so you can grab them locally) with:
```bash
hf upload <your-username>/kyrael-lora /workspace/output/kyrael --repo-type model
hf upload <your-username>/sorelith-lora /workspace/output/sorelith --repo-type model
```
Then locally: `hf download <your-username>/kyrael-lora --local-dir D:\ComfyModels\loras\kyrael`

### 6. Don't forget to stop the pod
RunPod bills by the hour while running -- stop/terminate it once training's
done so you're not paying for idle GPU time.

## Vaelith run: what changed after kyrael

- **Network volume: use at least 100GB**, not 50GB. Kyrael's pod hit disk-full
  twice at 50GB -- base models alone (~33GB) plus the Qwen2.5-VL-7B captioning
  model's cache (~16GB, easy to forget about) leave almost no margin at 50GB.
- **`num_repeats` dropped from 10 to 3** in `dataset_vaelith.toml`. Kyrael's
  LoRA came out overtrained/rigid (locked pose, completely resistant to style
  LoRA blending even at 5.0 weight) -- traced back to 34 images x 10 repeats x
  16 epochs = 5,440 total training exposures on a small dataset. Lower repeats
  should fix this without hurting identity retention.
- **`setup_pod_vaelith.sh` cleans up the captioning model cache immediately**
  after captioning finishes (`rm -rf /workspace/.cache`), instead of leaving
  it sitting there until disk fills up mid-training like last time.

## If something fails partway through
Each numbered section in `setup_pod.sh` is independent enough to re-run on its
own -- if training crashes on kyrael, you don't need to redo the downloads or
sorelith's caching, just re-run that one training command.