| # Training & Fine-tuning |
|
|
| ## Data scope |
|
|
| Train or fine-tune only with captcha images from PHPWind deployments you own |
| or are explicitly authorized to evaluate. Keep source, consent, and version |
| metadata with the dataset so results remain interpretable. |
|
|
| ## Environment |
|
|
| ```bash |
| pip install torch onnx onnxscript pillow numpy opencv-python |
| ``` |
|
|
| ## Training from scratch |
|
|
| ### 1. Prepare data |
|
|
| Label format `labels.json`: |
| ```json |
| { |
| "captcha_001.png": {"label": "4821"}, |
| "captcha_002.png": {"label": "9037"} |
| } |
| ``` |
|
|
| Image dir: ~150x60 RGB captcha PNGs; filenames must match label keys. |
|
|
| ### 2. Train |
|
|
| ```bash |
| python scripts/train_fixed.py \ |
| <labels.json> <image_dir> <output.onnx> \ |
| 0.08 # val fraction |
| 500 # epochs |
| 1 # augmentation on/off (1=on) |
| rgb # RGB input (important: do NOT use grayscale) |
| ``` |
|
|
| Example: |
| ```bash |
| python scripts/train_fixed.py \ |
| data/1000_labels.json data/1000_raw \ |
| models/my_model.onnx \ |
| 0.08 500 1 rgb |
| ``` |
|
|
| Outputs: |
| - `<output>.onnx` β final ONNX model |
| - `<output>.onnx.pt` β best-epoch PyTorch weights |
|
|
| ### 3. Verify inference |
|
|
| ```python |
| import numpy as np, onnxruntime as ort |
| from PIL import Image |
| sess = ort.InferenceSession("my_model.onnx", providers=["CPUExecutionProvider"]) |
| im = Image.open("captcha.png").convert("RGB").resize((160,64), Image.BILINEAR) |
| x = np.asarray(im, dtype=np.float32).transpose(2,0,1)[None] / 255.0 |
| logits = sess.run(None, {"input": x})[0] |
| code = "".join(str(int(logits[0,p].argmax())) for p in range(4)) |
| ``` |
|
|
| ## Fine-tuning |
|
|
| `train_fixed.py` auto-loads an existing `<output>.onnx.pt` and resumes from it: |
|
|
| ```bash |
| cp models/captcha_1000_raw.onnx.pt models/finetuned.onnx.pt |
| python scripts/train_fixed.py <new_labels.json> <new_image_dir> models/finetuned.onnx 0.08 200 1 rgb |
| ``` |
|
|
| - 50-200 new labeled samples are enough for small changes. |
| - Mix old + new data when the generator changes a lot, to avoid catastrophic forgetting. |
| - Keep the same preprocessing (RGB 160x64 /255, no denoise) and augmentation. |
|
|
| ## Parameters |
|
|
| | Param | Default | Description | |
| |---|---|---| |
| | val_frac | 0.1 | validation fraction | |
| | epochs | 400 | total epochs | |
| | aug | 1 | augmentation toggle | |
| | rgb | - | pass `rgb` for 3-channel input (recommended) | |
| |
| ## Key lessons |
| |
| 1. **Use RGB, not grayscale** β the digits are colored; grayscale loses information. |
| 2. **Do NOT denoise** β morphological opening removes thin strokes, dropping val from 88.6% to 79.8%. With enough data the model learns to ignore noise lines itself. |
| 3. **Position-preserving architecture, never global average pooling** β GAP destroys spatial position (val β 0%); per-column grouped pooling (20 cols β 4 positions) restores generalization. |
| 4. **Human labels >> auto vision labels** β auto labels with ~35% noise crashed the model to 0%; clean manual labels are essential. |
| 5. **No Gaussian noise augmentation** β the captcha already has noise lines; adding more obscures the training signal. |
| |