File size: 3,183 Bytes
8ad8408
5e34335
8ad8408
 
 
5e34335
8ad8408
5e34335
8ad8408
5e34335
8ad8408
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
# Configuration parity

The files in `configs/point2rbox_v3/` reproduce the official mmrotate
Point2RBox-v3 configuration. `tests/parity/test_L0_v3_config.py` compares the
flattened values with stored reference goldens.

## Framework mappings

| PyTorch/mmrotate | Jittor/JDet | Equivalent behavior |
|---|---|---|
| `mmdet.ResNet(out_indices=...)` | `Resnet50(return_stages=...)` | same feature stages |
| torchvision ResNet-50 initialization | `pretrained=True` | same pretrained source |
| `mmdet.FPN` | `FPN` | matching channels and levels |
| `mmdet.FocalLoss` | `MMDetFocalLoss` | mmdetection reduction semantics |
| `GWDLoss` | `GDLoss(loss_type='gwd')` | matching Gaussian distance |
| AdamW + `clip_grad` | JDet AdamW with `grad_clip` | global L2 clip at 35 |
| LinearLR + MultiStepLR | `LinearWarmupMultiStepLR` | pointwise-equal LR sequence |
| `SetEpochInfoHook` | runner `model.set_epoch(epoch)` | epoch switches preserved |
| mmrotate qbox transforms | `P2RV2DOTADataset` | qbox/rbox and point labels |
| mmrotate resize/flip | `MMRotateResize` / `MMRotateRandomFlip` | coordinate and angle parity |

## Locked training values

- 12 epochs, evaluation at epoch 12, checkpoint every epoch.
- AdamW learning rate `5e-5`, gradient clip `35`.
- Linear warmup from factor `1/3` for 500 iterations; learning-rate milestones
  at epochs 8 and 11 with factor 0.1.
- End-to-end batch size 2 and second-stage batch size 4 on one GPU.
- End-to-end weight decay 0.05; second-stage weight decay 0.005.
- Five FPN strides `[8, 16, 32, 64, 128]` for v3.
- Self-supervision probabilities `[0.68, 0.07, 0.25]`.
- Epoch 6 switches for edge supervision, pseudo-label assignment and copy-paste
  routing. The upstream key spelling
  `label_assign_pseudo_label_switch_eopch` is intentionally preserved.
- Validation points to the reference trainval split, matching the official
  diagnostic protocol.

## SAM filtering configuration

`configs/point2rbox_v3/_base_sam-dotav1-0.py` preserves the complete
class-specific filtering table from the reference. The L0 test compares every
key, value and tuple/list type. Notable intentional values include:

- classes 3, 8 and 10 use circularity weight `-3` with circularity penalty 100;
- prompt points outside a mask receive the reference hard center-alignment
  penalty;
- the internal fallback filter table is kept separate from the config table,
  because the reference values differ and the configured training path always
  passes the explicit table.

## Infrastructure-only adaptation

The validated Jittor setup uses `num_workers=0`. Jittor 1.3.8.5 can deadlock in
the multiprocessing dataset ring buffer for this variable-instance workload.
This changes loading concurrency only; sample definitions, transforms and
training math remain unchanged.

## Upstream MobileSAM cache behavior

The upstream builder enters evaluation mode before loading the TinyViT state
dict, which can leave a cached attention-bias tensor derived from initialization
values. The Jittor builder loads weights first and then refreshes the evaluation
cache. This is the deterministic checkpoint-loading behavior documented in
[porting_notes.md](porting_notes.md).