File size: 4,178 Bytes
ac02d47
bc15982
 
 
 
 
 
 
 
 
 
 
 
ac02d47
bc15982
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
---
license: agpl-3.0
library_name: pytorch
pipeline_tag: object-detection
tags:
- remote-sensing
- visual-grounding
- sar
- optical
- cross-domain
- mixture-of-experts
- contrastive-learning
- benchmark
---

# OptiSAR-Net++ — Official Weights

Official trained weights (`best.pt`) of **OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding**.

OptiSAR-Net++ is a **transformer-free** framework for Cross-Domain Remote Sensing Visual Grounding (CD-RSVG): a single unified model localizes targets described by free-form natural language in **both optical and SAR** remote sensing imagery, replacing a heavy Transformer decoder with **contrastive region–text matching**.

- 🤗 **Dataset**: [JunDong-dev/OptSAR-RSVG](https://huggingface.co/datasets/JunDong-dev/OptSAR-RSVG)
- 💻 **Code**: [GitHub — JunDong-dev/OptiSAR-Net-PlusPlus](https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus)

## 🏗️ Architecture

The checkpoint follows a YOLOE-style single-stage detector extended with:

| Component | Location | Role |
|---|---|---|
| **PLoRA-MoE** | Backbone | Patch-level low-rank-adaptation Mixture-of-Experts for optical/SAR feature disentanglement |
| **TGDG-SSA** | Neck (×3) | Language-guided multi-scale fusion of visual features with text embeddings |
| **OptiSARNetPlusPlusDetect** | Head | Region–text contrastive matching head with region-aware auxiliary supervision |
| MobileCLIP2-B (frozen) | Text encoder | Encodes referring expressions |

Key config of this checkpoint: `nc: 16` classes, `scale: m`, `reg_max: 16`.

## 📦 Files

| File | Description |
|---|---|
| `best.pt` | Trained weights (≈ 129 MB). The model architecture/modules are defined in the [GitHub repository](https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus). |

## 📈 Results — OptSAR-RSVG Test Split

Evaluated on the [OptSAR-RSVG](https://huggingface.co/datasets/JunDong-dev/OptSAR-RSVG) test split (4,434 images / 8,103 referring expressions). All values in %.

| Domain | Samples | Pr@0.5 | Pr@0.6 | Pr@0.7 | Pr@0.8 | Pr@0.9 | meanIoU | cumIoU |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| **All** | 8,103 | **93.61** | 93.19 | 91.94 | 87.75 | 66.43 | **85.94** | **92.11** |
| Optical | 6,027 | 93.01 | 92.58 | 91.67 | 88.58 | 72.99 | 86.48 | 92.25 |
| SAR | 2,076 | 95.33 | 94.94 | 92.73 | 85.31 | 47.40 | 84.38 | 83.21 |

Benchmark comparisons against TransVG, LQVG, TACMT, CSDNet, Grounding DINO, GLIP, etc. are reported in the [paper](https://arxiv.org/abs/2603.24876) and the [GitHub README](https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus).

## 🚀 Usage

`best.pt` contains **custom modules** (`PLoRA_MoE`, `TGDG_SSA`, `OptiSARNetPlusPlusDetect`), so it must be loaded with the project code rather than a stock Ultralytics install:

```bash
# 1. Clone the official repository (defines the custom modules & inference entry points)
git clone https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus.git
cd OptiSAR-Net-PlusPlus
pip install -r requirements.txt

# 2. Download the checkpoint and the dataset
hf download JunDong-dev/OptiSAR-Net-PlusPlus best.pt --local-dir .
hf download JunDong-dev/OptSAR-RSVG --repo-type dataset --local-dir OptSAR-RSVG

# 3. Run evaluation / inference with the scripts provided in the repository README
```

Evaluation metrics follow the standard CD-RSVG protocol: Pr@{0.5–0.9}, meanIoU, cumIoU, with per-domain (`optical` / `sar`) reporting.

## 🎯 Intended Use

- Cross-domain (optical ↔ SAR) referring expression comprehension / visual grounding in remote sensing
- Research on multi-modal fusion, parameter-efficient MoE adaptation, and contrastive region–text matching

## ⚖️ License

Code and weights are released under **AGPL-3.0**. The companion dataset is subject to its source datasets' licenses (see the dataset card).

## 📚 Citation

```bibtex
@article{tang2026optisar,
  title={OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding},
  author={Tang, Xiaoyu and Dong, Jun and Cheng, Jintao and Fan, Rui},
  journal={arXiv preprint arXiv:2603.24876},
  year={2026}
}
```