File size: 4,178 Bytes
ac02d47 bc15982 ac02d47 bc15982 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 | ---
license: agpl-3.0
library_name: pytorch
pipeline_tag: object-detection
tags:
- remote-sensing
- visual-grounding
- sar
- optical
- cross-domain
- mixture-of-experts
- contrastive-learning
- benchmark
---
# OptiSAR-Net++ — Official Weights
Official trained weights (`best.pt`) of **OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding**.
OptiSAR-Net++ is a **transformer-free** framework for Cross-Domain Remote Sensing Visual Grounding (CD-RSVG): a single unified model localizes targets described by free-form natural language in **both optical and SAR** remote sensing imagery, replacing a heavy Transformer decoder with **contrastive region–text matching**.
- 🤗 **Dataset**: [JunDong-dev/OptSAR-RSVG](https://huggingface.co/datasets/JunDong-dev/OptSAR-RSVG)
- 💻 **Code**: [GitHub — JunDong-dev/OptiSAR-Net-PlusPlus](https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus)
## 🏗️ Architecture
The checkpoint follows a YOLOE-style single-stage detector extended with:
| Component | Location | Role |
|---|---|---|
| **PLoRA-MoE** | Backbone | Patch-level low-rank-adaptation Mixture-of-Experts for optical/SAR feature disentanglement |
| **TGDG-SSA** | Neck (×3) | Language-guided multi-scale fusion of visual features with text embeddings |
| **OptiSARNetPlusPlusDetect** | Head | Region–text contrastive matching head with region-aware auxiliary supervision |
| MobileCLIP2-B (frozen) | Text encoder | Encodes referring expressions |
Key config of this checkpoint: `nc: 16` classes, `scale: m`, `reg_max: 16`.
## 📦 Files
| File | Description |
|---|---|
| `best.pt` | Trained weights (≈ 129 MB). The model architecture/modules are defined in the [GitHub repository](https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus). |
## 📈 Results — OptSAR-RSVG Test Split
Evaluated on the [OptSAR-RSVG](https://huggingface.co/datasets/JunDong-dev/OptSAR-RSVG) test split (4,434 images / 8,103 referring expressions). All values in %.
| Domain | Samples | Pr@0.5 | Pr@0.6 | Pr@0.7 | Pr@0.8 | Pr@0.9 | meanIoU | cumIoU |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| **All** | 8,103 | **93.61** | 93.19 | 91.94 | 87.75 | 66.43 | **85.94** | **92.11** |
| Optical | 6,027 | 93.01 | 92.58 | 91.67 | 88.58 | 72.99 | 86.48 | 92.25 |
| SAR | 2,076 | 95.33 | 94.94 | 92.73 | 85.31 | 47.40 | 84.38 | 83.21 |
Benchmark comparisons against TransVG, LQVG, TACMT, CSDNet, Grounding DINO, GLIP, etc. are reported in the [paper](https://arxiv.org/abs/2603.24876) and the [GitHub README](https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus).
## 🚀 Usage
`best.pt` contains **custom modules** (`PLoRA_MoE`, `TGDG_SSA`, `OptiSARNetPlusPlusDetect`), so it must be loaded with the project code rather than a stock Ultralytics install:
```bash
# 1. Clone the official repository (defines the custom modules & inference entry points)
git clone https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus.git
cd OptiSAR-Net-PlusPlus
pip install -r requirements.txt
# 2. Download the checkpoint and the dataset
hf download JunDong-dev/OptiSAR-Net-PlusPlus best.pt --local-dir .
hf download JunDong-dev/OptSAR-RSVG --repo-type dataset --local-dir OptSAR-RSVG
# 3. Run evaluation / inference with the scripts provided in the repository README
```
Evaluation metrics follow the standard CD-RSVG protocol: Pr@{0.5–0.9}, meanIoU, cumIoU, with per-domain (`optical` / `sar`) reporting.
## 🎯 Intended Use
- Cross-domain (optical ↔ SAR) referring expression comprehension / visual grounding in remote sensing
- Research on multi-modal fusion, parameter-efficient MoE adaptation, and contrastive region–text matching
## ⚖️ License
Code and weights are released under **AGPL-3.0**. The companion dataset is subject to its source datasets' licenses (see the dataset card).
## 📚 Citation
```bibtex
@article{tang2026optisar,
title={OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding},
author={Tang, Xiaoyu and Dong, Jun and Cheng, Jintao and Fan, Rui},
journal={arXiv preprint arXiv:2603.24876},
year={2026}
}
```
|