|
Download README.md from JunDong-dev/OptiSAR-Net-PlusPlus: direct link, hf CLI and curl.
- Browser
- Download file 4.18 kB
-
https://huggingface.co/JunDong-dev/OptiSAR-Net-PlusPlus/resolve/main/README.md
- Command line
-
hf download hf://JunDong-dev/OptiSAR-Net-PlusPlus/README.md
-
curl -L -o README.md https://huggingface.co/JunDong-dev/OptiSAR-Net-PlusPlus/resolve/main/README.md
4.18 kB
| license: agpl-3.0 | |
| library_name: pytorch | |
| pipeline_tag: object-detection | |
| tags: | |
| - remote-sensing | |
| - visual-grounding | |
| - sar | |
| - optical | |
| - cross-domain | |
| - mixture-of-experts | |
| - contrastive-learning | |
| - benchmark | |
| # OptiSAR-Net++ — Official Weights | |
| Official trained weights (`best.pt`) of **OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding**. | |
| OptiSAR-Net++ is a **transformer-free** framework for Cross-Domain Remote Sensing Visual Grounding (CD-RSVG): a single unified model localizes targets described by free-form natural language in **both optical and SAR** remote sensing imagery, replacing a heavy Transformer decoder with **contrastive region–text matching**. | |
| - 🤗 **Dataset**: [JunDong-dev/OptSAR-RSVG](https://huggingface.co/datasets/JunDong-dev/OptSAR-RSVG) | |
| - 💻 **Code**: [GitHub — JunDong-dev/OptiSAR-Net-PlusPlus](https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus) | |
| ## 🏗️ Architecture | |
| The checkpoint follows a YOLOE-style single-stage detector extended with: | |
| | Component | Location | Role | | |
| |---|---|---| | |
| | **PLoRA-MoE** | Backbone | Patch-level low-rank-adaptation Mixture-of-Experts for optical/SAR feature disentanglement | | |
| | **TGDG-SSA** | Neck (×3) | Language-guided multi-scale fusion of visual features with text embeddings | | |
| | **OptiSARNetPlusPlusDetect** | Head | Region–text contrastive matching head with region-aware auxiliary supervision | | |
| | MobileCLIP2-B (frozen) | Text encoder | Encodes referring expressions | | |
| Key config of this checkpoint: `nc: 16` classes, `scale: m`, `reg_max: 16`. | |
| ## 📦 Files | |
| | File | Description | | |
| |---|---| | |
| | `best.pt` | Trained weights (≈ 129 MB). The model architecture/modules are defined in the [GitHub repository](https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus). | | |
| ## 📈 Results — OptSAR-RSVG Test Split | |
| Evaluated on the [OptSAR-RSVG](https://huggingface.co/datasets/JunDong-dev/OptSAR-RSVG) test split (4,434 images / 8,103 referring expressions). All values in %. | |
| | Domain | Samples | Pr@0.5 | Pr@0.6 | Pr@0.7 | Pr@0.8 | Pr@0.9 | meanIoU | cumIoU | | |
| |---|---:|---:|---:|---:|---:|---:|---:|---:| | |
| | **All** | 8,103 | **93.61** | 93.19 | 91.94 | 87.75 | 66.43 | **85.94** | **92.11** | | |
| | Optical | 6,027 | 93.01 | 92.58 | 91.67 | 88.58 | 72.99 | 86.48 | 92.25 | | |
| | SAR | 2,076 | 95.33 | 94.94 | 92.73 | 85.31 | 47.40 | 84.38 | 83.21 | | |
| Benchmark comparisons against TransVG, LQVG, TACMT, CSDNet, Grounding DINO, GLIP, etc. are reported in the [paper](https://arxiv.org/abs/2603.24876) and the [GitHub README](https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus). | |
| ## 🚀 Usage | |
| `best.pt` contains **custom modules** (`PLoRA_MoE`, `TGDG_SSA`, `OptiSARNetPlusPlusDetect`), so it must be loaded with the project code rather than a stock Ultralytics install: | |
| ```bash | |
| # 1. Clone the official repository (defines the custom modules & inference entry points) | |
| git clone https://github.com/JunDong-dev/OptiSAR-Net-PlusPlus.git | |
| cd OptiSAR-Net-PlusPlus | |
| pip install -r requirements.txt | |
| # 2. Download the checkpoint and the dataset | |
| hf download JunDong-dev/OptiSAR-Net-PlusPlus best.pt --local-dir . | |
| hf download JunDong-dev/OptSAR-RSVG --repo-type dataset --local-dir OptSAR-RSVG | |
| # 3. Run evaluation / inference with the scripts provided in the repository README | |
| ``` | |
| Evaluation metrics follow the standard CD-RSVG protocol: Pr@{0.5–0.9}, meanIoU, cumIoU, with per-domain (`optical` / `sar`) reporting. | |
| ## 🎯 Intended Use | |
| - Cross-domain (optical ↔ SAR) referring expression comprehension / visual grounding in remote sensing | |
| - Research on multi-modal fusion, parameter-efficient MoE adaptation, and contrastive region–text matching | |
| ## ⚖️ License | |
| Code and weights are released under **AGPL-3.0**. The companion dataset is subject to its source datasets' licenses (see the dataset card). | |
| ## 📚 Citation | |
| ```bibtex | |
| @article{tang2026optisar, | |
| title={OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding}, | |
| author={Tang, Xiaoyu and Dong, Jun and Cheng, Jintao and Fan, Rui}, | |
| journal={arXiv preprint arXiv:2603.24876}, | |
| year={2026} | |
| } | |
| ``` | |