| --- |
| license: apache-2.0 |
| library_name: pytorch |
| pipeline_tag: image-segmentation |
| tags: |
| - affordance-segmentation |
| - instance-segmentation |
| - object-detection |
| - robotics |
| - manipulation |
| - rt-detr |
| - real-time |
| metrics: |
| - name: Weighted F-measure (beta^2=1) |
| type: f-measure |
| value: 0.8685 |
| --- |
| |
| # AffKernel: IIT-AFF affordance segmentation (R50vd, stride-2 + deep supervision) |
|
|
| Single-pass, NMS-free affordance segmentation. AffKernel couples RT-DETR object |
| queries with a CondInst-style per-query dynamic-convolution kernel that decodes |
| per-object affordance masks from one shared high-resolution affordance map. A |
| detector tells a robot *what* an object is; AffKernel also tells it *where* to |
| act on it, at a latency a control loop can afford. |
|
|
| - **Code:** https://github.com/anh0001/affkernel |
| - **Reproduction guide:** https://github.com/anh0001/affkernel/blob/main/docs/reproduction.md |
| - **Architecture:** RT-DETR-R50vd + affordance branch with stride-2 readout and deep supervision |
| - **Training data:** IIT-AFF (Nguyen et al., IROS 2017), 6,184 train images |
| - **Weights:** last-epoch (epoch 72) EMA parameters, fp32, 174 MB |
|
|
| ## Results |
|
|
| Evaluated on the IIT-AFF test split (2,651 images) with the Margolin weighted |
| F-measure. |
|
|
| | Metric | Value | |
| |---|---:| |
| | `F_beta^w` (beta^2 = 1), this checkpoint | **0.8685** | |
| | `F_beta^w` (beta^2 = 0.3), this checkpoint | 0.8582 | |
| | `F_beta^w` (beta^2 = 1), mean of 3 seeds | 0.8675 Β± 0.0009 | |
| | Mask quality on detected instances | 0.8933 | |
| | Latency, 640x640, fp32, RTX 6000 Ada | 16.3 ms median | |
| | Throughput | 61.5 img/s | |
|
|
| Against published baselines on the same benchmark: Mask R-CNN 0.844 at 45 ms, |
| deterministic Swin-T 0.883 at 42 ms, Bayesian Swin-T deep ensemble 0.906 at |
| roughly 1015 ms. AffKernel beats the Mask R-CNN baseline at about a third of its |
| latency and trails the deterministic Swin-T by 1.55 points at 39% of its |
| latency, with fully deterministic single-pass inference. |
|
|
| **On seed selection.** This is seed 42, which is both the primary anchor seed |
| used throughout the paper and the highest scoring of the three seeds trained |
| (seed 7: 0.8668, seed 123: 0.8673). IIT-AFF provides no validation split, so |
| the epoch was fixed a priori (last epoch, EMA weights) rather than selected on |
| the test set. |
|
|
| **On the beta convention.** Two conventions circulate in this literature. The |
| AffordanceNet lineage reports `beta^2 = 0.3`; recent transformer baselines |
| report `beta^2 = 1`. Both are given above so that comparisons can be made at a |
| matched convention. Do not compare a `beta^2 = 1` number against a |
| `beta^2 = 0.3` number. |
|
|
| ## Usage |
|
|
| ```bash |
| git clone https://github.com/anh0001/affkernel.git |
| cd affkernel |
| pip install -r requirements.txt |
| |
| pip install huggingface_hub |
| hf download anhrisn/affkernel-iit-aff \ |
| affkernel_iit_r50vd_stride2_deepsup_seed42.pth --local-dir weights/ |
| |
| python tools/infer.py \ |
| -c configs/rtdetr/rtdetr_r50vd_6x_iit_v3_stride2_deepsup.yml \ |
| -r weights/affkernel_iit_r50vd_stride2_deepsup_seed42.pth \ |
| --input path/to/image.jpg \ |
| --output outputs/ |
| ``` |
|
|
| `--input` takes a single image or a directory. `--output` is a **directory**; |
| results are written as `<stem>_aff.png`. Add `--device cpu` to run without a |
| GPU. Verified working on GPU and CPU, for both single-image and folder input. |
|
|
| Reproduce the reported metric (requires the IIT-AFF dataset; see |
| [`docs/datasets.md`](https://github.com/anh0001/affkernel/blob/main/docs/datasets.md)): |
|
|
| ```bash |
| python tools/decompose_fbw_gap.py \ |
| -c configs/rtdetr/rtdetr_r50vd_6x_iit_v3_stride2_deepsup.yml \ |
| -r weights/affkernel_iit_r50vd_stride2_deepsup_seed42.pth --beta2 1.0 |
| ``` |
|
|
| ### Checkpoint format |
|
|
| A single file containing the EMA weights only: |
|
|
| ```python |
| {"ema": {"module": <OrderedDict of 760 tensors>, "updates": 111312}} |
| ``` |
|
|
| The evaluation path uses the EMA parameters, so the raw (non-EMA) weights and |
| the optimizer state from training are not included. Do not repack this into a |
| bare `{"model": ...}` dict: under an EMA-enabled config the solver would then |
| evaluate a freshly initialised EMA module and score near zero. |
|
|
| The first run downloads ImageNet-pretrained ResNet-50vd backbone weights from |
| the RT-DETR release artefacts, so it needs network access. For an offline |
| deployment, set `PResNet.pretrained: False` in the config. |
|
|
| ## Deployment |
|
|
| All rows below use **these same weights** β fp16, CUDA graphs, BatchNorm folding |
| and the TensorRT backbone are runtime configurations, not separate checkpoints. |
|
|
| **RTX 6000 Ada**, 640x640, batch 1, end to end (preprocess, forward, mask decode, |
| device-to-host copy): |
|
|
| | Inference path | End to end | FPS | Peak GPU | `F_beta^w` (0.3 / 1) | |
| |---|---:|---:|---:|---:| |
| | fp32, as released | 35.7 ms | 28.0 | 0.43 GiB | 0.8582 / 0.8685 | |
| | fp16 + CUDA graph + BN fold | **16.0 ms** | **62.7** | 0.29 GiB | 0.8577 / 0.8680 | |
|
|
| **Jetson AGX Orin** Developer Kit (64 GB), JetPack 6.2, MAXN: |
|
|
| | Inference path | End to end | FPS | Peak GPU | `F_beta^w` (0.3 / 1) | |
| |---|---:|---:|---:|---:| |
| | fp32, as released | 199.5 ms | 5.0 | 1.29 GiB | 0.8598 / 0.8680 | |
| | `--half --cudagraph` | 53.2 ms | 18.8 | 0.21 GiB | 0.8596 / 0.8678 | |
| | `--trt-backbone` | **43.4 ms** | **23.1** | 0.23 GiB | 0.8591 / 0.8673 | |
|
|
| The fastest path costs **0.0007** `F_beta^w` at beta^2=1 β below the Β±0.0009 |
| seed-to-seed spread of the training run itself. Dropping the input to 512x512 is |
| *not* worth it by comparison: it costs 0.007, ten times more, to save less time. |
|
|
| ### `backbone_fp16.plan` β prebuilt TensorRT engine |
| |
| This repository also ships an fp16 TensorRT engine for the **backbone only**. The |
| dynamic-kernel affordance head cannot be exported to ONNX, but the backbone is |
| plain convolutions, exports cleanly, and is the largest single term of the forward |
| pass: 17.7 ms β 5.3 ms. The encoder, decoder and affordance head stay in PyTorch, |
| CUDA-graphed against the engine's output buffers. |
| |
| | | | |
| |---|---| |
| | Built on | Jetson AGX Orin (64 GB), JetPack 6.2 (L4T R36.4.7) | |
| | TensorRT / CUDA | 10.3.0 / 12.6 | |
| | Input | `images` β `(1, 3, 640, 640)`, fp16, fixed | |
| | Outputs | `feat0..feat3` β fp16 | |
| | Size | 46.4 MiB | |
| |
| > **Engine files are not portable.** A `.plan` is compiled for one specific GPU, |
| > TensorRT version and input shape. On any other device, JetPack/TensorRT version |
| > or input size it will fail to deserialize β or load and behave incorrectly. It |
| > is published here as a convenience for setups identical to the table above. |
| > **If yours differs in any respect, build your own** β one command, a few |
| > minutes. That is the supported path. |
| |
| ```bash |
| # use the prebuilt engine (identical setups only) |
| hf download anhrisn/affkernel-iit-aff backbone_fp16.plan --local-dir weights/ |
|
|
| # or build your own (recommended) |
| python tools/build_trt_backbone.py \ |
| -c configs/rtdetr/rtdetr_r50vd_6x_iit_v3_stride2_deepsup.yml \ |
| -r weights/affkernel_iit_r50vd_stride2_deepsup_seed42.pth \ |
| --out weights/backbone_fp16.plan |
|
|
| python tools/infer.py \ |
| -c configs/rtdetr/rtdetr_r50vd_6x_iit_v3_stride2_deepsup.yml \ |
| -r weights/affkernel_iit_r50vd_stride2_deepsup_seed42.pth \ |
| --input path/to/image.jpg --output outputs/ \ |
| --trt-backbone weights/backbone_fp16.plan --gpu-preprocess |
| ``` |
| |
| Keep the checkpoint either way β the engine replaces the backbone at inference |
| only, and training is unaffected. |
| |
| ## Files |
| |
| | File | What it is | |
| |---|---| |
| | `affkernel_iit_r50vd_stride2_deepsup_seed42.pth` | The model. fp32 EMA weights, 174 MB. | |
| | `backbone_fp16.plan` | Optional deployment artifact. TensorRT fp16 backbone engine for AGX Orin / JetPack 6.2 / TensorRT 10.3 @ 640x640. Not portable; not usable on its own. | |
| |
| ## Classes |
| |
| **Objects (10):** bowl, tvm, pan, hammer, knife, cup, drill, racket, spatula, bottle |
| |
| **Affordances (9):** contain, cut, display, engine, grasp, hit, pound, support, w-grasp |
| |
| ## Intended use and limitations |
| |
| Intended for research on affordance perception and perception-guided grasping. |
| |
| - Trained on a closed vocabulary of 10 object and 9 affordance classes; it will |
| not generalise to unseen categories. |
| - IIT-AFF is a tabletop dataset. Performance under other viewpoints, lighting |
| or clutter regimes is untested. |
| - The residual error is dominated by **missed detections**, not by mask quality. |
| On instances the detector does fire on, mask quality (0.893) already exceeds |
| the deterministic Swin-T baseline's overall score. |
| - The headline latency is fp32 on an RTX 6000 Ada. fp16, CUDA graphs and a |
| TensorRT backbone are now characterised on both a workstation and a Jetson AGX |
| Orin β see [Deployment](#deployment) β and cost at most 0.0007 `F_beta^w`. |
| - Not validated for safety-critical deployment. A predicted grasp region is a |
| perception cue, not a guarantee of a safe grasp. |
| |
| ## License and attribution |
| |
| These weights are released under Apache-2.0, matching the licence of the |
| RT-DETR components and the ImageNet-pretrained backbone they were initialised |
| from. The AffKernel repository's own source contributions are MIT licensed; see |
| [`THIRD_PARTY_LICENSES.md`](https://github.com/anh0001/affkernel/blob/main/THIRD_PARTY_LICENSES.md). |
| |
| The **IIT-AFF dataset is not redistributed** here or in the code repository. It |
| states no licence; its authors request citation of the original paper. Obtain |
| it from https://sites.google.com/site/iitaffdataset/. |
| |
| ## Citation |
| |
| ```bibtex |
| @article{risnumawan2026affkernel, |
| title = {AffKernel: Single-Pass Affordance Segmentation with Per-Query |
| Dynamic Convolution for Real-Time Robotic Manipulation}, |
| author = {Risnumawan, Anhar}, |
| journal = {IEEE Access}, |
| year = {2026}, |
| note = {Under review} |
| } |
| |
| @inproceedings{nguyen2017object, |
| title = {Object-Based Affordances Detection with Convolutional Neural |
| Networks and Dense Conditional Random Fields}, |
| author = {Nguyen, Anh and Kanoulas, Dimitrios and Caldwell, Darwin G. and |
| Tsagarakis, Nikos G.}, |
| booktitle = {IEEE/RSJ International Conference on Intelligent Robots and |
| Systems (IROS)}, |
| year = {2017} |
| } |
| |
| @inproceedings{lv2024detrs, |
| title = {DETRs Beat YOLOs on Real-time Object Detection}, |
| author = {Lv, Wenyu and Zhao, Yian and Xu, Shangliang and Wei, Jinman and |
| Wang, Guanzhong and Cui, Cheng and Du, Yuning and Dang, Qingqing |
| and Liu, Yi}, |
| booktitle = {CVPR}, |
| year = {2024} |
| } |
| ``` |
| |