File size: 2,550 Bytes
460d511
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d4ed48a
 
460d511
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4fe7dd3
 
 
 
 
 
 
 
460d511
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
---
license: cc-by-nc-4.0
library_name: torch-pointcloud
tags:
- point-cloud
- 3d
- pytorch
- torch-pointcloud
- concerto
- self-supervised
---

# Model card for concerto-large.pretrain.pointcept

A Concerto self-supervised pretraining model (joint 2D-3D representation encoder).

> **Non-commercial.** These weights are released by [Pointcept/Concerto](https://github.com/Pointcept/Concerto) under CC BY-NC 4.0 and may be used for research and evaluation only.

## Model Details

- **Model Type:** Self-supervised pretraining
- **Model Stats:**
  - Params (M): 207.7
  - Input channels: 9
- **Paper:** [Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations](https://arxiv.org/abs/2510.23607)
- **Converted from:** [Pointcept/Concerto](https://github.com/Pointcept/Concerto) (CC-BY-NC-4.0)
- **Library:** [torch-pointcloud](https://github.com/arthurdjn/pytorch-pointcloud)

## Install

```bash
pip install torch-pointcloud
```

This checkpoint also needs `spconv` and `flash-attn`, which need a build matching your torch and CUDA: see the [installation guide](https://pytorch-pointcloud.org/installation/).

## Usage

```python
import torch
import torch_pointcloud as tp
from torch_pointcloud.utils.data import collate

model, info = tp.create_model(
    "concerto-large.pretrain.pointcept",
    task="base",
    pretrained=True,
    return_info=True,
)
model = model.cuda().eval()  # GPU-only kernels

# synthetic sample with the keys a dataset provides
num_points = 8192
sample = {
    "pos": torch.randn(num_points, 3),
    "color": torch.rand(num_points, 3) * 255,
    "normal": torch.randn(num_points, 3),
    "segment": torch.zeros(num_points, dtype=torch.long),
    "instance": torch.zeros(num_points, dtype=torch.long),
}
data = info["transform"](sample)
data = collate([data])
data = {key: value.cuda() for key, value in data.items()}

with torch.no_grad():
    out = model(data.get("x"), data["pos_grid"], data["batch"], pos=data["pos"])
```

## Citation

```bibtex
@article{concerto2025,
  title   = {Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations},
  author  = {Yujia Zhang and Xiaoyang Wu and Yixing Lao and Chengyao Wang and Zhuotao Tian and Naiyan Wang and Hengshuang Zhao},
  journal = {arXiv preprint arXiv:2510.23607},
  year    = {2025}
}

@software{dujardin2026pytorchpointcloud,
  author  = {Arthur Dujardin},
  title   = {PyTorch PointCloud},
  year    = {2026},
  doi     = {10.5281/zenodo.22159632},
  url     = {https://github.com/arthurdjn/pytorch-pointcloud},
}
```