3D-PLOT-LLM-7B

Weights of 3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models (NeurIPS 2026).

3D-PLOT-LLM extends PointLLM (Vicuna-v1.5-7B with a frozen Point-BERT encoder on 8192-point Objaverse clouds). The encoder's patch tokens are partitioned into K=16 regions; each region is prefixed with a learnable marker and a reserved vocabulary token <part_k>, refined by a lightweight Marker-Space Refinement module. The model can answer which regions a part description refers to, describe a given set of regions, and caption whole objects, with fewer than one million added parameters.

Usage

git clone https://github.com/jintangxue/3D-PLOT-LLM && cd 3D-PLOT-LLM
source env.sh   # set POINTLLM_DATA, BCP_PACK_DIR, PARTVERSE_QA_DIR
python pointllm/eval/eval_partverse_caption2slots.py --model_name jintangx/3D-PLOT-LLM-7B \
    --anno_path $PARTVERSE_QA_DIR/eval_c2s.json \
    --data_path $POINTLLM_DATA/objaverse_data --bcp_pack_dir $BCP_PACK_DIR

The checkpoint is in PointLLM format (PointLLMLlamaForCausalLM, bf16) and loads with the pinned transformers commit used by the code. Each input object needs its partition pack (<object_id>.pack.npz), shipped with the dataset for the evaluation objects and generated by the repository tools otherwise. See the repository README for installation and the environment used for the paper.

Results

Benchmark Metric Score
PartVerse-QA caption-to-slots (392 queries) Jaccard / exact match 0.459 / 13.78%
PartVerse-QA slots-to-caption (196 queries) SBERT 62.0

Caption-to-slots uses greedy decoding and reproduces the paper's predictions exactly on an A100 with the paper's environment. Slots-to-caption is the mean over 5 sampled runs (temperature 1.0, top-p 0.95, top-k 50). See the paper for the full tables.

License

Non-commercial research use: the weights inherit the Llama 2 community license of the Vicuna backbone and the CC-BY-NC-4.0 license of the PointLLM initialization. The code is CC BY-NC-SA 4.0.

Citation

@article{xue20263d,
  title={3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models},
  author={Xue, Jintang and Wang, Xinyu and Wu, Yixing and Chen, Jingwen and Kuo, C-C Jay},
  journal={arXiv preprint arXiv:2606.19828},
  year={2026}
}

Accepted at NeurIPS 2026; the entry will be updated when the proceedings version is published.

Downloads last month
379
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jintangx/3D-PLOT-LLM-7B

Finetuned
(1)
this model

Dataset used to train jintangx/3D-PLOT-LLM-7B

Paper for jintangx/3D-PLOT-LLM-7B