godx7's picture
Update README.md
dc13e77 verified
|
Raw
History Blame Contribute Delete
1.78 kB
---
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- qwen3-vl
- samtok
- segmentation
- remote-code
---
# Qwen3-VL-4B-SAMTok Composite
This is a self-contained Hugging Face remote-code wrapper for Qwen3-VL-4B-SAMTok.
It loads both the Qwen3-VL language-vision model and the SAMTok VQ-SAM2 mask
decoder from this single model repository.
## Usage
```python
from transformers import AutoModel
model = AutoModel.from_pretrained(
"godx7/SAMTok-self-contained",
trust_remote_code=True,
torch_dtype="auto",
device_map="auto",
)
result = model.generate_with_masks(
image="example.jpg",
question=(
"Could you please give me a detailed description of the image? "
"Please respond with interleaved segmentation masks for the corresponding parts of the answer."
),
)
print(result["text"])
print(result["tags"])
print(result["masks"].shape if result["masks"] is not None else None)
```
For lower-level usage, the model forwards normal Qwen3-VL calls:
```python
processor = model.processor
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
generated_ids = model.generate(**inputs, max_new_tokens=512)
```
You can also decode SAMTok text tokens directly:
```python
masks, tags = model.decode_masks(output_text, image="example.jpg")
```
## Files
- `config.json`: composite remote-code config.
- `qwen_config.json`: original Qwen3-VL model config.
- `model-*.safetensors`: Qwen3-VL weights.
- `mask_tokenizer_256x2.pth`: SAMTok VQ decoder weights.
- `sam2.1_hiera_large.pt`: SAM2 image encoder checkpoint.
- `modeling_samtok.py`, `configuration_samtok.py`, `samtok/`: remote-code implementation.