Image-Text-to-Text
Transformers
Safetensors
samtok_composite
feature-extraction
qwen3-vl
samtok
segmentation
remote-code
conversational
custom_code
Instructions to use godx7/SAMTok-self-contained with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use godx7/SAMTok-self-contained with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="godx7/SAMTok-self-contained", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("godx7/SAMTok-self-contained", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use godx7/SAMTok-self-contained with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "godx7/SAMTok-self-contained" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "godx7/SAMTok-self-contained", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/godx7/SAMTok-self-contained
- SGLang
How to use godx7/SAMTok-self-contained with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "godx7/SAMTok-self-contained" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "godx7/SAMTok-self-contained", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "godx7/SAMTok-self-contained" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "godx7/SAMTok-self-contained", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use godx7/SAMTok-self-contained with Docker Model Runner:
docker model run hf.co/godx7/SAMTok-self-contained
File size: 1,779 Bytes
3805eff dc13e77 3805eff | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 | ---
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- qwen3-vl
- samtok
- segmentation
- remote-code
---
# Qwen3-VL-4B-SAMTok Composite
This is a self-contained Hugging Face remote-code wrapper for Qwen3-VL-4B-SAMTok.
It loads both the Qwen3-VL language-vision model and the SAMTok VQ-SAM2 mask
decoder from this single model repository.
## Usage
```python
from transformers import AutoModel
model = AutoModel.from_pretrained(
"godx7/SAMTok-self-contained",
trust_remote_code=True,
torch_dtype="auto",
device_map="auto",
)
result = model.generate_with_masks(
image="example.jpg",
question=(
"Could you please give me a detailed description of the image? "
"Please respond with interleaved segmentation masks for the corresponding parts of the answer."
),
)
print(result["text"])
print(result["tags"])
print(result["masks"].shape if result["masks"] is not None else None)
```
For lower-level usage, the model forwards normal Qwen3-VL calls:
```python
processor = model.processor
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
generated_ids = model.generate(**inputs, max_new_tokens=512)
```
You can also decode SAMTok text tokens directly:
```python
masks, tags = model.decode_masks(output_text, image="example.jpg")
```
## Files
- `config.json`: composite remote-code config.
- `qwen_config.json`: original Qwen3-VL model config.
- `model-*.safetensors`: Qwen3-VL weights.
- `mask_tokenizer_256x2.pth`: SAMTok VQ decoder weights.
- `sam2.1_hiera_large.pt`: SAM2 image encoder checkpoint.
- `modeling_samtok.py`, `configuration_samtok.py`, `samtok/`: remote-code implementation.
|