Instructions to use xcczach/xturnix-pt with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use xcczach/xturnix-pt with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="xcczach/xturnix-pt", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("xcczach/xturnix-pt", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("xcczach/xturnix-pt", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use xcczach/xturnix-pt with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "xcczach/xturnix-pt" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xcczach/xturnix-pt", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/xcczach/xturnix-pt
- SGLang
How to use xcczach/xturnix-pt with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "xcczach/xturnix-pt" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xcczach/xturnix-pt", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "xcczach/xturnix-pt" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xcczach/xturnix-pt", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use xcczach/xturnix-pt with Docker Model Runner:
docker model run hf.co/xcczach/xturnix-pt
English | ไธญๆ
XTurnix Pretrained (Qwen3 0.6B)
Predict turn-taking decisions from transcribed dialogue history: whether the AI should keep listening or start speaking while listening, and whether it should keep speaking or stop and listen when the user speaks while the AI is speaking.
The model maintains the AI's current state and makes a turn-taking decision whenever it receives user input:
| Current AI state | XTurnix decision | XTurnix output |
|---|---|---|
listening |
The user has not finished the current turn; keep listening | keep |
listening |
The user has finished the current turn; start responding | start |
speaking |
The user input does not require the AI to yield the current turn | keep |
speaking |
The user input requires the AI to stop speaking and listen | stop |
Quick Start
Transformers
from transformers import pipeline
pipe = pipeline(
model="xcczach/xturnix-pt",
trust_remote_code=True,
device=0,
dtype="auto",
)
result = pipe(
[
{
"role": "user",
"content": "Should the air purifier stay on continuously, or is it enough to run it for two hours before bed?",
}
],
state="listening",
)
print(result)
The inference pipeline automatically removes trailing punctuation from the last user message to preserve model performance.
Overlong dialogue history is truncated. The system prompt and the most recent dialogue messages are retained.
Example output:
{
"action": "<|start|>",
"scores": {
"<|start|>": 0.97,
"<|keep|>": 0.03,
},
}
The pipeline also accepts a string directly:
result = pipe("้ฃๆไปฌๆๅคฉๅ ็นๅบๅ๏ผ", state="listening")
Batch inference:
results = pipe(
[
{"messages": messages_a, "state": "listening"},
{"messages": messages_b, "state": "speaking"},
]
)
vLLM
Start the server:
vllm serve xcczach/xturnix-pt \
--served-model-name xturnix \
--host 0.0.0.0 \
--port 8000 \
--dtype auto \
--max-model-len 2048 \
--generation-config vllm
Use the included vllm_client.py as the reference client.
Complete Deployment
This model is supported by the X-Talk dialogue framework.
- Downloads last month
- 30