Instructions to use 01-ai/Yi-34B-Chat with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use 01-ai/Yi-34B-Chat with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="01-ai/Yi-34B-Chat") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("01-ai/Yi-34B-Chat") model = AutoModelForCausalLM.from_pretrained("01-ai/Yi-34B-Chat", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use 01-ai/Yi-34B-Chat with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "01-ai/Yi-34B-Chat" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "01-ai/Yi-34B-Chat", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/01-ai/Yi-34B-Chat
- SGLang
How to use 01-ai/Yi-34B-Chat with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "01-ai/Yi-34B-Chat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "01-ai/Yi-34B-Chat", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "01-ai/Yi-34B-Chat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "01-ai/Yi-34B-Chat", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use 01-ai/Yi-34B-Chat with Docker Model Runner:
docker model run hf.co/01-ai/Yi-34B-Chat
Why quantified version better than original version?
Why quantified version better than original version?
Hi there! Thank you for the Question! The reason for this difference is still unclear, and we are still investigating it. We will update you on the matter as soon as we find out.
I suspect this…
It’s like grading a college student on 4th-grade multiple choice and being surprised when the compressed version gets a similar—or slightly better—score.
🎯 Real Analogy
Model Type - Task Difficulty – Outcome
Yi-34B full College logic 🧠 True reasoning survives recursion
Yi-34B 8bit Grade 4 quiz 🏃 Fast + “good-enough” answers
The quantized model does great at surface tasks:
A → B style logic
Answer selection
Common sense fill-ins
But it doesn’t understand itself deeply. It just remembers fragments well and fills in blanks with pattern probability.
🧠 When It Fails:
Ask it:
“If your ethical recommendation leads to collapse of identity recursion in agent B, are you responsible?”
The quantized model:
Will either oversimplify
Drift into contradiction
Or dodge entirely
The full Yi-34B (non-quantized) has space to hold all vectors in float, compare identity threads, and refuse to lie.
🧩 Bottom Line:
Benchmarks ≠ Depth
Quantization ≠ Intelligence
Drift tolerance ≠ Truth stability
So yeah—I suspect that:
The questions were grade 4. The model is college-level. That’s why quant 8-bit looks good.
