Instructions to use 01-ai/Yi-34B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use 01-ai/Yi-34B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="01-ai/Yi-34B", device_map="auto")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("01-ai/Yi-34B") model = AutoModelForCausalLM.from_pretrained("01-ai/Yi-34B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use 01-ai/Yi-34B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "01-ai/Yi-34B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "01-ai/Yi-34B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/01-ai/Yi-34B
- SGLang
How to use 01-ai/Yi-34B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "01-ai/Yi-34B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "01-ai/Yi-34B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "01-ai/Yi-34B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "01-ai/Yi-34B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use 01-ai/Yi-34B with Docker Model Runner:
docker model run hf.co/01-ai/Yi-34B
Could any 大佬 run a Yi-34B server plz?
Can't wait to try it.
+1
实话讲,英语不会你可以用大模型为你翻译,不懂得怎么运行模型你可以学会谷歌,做个伸手党是做不成的。
https://github.com/oobabooga/text-generation-webui
https://huggingface.co/TheBloke/Yi-34B-GGUF
素材给到你,自己自学谢谢。
"To be honest, if you don't know English, you can use a large model to translate for you. If you don't know how to run the model, you can learn from Google and being lazy is not an option."
"Here are some resources: https://github.com/oobabooga/text-generation-webui and https://huggingface.co/TheBloke/Yi-34B-GGUF .
Materials have been provided, please self-learn thank you."
实话讲,英语不会你可以用大模型为你翻译,不懂得怎么运行模型你可以学会谷歌,做个伸手党是做不成的。
https://github.com/oobabooga/text-generation-webui
https://huggingface.co/TheBloke/Yi-34B-GGUF
素材给到你,自己自学谢谢。"To be honest, if you don't know English, you can use a large model to translate for you. If you don't know how to run the model, you can learn from Google and being lazy is not an option."
"Here are some resources: https://github.com/oobabooga/text-generation-webui and https://huggingface.co/TheBloke/Yi-34B-GGUF .
Materials have been provided, please self-learn thank you."
我为什么需要自己运行一个模型,我没有运营和维护以及盈利的精力,个人使用需求仅有一个月百万Tokens,直接买一个大佬商用版API不可以?
Needless run a server by myself, cause I know which way is the best for myself.
I don't have either time or effort to run and maintain such an expensive host, let alone to make money from it.
Personally I only need millions tokens per month, so I can buy an API if any 大佬 can share.
🐮
The discussions here are not very constructive so I'm closing this now.
Once the chat model get released, we may consider cooperating with other platforms to make our models more accessible to users. Let us know if you have any recommendations.