Instructions to use maywell/Yi-34B-Undertrained with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use maywell/Yi-34B-Undertrained with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="maywell/Yi-34B-Undertrained") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("maywell/Yi-34B-Undertrained") model = AutoModelForCausalLM.from_pretrained("maywell/Yi-34B-Undertrained", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use maywell/Yi-34B-Undertrained with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "maywell/Yi-34B-Undertrained" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maywell/Yi-34B-Undertrained", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/maywell/Yi-34B-Undertrained
- SGLang
How to use maywell/Yi-34B-Undertrained with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "maywell/Yi-34B-Undertrained" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maywell/Yi-34B-Undertrained", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "maywell/Yi-34B-Undertrained" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maywell/Yi-34B-Undertrained", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use maywell/Yi-34B-Undertrained with Docker Model Runner:
docker model run hf.co/maywell/Yi-34B-Undertrained
The model is a version of the SFT that trained on purpose to do DPO on Yi model.
I was unable to resolve the OOM issue while trying to train DPO, so I am only uploading the SFT.
If you would like to DPO on that model, please use the maywell/why_no_one_do_dpo_on_yi dataset.
It follows prompt format ChatML.
Below code used to load maywell/why_no_one_do_dpo_on_yi dataset on axolotl.
class SimpleShareGPTPromptTokenizingStrategy(ShareGPTPromptTokenizingStrategy):
_strict = True
@property
def strict(self):
return self._strict
@strict.setter
def strict(self, strict):
self._strict = strict
def get_conversation_thread(self, prompt):
conversations = prompt['chosen']
turns = [{"from": "assistant" if t["role"] == "assistant" else t["role"], "value": t["content"]} for t in conversations]
return turns
ํด๋น ๋ชจ๋ธ์ Yi ๋ชจ๋ธ์ DPOํ๊ธฐ ์ํด ํ๋ จ์์ผฐ๋ SFT ๋ฒ์ ์ ๋๋ค.
DPO ํ๋ จ์ ํ๋ ค๋ ์ค OOM ๋ฌธ์ ๋ฅผ ํด๊ฒฐํ์ง ๋ชปํ์ฌ SFT๋ง ์ ๋ก๋ํฉ๋๋ค.
ํด๋น ๋ชจ๋ธ์ DPO๋ฅผ ํ์๋ ค๋ฉด maywell/why_no_one_do_dpo_on_yi ๋ฐ์ดํฐ์ ์ ์ด์ฉํด์ฃผ์ธ์.
ํ๋กฌํํธ ํฌ๋งท์ ChatML์ ๋ฐ๋ฆ ๋๋ค.
- Downloads last month
- 5