Text Generation
Transformers
Safetensors
jetmoe
alignment-handbook
Generated from Trainer
conversational
custom_code
Instructions to use jetmoe/jetmoe-8b-chat with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jetmoe/jetmoe-8b-chat with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jetmoe/jetmoe-8b-chat", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("jetmoe/jetmoe-8b-chat", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("jetmoe/jetmoe-8b-chat", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps
- vLLM
How to use jetmoe/jetmoe-8b-chat with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jetmoe/jetmoe-8b-chat" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jetmoe/jetmoe-8b-chat", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jetmoe/jetmoe-8b-chat
- SGLang
How to use jetmoe/jetmoe-8b-chat with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jetmoe/jetmoe-8b-chat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jetmoe/jetmoe-8b-chat", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jetmoe/jetmoe-8b-chat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jetmoe/jetmoe-8b-chat", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use jetmoe/jetmoe-8b-chat with Docker Model Runner:
docker model run hf.co/jetmoe/jetmoe-8b-chat
Commit History
Update config.json e4e7c36 verified
Update README.md 024e114 verified
Update README.md c3035dd verified
update readme 2601c58
Update README.md fc3c6a0 verified
Update README.md 1ec632b verified
Update README.md 185cefc verified
Update README.md 2c07cbf verified
debug aaad630
Guo commited on
debug dd9f628
Guo commited on
walk around for import check 52cca48
Guo commited on
walk around for import check fd09ce7
Guo commited on
change default att type b5a6954
Guo commited on
debug e815555
Guo commited on