Instructions to use upstage/SOLAR-0-70b-16bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use upstage/SOLAR-0-70b-16bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="upstage/SOLAR-0-70b-16bit")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("upstage/SOLAR-0-70b-16bit") model = AutoModelForCausalLM.from_pretrained("upstage/SOLAR-0-70b-16bit", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use upstage/SOLAR-0-70b-16bit with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "upstage/SOLAR-0-70b-16bit" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/SOLAR-0-70b-16bit", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/upstage/SOLAR-0-70b-16bit
- SGLang
How to use upstage/SOLAR-0-70b-16bit with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "upstage/SOLAR-0-70b-16bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/SOLAR-0-70b-16bit", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "upstage/SOLAR-0-70b-16bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/SOLAR-0-70b-16bit", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use upstage/SOLAR-0-70b-16bit with Docker Model Runner:
docker model run hf.co/upstage/SOLAR-0-70b-16bit
Call w/ LiteLLM
Hi @hunkim / @yoonniverse
What's the best way for me to deploy this model? I'd love to make a demo of this with LiteLLM - https://github.com/BerriAI/litellm.
Lite currently works with Replicate, Azure, Together.ai and HF Inference Endpoints.
I'm facing issues with HF Inference endpoints due to quota limitations, so curious if you've tried any other provider.
We will soon host our model on Together.ai. We will keep you updated.
Do you know how to integrate our model with https://github.com/BerriAI/litellm? We will make it work. Let us know.
Hey @hunkim we made it easy to proxy openai with any deployment solution - should unlock any provider you choose. - https://github.com/BerriAI/litellm/issues/120
import litellm
def translate_function(model, messages, max_tokens):
prompt = " ".join(message["content"] for message in messages)
max_new_tokens = max_tokens
return {"model": model, "prompt": prompt, "max_new_tokens": max_new_tokens}
openai.api_base = litellm.translate_api_call(custom_api_base, translate_function)
We already have a custom integration with together.ai, which supports streaming. Excited to put out a demo notebook/etc. once it's deployed.