Instructions to use upstage/SOLAR-10.7B-Instruct-v1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use upstage/SOLAR-10.7B-Instruct-v1.0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="upstage/SOLAR-10.7B-Instruct-v1.0") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("upstage/SOLAR-10.7B-Instruct-v1.0") model = AutoModelForCausalLM.from_pretrained("upstage/SOLAR-10.7B-Instruct-v1.0", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use upstage/SOLAR-10.7B-Instruct-v1.0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "upstage/SOLAR-10.7B-Instruct-v1.0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/SOLAR-10.7B-Instruct-v1.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/upstage/SOLAR-10.7B-Instruct-v1.0
- SGLang
How to use upstage/SOLAR-10.7B-Instruct-v1.0 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "upstage/SOLAR-10.7B-Instruct-v1.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/SOLAR-10.7B-Instruct-v1.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "upstage/SOLAR-10.7B-Instruct-v1.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/SOLAR-10.7B-Instruct-v1.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use upstage/SOLAR-10.7B-Instruct-v1.0 with Docker Model Runner:
docker model run hf.co/upstage/SOLAR-10.7B-Instruct-v1.0
You know Mixtral, Llama 2 70b, GPT3.5... Are All Much Better
I used Solar instruct for hours across days, and while it scored slightly higher than Mistral 7bs in my testing, it didn't score near as high as Mixtrals, Llama 2 70bs or GPT3.5.
Usually the results of my testing roughly align with HF average scores, but since they were way off I looked into it.
It appears the discrepancy is primary due to 2 things.
(1) Solar instruct obsessively denies things are true, including countless millions of things which are in fact true, resulting in an absurdly high 71.5 TruthfulQA score (much higher than even GPT4). When I removed TruthfulQA from the HF average score it was a much better representation of Solar Instruct's true performance.
(2) Solar instruct gives unusually brief responses, even when contraindicated by the circumstances or the user's instruction. And because of automated eval limitations longer and more complex answers result in lower scores on numerous tests (more true answers falsely identified as false).
All things considered, the true HF score of Solar Instruct is no higher than 68, and certainly nowhere near 74. 74 would put it above GPT3.5, yet it isn't near as good (not my opinion). It's not even near as good as Llama 2 70b or Mixtral. It has far less knowledge and gets tripped up by much simpler questions than all three, yet has a higher score.
Thank you for the much needed analysis. Especially the point about TruthfulQA score was illuminating