Instructions to use Intel/neural-chat-7b-v3-1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Intel/neural-chat-7b-v3-1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Intel/neural-chat-7b-v3-1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Intel/neural-chat-7b-v3-1") model = AutoModelForCausalLM.from_pretrained("Intel/neural-chat-7b-v3-1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Intel/neural-chat-7b-v3-1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Intel/neural-chat-7b-v3-1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Intel/neural-chat-7b-v3-1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Intel/neural-chat-7b-v3-1
- SGLang
How to use Intel/neural-chat-7b-v3-1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Intel/neural-chat-7b-v3-1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Intel/neural-chat-7b-v3-1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Intel/neural-chat-7b-v3-1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Intel/neural-chat-7b-v3-1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Intel/neural-chat-7b-v3-1 with Docker Model Runner:
docker model run hf.co/Intel/neural-chat-7b-v3-1
About DROP results within the `lm-eval-harness`
Hi here! I'm curious about the huge gap w.r.t. Mistral in the DROP benchmark of the lm-eval-harness, did you use the same revision of EleutherAI/lm-eval-harness? Also did you run any other evaluation to see the reason why it excels that much at DROP compared to other SFT + DPO fine-tunes e.g. Zephyr? Is there any data contamination coming from the dataset used for training?
A bunch of questions π Feel free to answer in case you checked the issues with DROP, because the gap compared to other models seems huge and would be nice to investigate, maybe the data just has better quality!
I'm only chiming in so I will get a notice if someone answers this question because I'm also interested in why this LLM has a much higher DROP score than other SFT + DPO LLMs like Zephyr.
hi, the used datasets are listed at the model card. We find the metric of drop decreases during the training. So early stopping is needed.
True, I've submitted a PR at https://huggingface.co/Intel/neural-chat-7b-v3-1/discussions/15 to enrich the metadata within the Model Card in the README.md π€
Hmm did you evaluate the model during training using the lm-eval-harness? Was DROP within your evaluation set? Could you please elaborate more on that, I think it's a really interesting topic!