Instructions to use ChocoWu/nextgpt_7b_tiva_v0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ChocoWu/nextgpt_7b_tiva_v0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ChocoWu/nextgpt_7b_tiva_v0")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ChocoWu/nextgpt_7b_tiva_v0") model = AutoModelForCausalLM.from_pretrained("ChocoWu/nextgpt_7b_tiva_v0", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ChocoWu/nextgpt_7b_tiva_v0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ChocoWu/nextgpt_7b_tiva_v0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChocoWu/nextgpt_7b_tiva_v0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/ChocoWu/nextgpt_7b_tiva_v0
- SGLang
How to use ChocoWu/nextgpt_7b_tiva_v0 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ChocoWu/nextgpt_7b_tiva_v0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChocoWu/nextgpt_7b_tiva_v0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ChocoWu/nextgpt_7b_tiva_v0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChocoWu/nextgpt_7b_tiva_v0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use ChocoWu/nextgpt_7b_tiva_v0 with Docker Model Runner:
docker model run hf.co/ChocoWu/nextgpt_7b_tiva_v0
Not able to load the model using transformers
Could not load model ChocoWu/nextgpt_7b_tiva_v0 with any of the following classes: (<class 'transformers.models.llama.modeling_llama.LlamaForCausalLM'>,). See the original errors: while loading with LlamaForCausalLM, an error is thrown: Traceback (most recent call last): File "/src/transformers/src/transformers/pipelines/base.py", line 269, in infer_framework_load_model model = model_class.from_pretrained(model, **kwargs) File "/src/transformers/src/transformers/modeling_utils.py", line 3063, in from_pretrained raise EnvironmentError( OSError: ChocoWu/nextgpt_7b_tiva_v0 does not appear to have a file named pytorch_model.bin, tf_model.h5, model.ckpt or flax_model.msgpack.
same here.
Hi, @minar09 , @Prajwal231 , thx for interests.
We haven't integrated the model into the transformers framework yet, so you can't load the model directly by its name;
One way is to download it offline. For specific instructions, please refer to: https://github.com/NExT-GPT/NExT-GPT
In the future, we will consider enabling automatic loading of the model using transformers.
@ChocoWu Hi, I have tried navigating your model via the instructions you have provided (https://github.com/NExT-GPT/NExT-GPT) but with no success. After the setup is completed, I tried running a simple text-to-text inference but with no success. The model outputs repetitive word tokens that do not make sense (image below).
The only way to make this error go away is if the LoRA weights are ignored all together. However, then this of course cannot produce any other modality.
