Text Generation
Transformers
PyTorch
TensorBoard
Safetensors
llama
Generated from Trainer
text-generation-inference
Instructions to use flytech/devchat-llama-7b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use flytech/devchat-llama-7b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="flytech/devchat-llama-7b")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("flytech/devchat-llama-7b") model = AutoModelForCausalLM.from_pretrained("flytech/devchat-llama-7b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use flytech/devchat-llama-7b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "flytech/devchat-llama-7b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "flytech/devchat-llama-7b", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/flytech/devchat-llama-7b
- SGLang
How to use flytech/devchat-llama-7b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "flytech/devchat-llama-7b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "flytech/devchat-llama-7b", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "flytech/devchat-llama-7b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "flytech/devchat-llama-7b", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use flytech/devchat-llama-7b with Docker Model Runner:
docker model run hf.co/flytech/devchat-llama-7b
Commit History
Training in progress, step 900 95d1ba2
Training in progress, step 800 11e2935
Training in progress, step 700 9850581
Training in progress, step 600 a57e1a1
Training in progress, step 500 ad23f6f
Training in progress, step 400 bb696e0
Training in progress, step 300 439ac19
Training in progress, step 200 49cdab3
Training in progress, step 100 4f5f066
Training in progress, step 4 68a52f0
Training in progress, step 2 1565d64
Training in progress, step 200 75e21f3
sirr commited on
Training in progress, step 175 5ec101b
sirr commited on
Training in progress, step 150 5f487a7
sirr commited on
Training in progress, step 125 b9c0aac
sirr commited on
Training in progress, step 100 7dc0558
sirr commited on
Training in progress, step 75 b7d4dc5
sirr commited on
Training in progress, step 50 5abeaef
sirr commited on
Training in progress, step 25 c7937b5
sirr commited on
Training in progress, step 100 2cce779
sirr commited on
Model save 446e645
sirr commited on
Model save 915f3c6
sirr commited on