Instructions to use xiapk7/AQuilt_Eval_lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use xiapk7/AQuilt_Eval_lora with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="xiapk7/AQuilt_Eval_lora") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("xiapk7/AQuilt_Eval_lora", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use xiapk7/AQuilt_Eval_lora with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "xiapk7/AQuilt_Eval_lora" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xiapk7/AQuilt_Eval_lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/xiapk7/AQuilt_Eval_lora
- SGLang
How to use xiapk7/AQuilt_Eval_lora with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "xiapk7/AQuilt_Eval_lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xiapk7/AQuilt_Eval_lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "xiapk7/AQuilt_Eval_lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xiapk7/AQuilt_Eval_lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use xiapk7/AQuilt_Eval_lora with Docker Model Runner:
docker model run hf.co/xiapk7/AQuilt_Eval_lora
Improve model card: Add paper link, metadata, and enhanced description
#1
by nielsr HF Staff - opened
This PR enhances the model card for the AQuilt LoRA adapter by adding:
- The paper link: AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs.
- Essential metadata:
pipeline_tag: text-generationfor better discoverability andlibrary_name: transformersto indicate compatibility with the Hugging Face Transformers library. - A more comprehensive description of the model based on the paper's abstract, explaining the purpose of AQuilt and this LoRA adapter.
- A section for citing the associated paper.
These updates aim to improve the model's visibility, context, and usability on the Hugging Face Hub.
Thank you for your suggestions! We will improve this.