Text Generation
Transformers
Safetensors
English
Chinese
Russian
yue2
music-generation
orbitquant
quantization
4-bit precision
custom-code
8-bit precision
Instructions to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="WaveCut/YuE2-3B-OrbitQuant-W4A4")# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("WaveCut/YuE2-3B-OrbitQuant-W4A4", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WaveCut/YuE2-3B-OrbitQuant-W4A4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/YuE2-3B-OrbitQuant-W4A4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/WaveCut/YuE2-3B-OrbitQuant-W4A4
- SGLang
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WaveCut/YuE2-3B-OrbitQuant-W4A4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/YuE2-3B-OrbitQuant-W4A4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WaveCut/YuE2-3B-OrbitQuant-W4A4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/YuE2-3B-OrbitQuant-W4A4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with Docker Model Runner:
docker model run hf.co/WaveCut/YuE2-3B-OrbitQuant-W4A4
| from __future__ import annotations | |
| import torch | |
| from orbitquant_packed_matmul import matmul_packed_weight | |
| def _pack(values: torch.Tensor, bits: int) -> torch.Tensor: | |
| flat = values.detach().to(device="cpu", dtype=torch.uint8).flatten() | |
| packed = torch.zeros((flat.numel() * bits + 7) // 8, dtype=torch.uint8) | |
| for value_index, value in enumerate(flat.tolist()): | |
| bit_start = value_index * bits | |
| byte_index = bit_start // 8 | |
| shift = bit_start % 8 | |
| packed[byte_index] |= (value << shift) & 0xFF | |
| if shift + bits > 8: | |
| packed[byte_index + 1] |= value >> (8 - shift) | |
| return packed | |
| device = "cuda" if torch.cuda.is_available() else "mps" | |
| bits = 4 | |
| rows = 8 | |
| in_features = 16 | |
| out_features = 6 | |
| x = torch.randn(rows, in_features, device=device, dtype=torch.float16) | |
| indices = torch.arange(out_features * in_features, dtype=torch.uint8).reshape( | |
| out_features, in_features | |
| ) % (2**bits) | |
| packed = _pack(indices, bits).to(device) | |
| row_norms = torch.linspace(0.5, 1.5, out_features, device=device) | |
| centroids = torch.linspace(-1.0, 1.0, 2**bits, device=device) | |
| out = matmul_packed_weight( | |
| x, | |
| packed, | |
| row_norms, | |
| centroids, | |
| bits=bits, | |
| out_features=out_features, | |
| in_features=in_features, | |
| ) | |
| print(out.shape) | |