Text Generation
Transformers
Safetensors
PyTorch
MLX
llama
facebook
meta
llama-3
mlx-my-repo
text-generation-inference
4-bit precision
Instructions to use cyboghostginx/Llama3.1-8B-mlx-4Bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cyboghostginx/Llama3.1-8B-mlx-4Bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="cyboghostginx/Llama3.1-8B-mlx-4Bit")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("cyboghostginx/Llama3.1-8B-mlx-4Bit") model = AutoModelForCausalLM.from_pretrained("cyboghostginx/Llama3.1-8B-mlx-4Bit", device_map="auto") - MLX
How to use cyboghostginx/Llama3.1-8B-mlx-4Bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("cyboghostginx/Llama3.1-8B-mlx-4Bit") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- vLLM
How to use cyboghostginx/Llama3.1-8B-mlx-4Bit with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "cyboghostginx/Llama3.1-8B-mlx-4Bit" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cyboghostginx/Llama3.1-8B-mlx-4Bit", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/cyboghostginx/Llama3.1-8B-mlx-4Bit
- SGLang
How to use cyboghostginx/Llama3.1-8B-mlx-4Bit with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "cyboghostginx/Llama3.1-8B-mlx-4Bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cyboghostginx/Llama3.1-8B-mlx-4Bit", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "cyboghostginx/Llama3.1-8B-mlx-4Bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cyboghostginx/Llama3.1-8B-mlx-4Bit", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - MLX LM
How to use cyboghostginx/Llama3.1-8B-mlx-4Bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "cyboghostginx/Llama3.1-8B-mlx-4Bit" --prompt "Once upon a time"
- Docker Model Runner
How to use cyboghostginx/Llama3.1-8B-mlx-4Bit with Docker Model Runner:
docker model run hf.co/cyboghostginx/Llama3.1-8B-mlx-4Bit
- Atomic Chat
Rewrite model card
Browse files
README.md
CHANGED
|
@@ -189,31 +189,30 @@ extra_gated_description: The information you provide will be collected, stored,
|
|
| 189 |
and shared in accordance with the [Meta Privacy Policy](https://www.facebook.com/privacy/policy/).
|
| 190 |
extra_gated_button_content: Submit
|
| 191 |
library_name: transformers
|
| 192 |
-
base_model:
|
|
|
|
| 193 |
---
|
|
|
|
| 194 |
|
| 195 |
-
|
| 196 |
|
| 197 |
-
|
| 198 |
-
|
| 199 |
-
## Use with mlx
|
| 200 |
|
| 201 |
```bash
|
| 202 |
pip install mlx-lm
|
|
|
|
|
|
|
| 203 |
```
|
| 204 |
|
| 205 |
```python
|
| 206 |
from mlx_lm import load, generate
|
| 207 |
|
| 208 |
-
model, tokenizer = load("
|
|
|
|
|
|
|
| 209 |
|
| 210 |
-
|
| 211 |
|
| 212 |
-
|
| 213 |
-
messages = [{"role": "user", "content": prompt}]
|
| 214 |
-
prompt = tokenizer.apply_chat_template(
|
| 215 |
-
messages, tokenize=False, add_generation_prompt=True
|
| 216 |
-
)
|
| 217 |
|
| 218 |
-
|
| 219 |
-
```
|
|
|
|
| 189 |
and shared in accordance with the [Meta Privacy Policy](https://www.facebook.com/privacy/policy/).
|
| 190 |
extra_gated_button_content: Submit
|
| 191 |
library_name: transformers
|
| 192 |
+
base_model: cyboghostginx/Llama3.1-8B
|
| 193 |
+
base_model_relation: quantized
|
| 194 |
---
|
| 195 |
+
# Llama3.1-8B-mlx-4Bit
|
| 196 |
|
| 197 |
+
Built with Llama. 4-bit MLX conversion of Meta's **Llama 3.1 8B base** model for Apple silicon. 4.5 GB, single shard, converted with mlx-lm 0.26.4 from [cyboghostginx/Llama3.1-8B](https://huggingface.co/cyboghostginx/Llama3.1-8B). Weights only, no fine-tuning.
|
| 198 |
|
| 199 |
+
**This is the base model, not Instruct.** It ships no chat template, so it completes text rather than answering turns. For chat, quantize an Instruct checkpoint instead.
|
|
|
|
|
|
|
| 200 |
|
| 201 |
```bash
|
| 202 |
pip install mlx-lm
|
| 203 |
+
mlx_lm.generate --model cyboghostginx/Llama3.1-8B-mlx-4Bit \
|
| 204 |
+
--prompt "The three laws of robotics are" --max-tokens 256
|
| 205 |
```
|
| 206 |
|
| 207 |
```python
|
| 208 |
from mlx_lm import load, generate
|
| 209 |
|
| 210 |
+
model, tokenizer = load("cyboghostginx/Llama3.1-8B-mlx-4Bit")
|
| 211 |
+
print(generate(model, tokenizer, prompt="The three laws of robotics are", verbose=True))
|
| 212 |
+
```
|
| 213 |
|
| 214 |
+
Higher precision: [8-bit MLX](https://huggingface.co/cyboghostginx/Llama3.1-8B-mlx-8Bit), 8.5 GB.
|
| 215 |
|
| 216 |
+
## License
|
|
|
|
|
|
|
|
|
|
|
|
|
| 217 |
|
| 218 |
+
Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved. The full agreement is reproduced in the gate above.
|
|
|