Instructions to use mlx-community/c4ai-command-r-plus-08-2024-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mlx-community/c4ai-command-r-plus-08-2024-4bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mlx-community/c4ai-command-r-plus-08-2024-4bit") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("mlx-community/c4ai-command-r-plus-08-2024-4bit") model = AutoModelForCausalLM.from_pretrained("mlx-community/c4ai-command-r-plus-08-2024-4bit", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - MLX
How to use mlx-community/c4ai-command-r-plus-08-2024-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("mlx-community/c4ai-command-r-plus-08-2024-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- vLLM
How to use mlx-community/c4ai-command-r-plus-08-2024-4bit with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mlx-community/c4ai-command-r-plus-08-2024-4bit" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlx-community/c4ai-command-r-plus-08-2024-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/mlx-community/c4ai-command-r-plus-08-2024-4bit
- SGLang
How to use mlx-community/c4ai-command-r-plus-08-2024-4bit with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mlx-community/c4ai-command-r-plus-08-2024-4bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlx-community/c4ai-command-r-plus-08-2024-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mlx-community/c4ai-command-r-plus-08-2024-4bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlx-community/c4ai-command-r-plus-08-2024-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - MLX LM
How to use mlx-community/c4ai-command-r-plus-08-2024-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "mlx-community/c4ai-command-r-plus-08-2024-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "mlx-community/c4ai-command-r-plus-08-2024-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlx-community/c4ai-command-r-plus-08-2024-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Docker Model Runner
How to use mlx-community/c4ai-command-r-plus-08-2024-4bit with Docker Model Runner:
docker model run hf.co/mlx-community/c4ai-command-r-plus-08-2024-4bit
Model not functioning as expected
Hello, Thank you for converting this model.
However, this is not working correctly in my environment. The model loads without issue, but the generate function outputs meaningless strings. Below, I have pasted the current version of mlx_lm and the output of mlx_lm.generate.
Is it working well in your environments?
% pip show mlx_lm
Name: mlx-lm
Version: 0.18.1
Summary: LLMs on Apple silicon with MLX and the Hugging Face Hub
Home-page: https://github.com/ml-explore/mlx-examples
Author: MLX Contributors
Author-email: mlx@group.apple.com
License: MIT
Location: /Volumes/CT4000P3PSSD8JP/github/mlx_gguf_server/.venv/lib/python3.12/site-packages
Requires: jinja2, mlx, numpy, protobuf, pyyaml, transformers
Required-by:
% python -m mlx_lm.generate --model models/mlx-community/c4ai-command-r-plus-08-2024-4bit --prompt "hello"
Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.
==========
Prompt: <BOS_TOKEN><|START_OF_TURN_TOKEN|><|USER_TOKEN|>hello<|END_OF_TURN_TOKEN|><|START_OF_TURN_TOKEN|><|CHATBOT_TOKEN|>
Smir thru ope hitch Tomo seventy seventy seventy seventy seventy seventy seventy etnic Smir thru relanç勾勾 meis stump ope¹ relanç thru gib gib gib thru seventy seventy meis rééd扭 relanç扭모리 Smir Smir gib Laff Laff Laff Laff rééd etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic etnic
==========
Prompt: 8 tokens, 6.411 tokens-per-sec
Generation: 100 tokens, 10.814 tokens-per-sec
Peak memory: 54.463 GB
Thank you,
Thank you for reaching out about the model conversion issue. I'm experiencing a similar situation in my environment as well. The model loads without problems, but the generate function produces meaningless output, just as you described.
Unfortunately, I haven't been able to identify the root cause of this problem yet. I'm currently working on potential fixes, but so far without success.
If anyone else has encountered this issue and found a solution, or has any additional insights, please share them with the community. Any information that could help us resolve this would be greatly appreciated.