Text Generation
MLX
Safetensors
English
Chinese
deepseek_v4
deepseek-v4
deepseek
quantized
apple-silicon
mixture-of-experts
Mixture of Experts
affine
conversational
5-bit
Instructions to use osmapi/DeepSeek-V4-Flash-5bit-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use osmapi/DeepSeek-V4-Flash-5bit-mlx with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("osmapi/DeepSeek-V4-Flash-5bit-mlx") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps
- LM Studio
- MLX LM
How to use osmapi/DeepSeek-V4-Flash-5bit-mlx with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "osmapi/DeepSeek-V4-Flash-5bit-mlx"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "osmapi/DeepSeek-V4-Flash-5bit-mlx" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "osmapi/DeepSeek-V4-Flash-5bit-mlx", "messages": [ {"role": "user", "content": "Hello"} ] }'
| { | |
| "tools": [ | |
| { | |
| "type": "function", | |
| "function": { | |
| "name": "get_weather", | |
| "description": "Get the weather for a specific location", | |
| "parameters": { | |
| "type": "object", | |
| "properties": { | |
| "location": { | |
| "type": "string", | |
| "description": "The city name" | |
| }, | |
| "unit": { | |
| "type": "string", | |
| "enum": ["celsius", "fahrenheit"], | |
| "description": "Temperature unit" | |
| } | |
| }, | |
| "required": ["location"] | |
| } | |
| } | |
| }, | |
| { | |
| "type": "function", | |
| "function": { | |
| "name": "search", | |
| "description": "Search the web for information", | |
| "parameters": { | |
| "type": "object", | |
| "properties": { | |
| "query": { | |
| "type": "string", | |
| "description": "Search query" | |
| }, | |
| "num_results": { | |
| "type": "integer", | |
| "description": "Number of results to return" | |
| } | |
| }, | |
| "required": ["query"] | |
| } | |
| } | |
| } | |
| ], | |
| "messages": [ | |
| { | |
| "role": "system", | |
| "content": "You are a helpful assistant." | |
| }, | |
| { | |
| "role": "user", | |
| "content": "What's the weather in Beijing?" | |
| }, | |
| { | |
| "role": "assistant", | |
| "reasoning_content": "The user wants to know the weather in Beijing. I should use the get_weather tool.", | |
| "tool_calls": [ | |
| { | |
| "id": "call_001", | |
| "type": "function", | |
| "function": { | |
| "name": "get_weather", | |
| "arguments": "{\"location\": \"Beijing\", \"unit\": \"celsius\"}" | |
| } | |
| } | |
| ] | |
| }, | |
| { | |
| "role": "tool", | |
| "tool_call_id": "call_001", | |
| "content": "{\"temperature\": 22, \"condition\": \"sunny\", \"humidity\": 45}" | |
| }, | |
| { | |
| "role": "assistant", | |
| "reasoning_content": "Got the weather data. Let me format a nice response.", | |
| "content": "The weather in Beijing is currently sunny with a temperature of 22°C and 45% humidity." | |
| } | |
| ] | |
| } | |