Text Generation
MLX
Safetensors
English
qwen3
Explorer SubAgent
Repository Exploration
conversational
4-bit precision
Instructions to use rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32
Run Hermes
hermes
- OpenClaw new
How to use rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rubybear/FastContext-1.0-4B-SFT-mlx-4bit-g32", "messages": [ {"role": "user", "content": "Hello"} ] }'
File size: 1,419 Bytes
25a0466 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 | ---
language:
- en
license: mit
tags:
- Explorer SubAgent
- Repository Exploration
- mlx
library_name: mlx
base_model: microsoft/FastContext-1.0-4B-SFT
pipeline_tag: text-generation
---
# FastContext-1.0-4B-SFT-mlx-4bit-g32
4-bit MLX quantization of [microsoft/FastContext-1.0-4B-SFT](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT) with group_size=32 for Apple Silicon.
## Quantization details
- **Method:** Affine 4-bit
- **Group size:** 32 (finer than the default 64)
- **Effective bits per weight:** 5.0
- **Model size:** 2.4 GB (vs 7.5 GB bf16)
## Benchmark results
Tested on 10 SWE-bench Multilingual instances against other quantization variants:
| Model | Bits/Wt | Size | File F1 | Line F1 |
|-------|---------|------|---------|---------|
| affine 8-bit g64 | 8.5 | 4.0G | 0.507 | 0.140 |
| **affine 4-bit g32 (this model)** | **5.0** | **2.4G** | **0.300** | **0.090** |
| affine 3-bit g64 | 3.5 | 1.7G | 0.100 | 0.000 |
| affine 4-bit g64 | 4.5 | 2.1G | 0.050 | 0.005 |
| mattrobenolt 4-bit g64 | 4.5 | 2.1G | 0.025 | 0.008 |
The finer group_size=32 delivers **12x better File F1** than standard 4-bit g64 quantization with only 300MB additional size.
## Usage
```python
from mlx_lm import load, generate
model, tokenizer = load("rubybear-lgtm/FastContext-1.0-4B-SFT-mlx-4bit-g32")
```
Or with [fastcontext-mcp](https://github.com/rubybear-lgtm/fastcontext) for Claude Code integration.
|