FastContext-1.0-4B-SFT — MLX 4-bit

MLX 4-bit quantization (4.501 bits per weight, ~2.1 GB) of ShaunGves/FastContext-1.0-4B-SFT, a surviving mirror of Microsoft's FastContext repository-exploration subagent (arXiv:2606.14066), removed from the official listings on 2026-06-30. Base model: Qwen3-4B-Instruct-2507. Converted with mlx-lm 0.29.1 (default group size 64).

FastContext is a small explorer model for coding agents: given a natural-language query about a repository, it explores with read-only tools (Read / Glob / Grep, called in parallel) and returns a compact <final_answer> block of file:line-range citations, keeping broad exploration out of the main agent's context window.

⚠️ Quality note (tested)

In end-to-end tests with the fastcontext CLI on an M4 Mac, this 4-bit quant showed noticeably degraded path grounding: it repeatedly assumed wrong repository root paths, failed to recover from tool errors, and hallucinated citations to nonexistent files — on the same queries the 8-bit quant answered accurately. It is published for memory-constrained setups and further experimentation; prefer the 8-bit quant (~4 GB) if your machine allows. Run with temperature ≤ 0.6 and cap max_tokens (missed stop tokens otherwise generate for minutes).

Use with LM Studio

Search for FastContext-1.0-4B-SFT-4bit in LM Studio (MLX runtime), or:

lms get mlx-community/FastContext-1.0-4B-SFT-4bit

Use with mlx-lm

uv tool install "mlx-lm==0.29.1" --with "transformers<5" --with "mlx<0.31"
mlx_lm.server --model mlx-community/FastContext-1.0-4B-SFT-4bit --port 8080

Use with the fastcontext CLI

Install from the preserved mirror (includes fixes for local OpenAI-compatible servers), then from the repo you want to explore:

export BASE_URL="http://localhost:8080/v1"   # or http://localhost:1234/v1 for LM Studio
export MODEL="mlx-community/FastContext-1.0-4B-SFT-4bit"
export API_KEY="local"
export TEMPERATURE=0.6
export MAX_TOKENS=4000

fastcontext -q "Where is the retry logic for failed API calls?" --citation
Downloads last month
62
Safetensors
Model size
0.6B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/FastContext-1.0-4B-SFT-4bit

Quantized
(5)
this model

Paper for mlx-community/FastContext-1.0-4B-SFT-4bit