File size: 1,340 Bytes
e9ec065
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
---
language:
- en
license: mit
tags:
- Explorer SubAgent
- Repository Exploration
- mlx
library_name: mlx
base_model: microsoft/FastContext-1.0-4B-SFT
pipeline_tag: text-generation
---

# FastContext-1.0-4B-SFT-mlx-8bit

8-bit MLX quantization of [microsoft/FastContext-1.0-4B-SFT](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT) for Apple Silicon.

## Quantization details

- **Method:** Affine 8-bit
- **Group size:** 64
- **Effective bits per weight:** 8.5
- **Model size:** 4.0 GB (vs 7.5 GB bf16)

## Benchmark results

Tested on 10 SWE-bench Multilingual instances against other quantization variants:

| Model | Bits/Wt | Size | File F1 | Line F1 |
|-------|---------|------|---------|---------|
| **affine 8-bit g64 (this model)** | **8.5** | **4.0G** | **0.507** | **0.140** |
| affine 4-bit g32 | 5.0 | 2.4G | 0.300 | 0.090 |
| affine 3-bit g64 | 3.5 | 1.7G | 0.100 | 0.000 |
| affine 4-bit g64 | 4.5 | 2.1G | 0.050 | 0.005 |
| mattrobenolt 4-bit g64 | 4.5 | 2.1G | 0.025 | 0.008 |

Highest quality quantization — best File F1 and Line F1 at the cost of larger size and slower inference.

## Usage

```python
from mlx_lm import load, generate

model, tokenizer = load("rubybear/FastContext-1.0-4B-SFT-mlx-8bit")
```

Or with [fastcontext-mcp](https://github.com/rubybear-lgtm/fastcontext) for Claude Code integration.