CodeLens-7B-MLX / README.md
xunker's picture
Upload README
25c04ed verified
|
Raw
History Blame
2.08 kB
metadata
license: apache-2.0
base_model: Qwen/Qwen2.5-7B-Instruct
tags:
  - code
  - code-review
  - programming
  - qwen2.5
  - bug-detection
  - mlx
datasets:
  - sahil2801/CodeAlpaca-20k
language:
  - en
pipeline_tag: text-generation
library_name: transformers
model-index:
  - name: CodeLens-7B
    results: []

CodeLens-7B-MLX

MLX version of sriksven/CodeLens-7B in various oQ levels and dtypes.

Directory oQ Level dtype size
CodeLens-7B-oQ4-bf16 4-bit bfloat16 4.2GB
CodeLens-7B-oQ4-fp16 4-bit fp16 4.2GB
CodeLens-7B-oQ6-bf16 6-bit bfloat16 5.9GB
CodeLens-7B-oQ6-fp16 6-bit fp16 5.9GB
CodeLens-7B-oQ8-bf16 8-bit bfloat16 7.5GB
CodeLens-7B-oQ8-fp16 8-bit fp16 7.5GB

Why choose FP16 over BFLOAT16/BF16?

On older Apple Silicon (M1 and M2), fp16 can be faster. Here are the details from Muhammad Raza:

A lot of MLX builds ship as bf16, and on the M1 and M2 that data type does not get the accelerated path that fp16 does. During prefill those weights run un-accelerated and the penalty multiplies across every input token, which is part of why some “MLX is slow” reports come from older hardware. [...]

If you are on an M1 or M2 and MLX feels sluggish, check this before you blame the format.

Hardware and Software

These were converted to MLX using oMLX 0.4.4 on a 32GB Macbook Pro 2021 (M1 Pro). I cleared all my RAM so you don't have to.

License

Apache 2.0, as per original model.