DrontGPTQ / README.md
MishaGGG's picture
Update README.md
0fc188e verified
|
Raw
History Blame Contribute Delete
739 Bytes
---
license: mit
language:
- en
tags:
- Q
- Q2
- Q3
- Q4
- Q5
- Q6
- Q12
- Q8
- Q16
- Q32
- Dront
---
DrontGPTQ is a simple quantizer for GPT-like models.
Supports:
2 bit ✅
3 bit ✅
4 bit ✅
5 bit ✅
6 bit ✅
7 bit ❌
8 bit ✅
12 bit ✅
16 bit ✅
32 bit ✅
Consists of just two files.
DrontQ.py – a very simple quantizer. Converts HF F32 to .dq format (Dront quant).
DQuse.py – a simple way to run .dq models.
*To use this, scroll to the end of the files; you will find the main variables there, which you should adjust to your needs.*
*Quantizes local models—specifically GPT-like ones.*
**Most Dront gpt-like models will also be released in DQ quantization! Especially the new VdrontV**