dadni's picture
Upload README.md with huggingface_hub
018a028 verified
|
Raw
History Blame Contribute Delete
1.91 kB
# GPT-2 Medium (2-Bit Quantized)
This repository contains a demonstration of a highly advanced, proprietary 2-bit quantization algorithm applied to the `gpt2-medium` model.
Despite being aggressively compressed to exactly 2 bits per weight, the model preserves strong grammatical structure, contextual awareness, and high-dimensional logic. This serves as a proof-of-concept for next-generation extreme compression technologies.
**Note:** The exact quantization method, clustering techniques, and matrix transformation parameters used to achieve this result are kept strictly confidential.
## How to Test the Model
You can load and test this model using standard `transformers` and `torch` libraries. No custom or proprietary layers are required to perform inference.
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from huggingface_hub import hf_hub_download
# 1. Download the quantized weights
weights_path = hf_hub_download(repo_id="dadni/GPT2-Medium-2Bit-Demo", filename="gpt2-2bit-matmulfree.pt")
# 2. Load the base architecture and tokenizer
model_id = "gpt2-medium"
model = AutoModelForCausalLM.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
# 3. Load the 2-bit weights into the model
state_dict = torch.load(weights_path, map_location="cpu")
model.load_state_dict(state_dict)
model.eval()
# 4. Generate Text
prompt = "The true secret to human intelligence is"
input_ids = tokenizer.encode(prompt, return_tensors="pt")
with torch.no_grad():
output_ids = model.generate(
input_ids,
max_new_tokens=100,
do_sample=True,
top_p=0.9,
temperature=0.8,
repetition_penalty=1.3,
pad_token_id=tokenizer.eos_token_id
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
```
*For inquiries or demonstrations on larger models, please refer to the author.*