YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
GPT-2 Medium (2-Bit Quantized)
This repository contains a demonstration of a highly advanced, proprietary 2-bit quantization algorithm applied to the gpt2-medium model.
Despite being aggressively compressed to exactly 2 bits per weight, the model preserves strong grammatical structure, contextual awareness, and high-dimensional logic. This serves as a proof-of-concept for next-generation extreme compression technologies.
Note: The exact quantization method, clustering techniques, and matrix transformation parameters used to achieve this result are kept strictly confidential.
How to Test the Model
You can load and test this model using standard transformers and torch libraries. No custom or proprietary layers are required to perform inference.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from huggingface_hub import hf_hub_download
# 1. Download the quantized weights
weights_path = hf_hub_download(repo_id="dadni/GPT2-Medium-2Bit-Demo", filename="gpt2-2bit-matmulfree.pt")
# 2. Load the base architecture and tokenizer
model_id = "gpt2-medium"
model = AutoModelForCausalLM.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
# 3. Load the 2-bit weights into the model
state_dict = torch.load(weights_path, map_location="cpu")
model.load_state_dict(state_dict)
model.eval()
# 4. Generate Text
prompt = "The true secret to human intelligence is"
input_ids = tokenizer.encode(prompt, return_tensors="pt")
with torch.no_grad():
output_ids = model.generate(
input_ids,
max_new_tokens=100,
do_sample=True,
top_p=0.9,
temperature=0.8,
repetition_penalty=1.3,
pad_token_id=tokenizer.eos_token_id
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
For inquiries or demonstrations on larger models, please refer to the author.