Instructions to use harpertoken/chat with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use harpertoken/chat with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="harpertoken/chat")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("harpertoken/chat") model = AutoModelForCausalLM.from_pretrained("harpertoken/chat", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use harpertoken/chat with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "harpertoken/chat" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harpertoken/chat", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/harpertoken/chat
- SGLang
How to use harpertoken/chat with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "harpertoken/chat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harpertoken/chat", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "harpertoken/chat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harpertoken/chat", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use harpertoken/chat with Docker Model Runner:
docker model run hf.co/harpertoken/chat
chat
openai-community/gpt2 further fine-tuned on DailyDialog, a corpus of short multi-turn English conversations about everyday topics. The base model is the 124M-parameter GPT-2: config.json records _name_or_path: gpt2, twelve layers, 768 embedding dimensions. The previous version of this card declared its base model as gpt2, an identifier that has since moved; openai-community/gpt2 is the current one.
This is a continuation, not a chat model. Nothing here establishes turn-taking, instruction following or any conversational protocol: there is no chat template in tokenizer_config.json, and the only token added beyond GPT-2's vocabulary is a [PAD] entry at id 50257, which takes vocab_size to 50258 against GPT-2's 50257. It produces fluent English that resembles dialogue because that is what it was trained to imitate, which is a weaker property than it sounds.
One detail about the checkpoint itself. It carries twelve transformer.h.*.attn.masked_bias tensors that current Transformers does not define: the key was removed from GPT2Attention years ago, and the upstream gpt2 weights contain no such tensor. Loading this repository therefore prints an UNEXPECTED warning for those twelve keys on every run. They are inert, since nothing reads them, and generation is unaffected. They are left in place rather than silently stripped, so the file matches whatever produced it.
The card this replaces documented a generate_response.py script and a FastAPI service on localhost:8000 with /status, /chat and interactive docs endpoints. None of that is in the repository, which contains weights and tokenizer files only, so those sections have been removed. The declared perplexity, BLEU and F1 metrics likewise had no values behind them; no evaluation is recorded for this checkpoint.
Usage
from transformers import pipeline
generator = pipeline("text-generation", model="harpertoken/chat")
print(generator("Hello, how are you?", max_new_tokens=60)[0]["generated_text"])
A pad_token is not configured, so batching prompts together will warn. Set tokenizer.pad_token = tokenizer.eos_token if you need it.
Limitations
DailyDialog is scripted, crowd-sourced, and narrow: 13k conversations of polite small talk, annotated for emotion and communication acts. A model trained on it will handle greetings and farewells far better than disagreement, and will reproduce the register of that corpus rather than anything wider. Because GPT-2 is a 124M-parameter base and the fine-tuning is small, output is frequently fluent and wrong. The biases of both GPT-2 and DailyDialog carry through unchanged.
Attribution
GPT-2 follows Radford et al. (2018). DailyDialog is described in Li et al., DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset, IJCNLP 2017.
- Downloads last month
- 512
Model tree for harpertoken/chat
Base model
openai-community/gpt2