| --- |
| license: apache-2.0 |
| --- |
| # NanoRush Chat |
|
|
| NanoRush Chat is a 283M parameter GPT-style causal language model fine-tuned for conversational AI. |
|
|
| Github- https://github.com/Amogh1221/NanoRush |
|
|
| Live- https://nano-chat-web.vercel.app |
|
|
| ## Model Details & Configuration |
| | Detail | Value | |
| | --- | --- | |
| | **Parameters** | 283M | |
| | **Architecture** | GPT-2 style | |
| | **Precision** | FP16 / BFloat16 | |
| | **Context Window** | Up to 4096 tokens | |
| | **Vocabulary Size** | 32,768 | |
| | **Embedding Dimension (n_embd)** | 768 | |
| | **Number of Heads (n_head)** | 12 | |
| | **Number of Layers (n_layer)** | 36 | |
| | **Base Model** | Custom pre-trained | |
| | **Pre-Training Dataset** | `HuggingFaceTB/cosmopedia` | |
| | **Fine-tuning Dataset** | `HuggingFaceTB/smoltalk` | |
| |
| ## Evaluation Results |
| |
| The model was evaluated using standard zero-shot accuracy metrics. |
| |
| | Groups / Tasks | Version | n-shot | Metric | Value | Stderr | |
| |---|---|---|---|---|---| |
| | **mmlu** | 2 | 0 | acc | 0.2297 | ± 0.0035 | |
| | - humanities | 2 | 0 | acc | 0.2438 | ± 0.0063 | |
| | - other | 2 | 0 | acc | 0.2375 | ± 0.0076 | |
| | - social sciences | 2 | 0 | acc | 0.2184 | ± 0.0074 | |
| | - stem | 2 | 0 | acc | 0.2119 | ± 0.0073 | |
| | **arc_challenge** | 1 | 0 | acc | 0.2295 | ± 0.0123 | |
| | **hellaswag** | 1 | 0 | acc | 0.3116 | ± 0.0046 | |
| | **truthfulqa_mc2** | 3 | 0 | acc | 0.4320 | ± 0.0153 | |
| | **winogrande** | 1 | 0 | acc | 0.5107 | ± 0.0140 | |
| |
| |
| ## Usage |
| |
| This model has been exported to be fully compatible with the Hugging Face `transformers` library. |
| You can load it using the standard `AutoModelForCausalLM` pipeline. |
| |
| ### Installation |
| Make sure you have the latest version of the `transformers` and `torch` libraries installed: |
| ```bash |
| pip install torch transformers |
| ``` |
| |
| ### Example Code |
| |
| |
| ```python |
| import torch |
| from transformers import AutoModelForCausalLM, AutoTokenizer |
| from transformers import StoppingCriteria, StoppingCriteriaList |
| from transformers.generation.streamers import TextIteratorStreamer |
| import threading |
| |
| # Load the model and tokenizer from Hugging Face |
| model_id = "Amogh1221/nano-chat" |
| tokenizer = AutoTokenizer.from_pretrained(model_id) |
| model = AutoModelForCausalLM.from_pretrained( |
| model_id, |
| torch_dtype=torch.float32, |
| device_map="cpu" |
| ) |
| system_prompt = |
| """ |
| You are NanoRush, an AI assistant, you are a helpful, respectful, and intelligent conversational |
| partner. You must never pretend to be a human, and you must carefully pay attention to the conversation history. |
| """ |
| |
| class StopOnUser(StoppingCriteria): |
| def __init__(self, prompt_length): |
| self.prompt_length = prompt_length |
| def __call__(self, input_ids, scores, **kwargs): |
| generated_tokens = input_ids[0][self.prompt_length:] |
| tail = tokenizer.decode(generated_tokens[-10:]) |
| return "\nUser:" in tail or "User:" in tail |
| |
| # Format your prompt |
| prompt = f"System: {system_prompt}\\n\\nUser: What is Quantum Computing?\\nAssistant:" |
| inputs = tokenizer(prompt, return_tensors="pt").to(model.device) |
|
|
| stop_criteria = StoppingCriteriaList([StopOnUser(prompt_length=inputs["input_ids"].shape[1])]) |
| streamer = TextIteratorStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True) |
|
|
| generation_kwargs = dict( |
| **inputs, |
| max_new_tokens=512, |
| temperature=0.7, |
| top_k=50, |
| top_p=0.9, |
| do_sample=True, |
| repetition_penalty=1.15, |
| pad_token_id=tokenizer.eos_token_id, |
| prompt_lookup_num_tokens=3, |
| stopping_criteria=stop_criteria, |
| streamer=streamer, |
| ) |
| |
| # Run generation in a background thread |
| thread = threading.Thread(target=model.generate, kwargs=generation_kwargs) |
| thread.start() |
| |
| print("Assistant: ", end="") |
| for text in streamer: |
| print(text, end="", flush=True) |
| print() |
| ``` |
| |