--- base_model: nphearum/psarai-2b tags: - transformers - safetensors - unsloth - gemma4 - psarai - conversational - multimodal --- # PsarAI-2B **PsarAI-2B** is a PsarAI chat model exported in Hugging Face format. The model uses a Gemma4-style architecture and a PsarAI chat template. The assistant identity in the template is: > You are PsarAI, created by the PsarAI team under the leadership of an ITC lecturer. ## Files This repository contains the standard Hugging Face model export: | File | Purpose | |---|---| | `model.safetensors` | model weights | | `config.json` | model architecture/config | | `tokenizer.json` | tokenizer | | `tokenizer_config.json` | tokenizer metadata and special tokens | | `processor_config.json` | multimodal processor config | | `chat_template.jinja` | chat formatting template | | `generation_config.json` | generation defaults | ## Quick Start ```python import torch from transformers import AutoProcessor, AutoModelForCausalLM repo_id = "nphearum/PsarAI-2B" processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( repo_id, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True, ) messages = [ {"role": "user", "content": "Who created you?"} ] prompt = processor.tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True, enable_thinking=False, ) inputs = processor.tokenizer(prompt, return_tensors="pt").to(model.device) outputs = model.generate( **inputs, max_new_tokens=256, temperature=0.7, top_p=0.9, ) print(processor.tokenizer.decode(outputs[0], skip_special_tokens=False)) ``` ## Chat Template The template uses Gemma-style tokens: - `<|turn>system` - `<|turn>user` - `<|turn>model` - `` - `<|channel>thought` - `<|tool_call>` - `<|tool_response>` For normal chatbot use, disable visible thinking when your runtime supports template kwargs: ```python enable_thinking=False ``` ## Suggested Generation Settings ```python temperature = 0.7 top_p = 0.9 max_new_tokens = 512 ``` Use lower temperature, such as `0.2`, for factual or deterministic answers. ## Multimodal Notes The config includes image, audio, and video processor metadata. Runtime support depends on the installed `transformers` version and model implementation availability. For GGUF/llama.cpp usage, use the sibling GGUF export repo instead: ```text nphearum/PsarAI-2B-GGUF ``` ## Attribution Base model metadata in this export is: ```text nphearum/psarai-2b ``` Keep this metadata for traceability when publishing derived formats.