PsarAI-2B / README.md
nphearum's picture
Update README.md
85151d6 verified
|
Raw
History Blame Contribute Delete
2.64 kB
---
base_model: nphearum/psarai-2b
tags:
- transformers
- safetensors
- unsloth
- gemma4
- psarai
- conversational
- multimodal
---
# PsarAI-2B
**PsarAI-2B** is a PsarAI chat model exported in Hugging Face format.
The model uses a Gemma4-style architecture and a PsarAI chat template. The assistant identity in the template is:
> You are PsarAI, created by the PsarAI team under the leadership of an ITC lecturer.
## Files
This repository contains the standard Hugging Face model export:
| File | Purpose |
|---|---|
| `model.safetensors` | model weights |
| `config.json` | model architecture/config |
| `tokenizer.json` | tokenizer |
| `tokenizer_config.json` | tokenizer metadata and special tokens |
| `processor_config.json` | multimodal processor config |
| `chat_template.jinja` | chat formatting template |
| `generation_config.json` | generation defaults |
## Quick Start
```python
import torch
from transformers import AutoProcessor, AutoModelForCausalLM
repo_id = "nphearum/PsarAI-2B"
processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
messages = [
{"role": "user", "content": "Who created you?"}
]
prompt = processor.tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
inputs = processor.tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=256,
temperature=0.7,
top_p=0.9,
)
print(processor.tokenizer.decode(outputs[0], skip_special_tokens=False))
```
## Chat Template
The template uses Gemma-style tokens:
- `<|turn>system`
- `<|turn>user`
- `<|turn>model`
- `<turn|>`
- `<|channel>thought`
- `<|tool_call>`
- `<|tool_response>`
For normal chatbot use, disable visible thinking when your runtime supports template kwargs:
```python
enable_thinking=False
```
## Suggested Generation Settings
```python
temperature = 0.7
top_p = 0.9
max_new_tokens = 512
```
Use lower temperature, such as `0.2`, for factual or deterministic answers.
## Multimodal Notes
The config includes image, audio, and video processor metadata. Runtime support depends on the installed `transformers` version and model implementation availability.
For GGUF/llama.cpp usage, use the sibling GGUF export repo instead:
```text
nphearum/PsarAI-2B-GGUF
```
## Attribution
Base model metadata in this export is:
```text
nphearum/psarai-2b
```
Keep this metadata for traceability when publishing derived formats.