SixpertK1 / docs /usage_guide.md
SixpertAI's picture
Upload docs/usage_guide.md with huggingface_hub
d5711b0 verified
|
Raw
History Blame Contribute Delete
3.82 kB
# Sixpert K1 - Complete Usage Guide
## Quick Start
### Option 1: Ollama (Easiest)
```bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# Download and import the model
ollama create sixpert-k1 -f OllamaModelfile
# Or if GGUF is in Ollama library:
# ollama run sixpert-k1
# Chat
ollama run sixpert-k1
```
### Option 2: llama-cpp-python (Python)
```bash
pip install llama-cpp-python
python examples/generate.py --prompt "Hello, who are you?"
```
### Option 3: API Server
```bash
pip install llama-cpp-python
python examples/api_server.py --model SixpertK1.gguf
```
### Option 4: LM Studio
1. Download LM Studio from https://lmstudio.ai
2. Import `SixpertK1.gguf`
3. Start chatting with the Sixpert K1 preset
## Chat Format
Sixpert K1 uses the following chat template:
```
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
What is quantum computing?<|im_end|>
<|im_start|>assistant
Quantum computing uses quantum mechanical phenomena...<|im_end|>
```
## Recommended Settings
| Parameter | Value | Notes |
|---|---|---|
| temperature | 0.7 | Good balance of creativity and accuracy |
| top_p | 0.8 | Nucleus sampling |
| top_k | 40 | Limit token selection |
| repeat_penalty | 1.05 | Prevent repetition |
| max_tokens | 8192 | Max output length |
| context_size | 131072 | Full context window |
## Function Calling
Sixpert K1 supports native function calling. See `examples/function_calling.py` for a complete implementation.
### Tool Format
```json
{
"type": "function",
"function": {
"name": "search",
"description": "Search for information",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string"}
},
"required": ["query"]
}
}
}
```
## Vision / Multimodal
Sixpert K1 can understand images. See `examples/vision_example.py` for implementation details.
```python
response = llm.create_chat_completion(
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}},
]
}]
)
```
## Integration Examples
### OpenAI-Compatible Client
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
response = client.chat.completions.create(
model="sixpert-k1",
messages=[{"role": "user", "content": "Explain recursion"}],
temperature=0.7,
)
print(response.choices[0].message.content)
```
### LangChain Integration
```python
from langchain.llms import LlamaCpp
llm = LlamaCpp(
model_path="SixpertK1.gguf",
temperature=0.7,
n_ctx=131072,
n_gpu_layers=-1,
)
result = llm.invoke("What is machine learning?")
print(result)
```
### CrewAI Agent
```python
from crewai import Agent, Task, Crew
agent = Agent(
role="Research Analyst",
backstory="You are Sixpert K1, a precision logic engine",
goal="Provide accurate, detailed analysis",
llm=LlamaCpp(model_path="SixpertK1.gguf", temperature=0.7),
allow_delegation=False,
)
```
## Performance Tips
1. **GPU Offloading**: Set `n_gpu_layers=-1` to offload all layers to GPU
2. **Context Pruning**: Use smaller context windows (8192-32768) for faster inference
3. **Batch Processing**: Use the API server for batch inference
4. **Quantization**: Q4_K_M is the sweet spot; upgrade to Q6_K if quality matters more
## Troubleshooting
| Issue | Solution |
|---|---|
| Out of memory | Reduce context size or use CPU-only inference |
| Slow generation | Enable GPU offloading (`n_gpu_layers=-1`) |
| Repetitive output | Increase `repeat_penalty` to 1.1-1.2 |
| Hallucinations | Lower temperature to 0.3-0.5 |
| Context overflow | Use 4096 context for testing, 131072 for production |