File size: 3,817 Bytes
d5711b0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
# Sixpert K1 - Complete Usage Guide

## Quick Start

### Option 1: Ollama (Easiest)

```bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Download and import the model
ollama create sixpert-k1 -f OllamaModelfile

# Or if GGUF is in Ollama library:
# ollama run sixpert-k1

# Chat
ollama run sixpert-k1
```

### Option 2: llama-cpp-python (Python)

```bash
pip install llama-cpp-python
python examples/generate.py --prompt "Hello, who are you?"
```

### Option 3: API Server

```bash
pip install llama-cpp-python
python examples/api_server.py --model SixpertK1.gguf
```

### Option 4: LM Studio

1. Download LM Studio from https://lmstudio.ai
2. Import `SixpertK1.gguf`
3. Start chatting with the Sixpert K1 preset

## Chat Format

Sixpert K1 uses the following chat template:

```
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
What is quantum computing?<|im_end|>
<|im_start|>assistant
Quantum computing uses quantum mechanical phenomena...<|im_end|>
```

## Recommended Settings

| Parameter | Value | Notes |
|---|---|---|
| temperature | 0.7 | Good balance of creativity and accuracy |
| top_p | 0.8 | Nucleus sampling |
| top_k | 40 | Limit token selection |
| repeat_penalty | 1.05 | Prevent repetition |
| max_tokens | 8192 | Max output length |
| context_size | 131072 | Full context window |

## Function Calling

Sixpert K1 supports native function calling. See `examples/function_calling.py` for a complete implementation.

### Tool Format

```json
{
  "type": "function",
  "function": {
    "name": "search",
    "description": "Search for information",
    "parameters": {
      "type": "object",
      "properties": {
        "query": {"type": "string"}
      },
      "required": ["query"]
    }
  }
}
```

## Vision / Multimodal

Sixpert K1 can understand images. See `examples/vision_example.py` for implementation details.

```python
response = llm.create_chat_completion(
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Describe this image"},
            {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}},
        ]
    }]
)
```

## Integration Examples

### OpenAI-Compatible Client

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

response = client.chat.completions.create(
    model="sixpert-k1",
    messages=[{"role": "user", "content": "Explain recursion"}],
    temperature=0.7,
)
print(response.choices[0].message.content)
```

### LangChain Integration

```python
from langchain.llms import LlamaCpp

llm = LlamaCpp(
    model_path="SixpertK1.gguf",
    temperature=0.7,
    n_ctx=131072,
    n_gpu_layers=-1,
)

result = llm.invoke("What is machine learning?")
print(result)
```

### CrewAI Agent

```python
from crewai import Agent, Task, Crew

agent = Agent(
    role="Research Analyst",
    backstory="You are Sixpert K1, a precision logic engine",
    goal="Provide accurate, detailed analysis",
    llm=LlamaCpp(model_path="SixpertK1.gguf", temperature=0.7),
    allow_delegation=False,
)
```

## Performance Tips

1. **GPU Offloading**: Set `n_gpu_layers=-1` to offload all layers to GPU
2. **Context Pruning**: Use smaller context windows (8192-32768) for faster inference
3. **Batch Processing**: Use the API server for batch inference
4. **Quantization**: Q4_K_M is the sweet spot; upgrade to Q6_K if quality matters more

## Troubleshooting

| Issue | Solution |
|---|---|
| Out of memory | Reduce context size or use CPU-only inference |
| Slow generation | Enable GPU offloading (`n_gpu_layers=-1`) |
| Repetitive output | Increase `repeat_penalty` to 1.1-1.2 |
| Hallucinations | Lower temperature to 0.3-0.5 |
| Context overflow | Use 4096 context for testing, 131072 for production |