| # Tool calling with these APEX quants (reproducible) |
|
|
| Captured against the `IQ3_XXS` quant. Turn 2 depends on turn 1's result (a chained call, not a one-off), and a distractor tool (`search_web`) is offered but correctly never called. |
|
|
| ## Server |
|
|
| ```bash |
| llama-server --jinja --ctx-size 8192 -fa on --temp 0.6 --top-p 0.9 --repeat-penalty 1.05 --model <this-quant>.gguf --host 127.0.0.1 --port 8092 |
| ``` |
|
|
| (this GGUF carries its own chat template, so `--jinja` alone is enough — no `--chat-template-file` needed.) |
|
|
| ## System prompt |
|
|
| ``` |
| You are a helpful assistant with access to tools. When a tool is needed, call it; when it returns a result, answer the user directly using that result. |
| ``` |
|
|
| ## Tools offered |
|
|
| ```json |
| [ |
| { |
| "type": "function", |
| "function": { |
| "name": "get_user_location", |
| "description": "Get the user's current city.", |
| "parameters": { |
| "type": "object", |
| "properties": {} |
| } |
| } |
| }, |
| { |
| "type": "function", |
| "function": { |
| "name": "get_weather", |
| "description": "Get the current weather for a location.", |
| "parameters": { |
| "type": "object", |
| "properties": { |
| "location": { |
| "type": "string", |
| "description": "City name" |
| }, |
| "unit": { |
| "type": "string", |
| "enum": [ |
| "celsius", |
| "fahrenheit" |
| ] |
| } |
| }, |
| "required": [ |
| "location" |
| ] |
| } |
| } |
| }, |
| { |
| "type": "function", |
| "function": { |
| "name": "search_web", |
| "description": "Search the web for a query.", |
| "parameters": { |
| "type": "object", |
| "properties": { |
| "query": { |
| "type": "string" |
| } |
| }, |
| "required": [ |
| "query" |
| ] |
| } |
| } |
| } |
| ] |
| ``` |
|
|
| ## Conversation |
|
|
| **User:** What is the weather at my current location? Use celsius. |
|
|
| **Turn 1:** LLM sent properly formatted JSON to the test server |
|
|
| <details><summary>raw response JSON</summary> |
|
|
| ```json |
| { |
| "role": "assistant", |
| "content": "", |
| "reasoning_content": "I need to get the user's current location first, and then get the weather for that location in celsius. Let me start by getting the user's location.\n", |
| "tool_calls": [ |
| { |
| "type": "function", |
| "function": { |
| "name": "get_user_location", |
| "arguments": "{}" |
| }, |
| "id": "ubllaF0cqjmv496J8GurXOhRxQ1X0VDw" |
| } |
| ] |
| } |
| ``` |
|
|
| </details> |
|
|
| **Tool result (`get_user_location`):** `{"city": "Tokyo"}` |
|
|
| **Turn 2:** LLM sent properly formatted JSON to the test server |
|
|
| <details><summary>raw response JSON</summary> |
|
|
| ```json |
| { |
| "role": "assistant", |
| "content": "", |
| "reasoning_content": "Now I have the user's location: Tokyo. Let me get the weather for Tokyo in celsius.\n", |
| "tool_calls": [ |
| { |
| "type": "function", |
| "function": { |
| "name": "get_weather", |
| "arguments": "{\"location\":\"Tokyo\",\"unit\":\"celsius\"}" |
| }, |
| "id": "9LZjmuRGczb44iPrUACoU9LIqQDdK4Op" |
| } |
| ] |
| } |
| ``` |
|
|
| </details> |
|
|
| **Tool result (`get_weather`):** `{"temp_c": 18, "conditions": "cloudy"}` |
| |
| **Turn 3:** LLM correctly gave a plain-language final answer, no further tool call |
|
|
| > The weather at your current location (Tokyo) is **18°C** and **cloudy**. |
|
|
| <details><summary>raw response JSON</summary> |
|
|
| ```json |
| { |
| "role": "assistant", |
| "content": "The weather at your current location (Tokyo) is **18\u00b0C** and **cloudy**.", |
| "reasoning_content": "I got the weather for Tokyo. The temperature is 18\u00b0C and it's cloudy. Let me provide this information to the user.\n" |
| } |
| ``` |
|
|
| </details> |
|
|
|
|