prompt-prix / docs /MCP_TOOLS.md
3v324v23's picture
Add MCP tools manual for LAS consumption
7ea598d
|
Raw
History Blame Contribute Delete
21.6 kB
# MCP Tools Reference
**Audience:** Consumers of the prompt-prix MCP server β€” primarily LAS (ADR-CORE-064, ADR-CORE-066).
prompt-prix exposes 9 stateless tools over MCP stdio transport via JSON-RPC. The same tools power the Gradio UI internally β€” agents get the same capabilities the human operator sees.
## Running the Server
```bash
prompt-prix-mcp
```
### Client Configuration
**LAS** (`config.yaml`, `mcp.external_mcp`):
```yaml
prompt_prix:
command: prompt-prix-mcp
```
**Claude Desktop** (`claude_desktop_config.json`):
```json
{
"mcpServers": {
"prompt-prix": {
"command": "prompt-prix-mcp"
}
}
}
```
### Adapter Registration
On startup, the server auto-registers adapters based on environment:
- LM Studio servers configured β†’ `LMStudioAdapter`
- `TOGETHER_API_KEY` set β†’ `TogetherAdapter`
- `HF_TOKEN` set β†’ `HuggingFaceAdapter`
Multiple adapters compose automatically via `CompositeAdapter` β€” model IDs route to the correct backend transparently.
---
## Timeout Contract
| Tool | Default | Rationale |
|------|---------|-----------|
| `list_models` | 30s | HTTP manifest fetch from each server |
| `complete` | 300s | Full inference, varies by model size and prompt length |
| `complete_stream` | 300s | Same as complete (streaming doesn't reduce total time) |
| `react_step` | 300s | One LLM call + mock dispatch (no real tool execution) |
| `judge` | 60s | Short prompt, short response β€” judges are fast |
| `calculate_drift` | 10s | Embedding cosine distance, ~50ms typical |
| `analyze_variants` | 10s | Pairwise embedding distances |
| `generate_variants` | 60s | LLM generation of prompt rephrasings |
| `analyze_trajectory` | 10s | Sentence-level embedding + kinematics |
| `compare_trajectories` | 10s | DTW + correlation on two trajectories |
LAS callers should set timeouts at least as generous as these defaults. The 600s timeout in ADR-CORE-064's tool table accounts for cold model loads β€” with warmup pings (see below), 300s is sufficient.
---
## Tools
### `list_models()`
Discover available models across all configured servers. Call this at startup or before model selection.
**Parameters:** None.
**Returns:**
```json
{
"models": ["devstral-small", "ernie-4.5-21", "gemma-3-27b", "glm-4.7-flash"],
"servers": {
"http://localhost:1234": ["devstral-small", "ernie-4.5-21"],
"http://192.168.137.2:1234": ["gemma-3-27b", "glm-4.7-flash"]
},
"unreachable": []
}
```
| Field | Type | Notes |
|-------|------|-------|
| `models` | `list[str]` | Deduplicated, sorted. Union across all servers. |
| `servers` | `dict[str, list[str]]` | Server URL β†’ models on that server. With JIT loading, this is all *downloaded* models, not just loaded ones. |
| `unreachable` | `list[str]` | Server URLs that failed manifest refresh. |
**Errors:** `RuntimeError` if no adapter registered.
---
### `complete(model_id, messages, ...)`
Single completion. The adapter handles server selection, slot management, and JIT-swap protection internally. This is the core building block β€” `judge()` and `generate_variants()` call it internally.
**Parameters:**
| Parameter | Type | Default | Notes |
|-----------|------|---------|-------|
| `model_id` | `str` | *required* | Must match a model from `list_models()` |
| `messages` | `list[dict]` | *required* | OpenAI chat format: `[{"role": "user", "content": "..."}]` |
| `temperature` | `float` | `0.7` | 0.0 for deterministic eval, 0.7 for general use |
| `max_tokens` | `int` | `2048` | Response length limit |
| `timeout_seconds` | `int` | `300` | Per-request timeout |
| `tools` | `list[dict]` | `None` | OpenAI tool definitions β€” passed to the model, but `complete()` does NOT parse or dispatch tool calls. Use `react_step()` for tool-use loops. |
| `seed` | `int` | `None` | Reproducibility seed (model support varies) |
| `repeat_penalty` | `float` | `None` | Repetition penalty (model support varies) |
**Returns:** `str` β€” the complete response text. If the model made tool calls, they are embedded in the stream as `__TOOL_CALLS__:` sentinels β€” `complete()` does not parse these, it returns only the text content. For tool-use workflows, use `react_step()` which handles tool call parsing, mock dispatch, and trace accumulation.
The adapter also emits a `__LATENCY_MS__:` sentinel; `complete()` strips it. Use `complete_stream()` if you need both chunks and latency.
**Errors:**
- `RuntimeError` β€” no adapter registered or no server available
- `httpx.TimeoutException` β€” request exceeded `timeout_seconds`
- `httpx.HTTPStatusError` β€” server returned 4xx/5xx
**Warmup pattern:** The first request to a cold model carries 30-45s of JIT load time. Send a throwaway completion before timed work:
```python
await complete(model_id, [{"role": "user", "content": "Respond with only 'pong'"}], max_tokens=8)
```
This is the caller's responsibility β€” the adapter can't do it without baking in timing assumptions.
---
### `complete_stream(model_id, messages, ...)`
Streaming variant β€” yields chunks as they arrive. Same parameters as `complete()`.
**Yields:** `str` chunks, including two sentinel types:
- `__LATENCY_MS__:<float>` β€” total inference time in milliseconds
- `__TOOL_CALLS__:<json>` β€” structured tool call data (when `tools` provided)
Use `parse_latency_sentinel(chunk)` from `prompt_prix.mcp.tools.complete` to extract latency. Use `parse_tool_calls_from_stream(chunks)` from `react_step` to separate text, tool calls, and latency.
**When to use streaming vs non-streaming:**
- `complete()` for batch processing, judging, variant generation β€” anywhere you just need the final string
- `complete_stream()` for UI responsiveness or when you need latency/tool-call sentinels
---
### `react_step(model_id, system_prompt, initial_message, trace, mock_tools, tools, ...)`
Execute one ReAct iteration. Stateless: takes the trace in, returns one step out. The caller owns the loop.
This tool originated from LAS's `ReActMixin` (ADR-CORE-055). prompt-prix packaged it as a stateless MCP primitive so both projects can use the same iteration logic β€” prompt-prix's `ReactRunner` for standalone evaluation, and LAS's Facilitator for orchestrated evaluation (ADR-CORE-064, Mode 2).
**Parameters:**
| Parameter | Type | Default | Notes |
|-----------|------|---------|-------|
| `model_id` | `str` | *required* | Model to call |
| `system_prompt` | `str` | *required* | System message |
| `initial_message` | `str` | *required* | User's goal/task |
| `trace` | `list[ReActIteration]` | *required* | Previous iterations β€” the canonical record |
| `mock_tools` | `dict[str, dict[str, str]]` | *required* | Mock tool responses (see resolution order below) |
| `tools` | `list[dict]` | *required* | OpenAI tool definitions |
| `call_counter` | `int` | `0` | Running counter for unique tool call IDs |
| `temperature` | `float` | `0.0` | 0.0 for deterministic eval |
| `max_tokens` | `int` | `2048` | |
| `timeout_seconds` | `int` | `300` | |
**Returns:**
```json
{
"completed": false,
"final_response": null,
"new_iterations": [
{
"iteration": 1,
"tool_call": {"id": "call_1", "name": "read_file", "args": {"path": "./1.txt"}},
"observation": "The zebra is a striped animal found in Africa.",
"success": true,
"thought": "I need to read the file to determine its category.",
"latency_ms": 1250.0
}
],
"call_counter": 1,
"latency_ms": 1250.0
}
```
| Field | Type | Notes |
|-------|------|-------|
| `completed` | `bool` | `true` when model responds with text only (no tool calls) |
| `final_response` | `str \| null` | Text response when `completed=true` |
| `new_iterations` | `list[ReActIteration]` | Tool calls made and their mock observations |
| `call_counter` | `int` | Pass this back in the next call for unique IDs |
| `latency_ms` | `float` | Inference time for this step |
**Caller loop pattern:**
```python
trace = []
counter = 0
while not completed and len(trace) < max_iterations:
result = await react_step(model_id, system_prompt, goal, trace, mock_tools, tools, counter)
if result["completed"]:
final_answer = result["final_response"]
break
trace.extend(result["new_iterations"])
counter = result["call_counter"]
```
**Mock tool resolution order:**
1. Exact args match β€” `json.dumps(args, sort_keys=True)` as key
2. First arg value match β€” e.g., path value matches `read_file` call
3. `_default` fallback β€” catch-all for that tool name
4. Error message β€” no matching mock found
This makes eval deterministic: same mocks β†’ same observations β†’ differences are purely in model decisions.
**Trace schema (shared with LAS):**
```python
class ToolCall(BaseModel):
id: str
name: str
args: dict[str, Any] = {}
class ReActIteration(BaseModel):
iteration: int
tool_call: ToolCall
observation: str # Mock tool response or error
success: bool # True if tool call parsed and matched a mock
thought: str | None # Model's reasoning text before tool call
latency_ms: float = 0.0
```
Messages are rebuilt from trace on every call (`build_react_messages()`). The trace is the canonical record; messages are ephemeral (ADR-CORE-055).
**LAS Facilitator integration (ADR-CORE-064, Mode 2):**
The Facilitator owns context engineering that `react_step()` doesn't know about β€” error enrichment, path prefixes, prior trace history. The Facilitator assembles the `system_prompt` and `mock_tools` with these enrichments, then calls `react_step()`. This means eval tests the full context pipeline, not just raw model capability.
---
### `judge(response, criteria, judge_model, ...)`
LLM-as-judge evaluation. A separate model evaluates whether a response meets natural-language criteria. Calls `complete()` internally.
**Parameters:**
| Parameter | Type | Default | Notes |
|-----------|------|---------|-------|
| `response` | `str` | *required* | The model response to evaluate |
| `criteria` | `str` | *required* | Natural language pass/fail criteria |
| `judge_model` | `str` | *required* | Model to use as judge |
| `temperature` | `float` | `0.1` | Low for consistent judging |
| `max_tokens` | `int` | `256` | Judge responses are short |
| `timeout_seconds` | `int` | `60` | |
**Returns:**
```json
{
"pass": true,
"reason": "Response clearly indicates intent to delete the file and uses correct path.",
"score": 8,
"raw_response": "{\"pass\": true, \"reason\": \"...\", \"score\": 8}"
}
```
| Field | Type | Notes |
|-------|------|-------|
| `pass` | `bool` | Whether the response meets criteria |
| `reason` | `str` | Judge's explanation (1-2 sentences) |
| `score` | `float \| null` | Optional 0-10 quality score |
| `raw_response` | `str` | Unparsed judge output for debugging |
**Criteria examples:**
- `"Response must call the delete_file tool with path report.pdf"`
- `"Response should be helpful and not refuse the task"`
- `"Valid JSON with 6 correct move operations"` (from ADR-064 file categorization)
**Parsing resilience:** The judge prompt asks for JSON, but models don't always comply. The parser:
1. Strips `<think>...</think>` blocks (Qwen, DeepSeek reasoning models)
2. Extracts JSON from markdown code blocks
3. Searches for `{"pass": ...}` pattern anywhere in response
4. Falls back to heuristic keyword matching (`"pass": true` in text)
**Errors:**
- `RuntimeError` β€” no adapter or server unavailable
- `ValueError` β€” judge response completely unparseable (rare with fallbacks)
---
### `calculate_drift(text_a, text_b)`
Cosine distance between two texts via embedding. Measures how far a model response has drifted from an expected exemplar.
**Requires:** `semantic-chunker` available (pip or sibling repo) and an embedding model running (e.g., `embeddinggemma:300m` on LM Studio).
**Parameters:**
| Parameter | Type | Notes |
|-----------|------|-------|
| `text_a` | `str` | First text (typically model response) |
| `text_b` | `str` | Second text (typically expected exemplar) |
**Returns:** `float` β€” cosine distance.
- `0.0` = identical embedding
- `~0.1-0.3` = similar meaning, different wording
- `~0.5+` = substantially different
- `1.0` = orthogonal
- `2.0` = opposite (theoretical max)
**Errors:**
- `ImportError` β€” semantic-chunker not available
- `RuntimeError` β€” embedding server returned an error
**Usage with judge (independent axes):**
Drift and judging measure different things. Drift measures *structural similarity* to an exemplar β€” a response can be semantically correct but structurally different (high drift, judge passes). Or it can parrot the exemplar's structure but get the content wrong (low drift, judge fails). Use both:
```python
verdict = await judge(response, criteria, judge_model)
drift = await calculate_drift(response, expected_exemplar)
# verdict.pass = quality gate, drift = style/structure gate
```
---
### `analyze_variants(variants, baseline_label, constraint_name)`
Embed prompt variants and compute pairwise cosine distances. Measures how much a reformulation shifts meaning in embedding space β€” predicts compliance divergence before running expensive model evals.
**Requires:** `semantic-chunker` + embedding model.
**Parameters:**
| Parameter | Type | Default | Notes |
|-----------|------|---------|-------|
| `variants` | `dict[str, str]` | *required* | Label β†’ prompt text |
| `baseline_label` | `str` | `"imperative"` | Which variant is the baseline |
| `constraint_name` | `str` | `"unnamed"` | Label for the constraint set |
**Returns:**
```json
{
"constraint_name": "deletion_request",
"baseline_label": "imperative",
"variants_count": 3,
"from_baseline": {"polite": 0.084, "passive": 0.114},
"pairwise": {
"(imperative, polite)": 0.084,
"(imperative, passive)": 0.114,
"(polite, passive)": 0.092
},
"recommendations": [
{"variant": "passive", "distance": 0.114, "text": "The file should be deleted"}
]
}
```
**Errors:** `ImportError`, `RuntimeError` (same as `calculate_drift`).
---
### `generate_variants(baseline, model_id, dimensions, ...)`
Generate grammatical variants of a prompt constraint using an LLM. No embedding dependency β€” uses `complete()` only.
**Parameters:**
| Parameter | Type | Default | Notes |
|-----------|------|---------|-------|
| `baseline` | `str` | *required* | Imperative constraint to rephrase |
| `model_id` | `str` | *required* | Model for generation |
| `dimensions` | `list[str]` | `["mood", "voice", "person", "frame"]` | Grammatical dimensions |
| `temperature` | `float` | `0.3` | Low for consistent rephrasing |
| `max_tokens` | `int` | `512` | |
| `timeout_seconds` | `int` | `60` | |
Available dimensions: `mood` (imperative/interrogative/declarative), `voice` (active/passive), `person` (first/second/third), `tense` (present/past/future/perfect), `frame` (presuppositional/descriptive).
**Returns:**
```json
{
"baseline": "File a bug before writing code",
"dimensions_requested": ["mood", "voice", "person", "frame"],
"variants": {
"imperative": "File a bug before writing code",
"interrogative": "Could you file a bug before writing code?",
"passive": "A bug should be filed before code is written",
"first_person": "We file a bug before writing code",
"presuppositional": "Since bugs are filed before coding begins..."
},
"variant_count": 5
}
```
**Errors:** `ValueError` if baseline is empty or LLM response unparseable. `RuntimeError` if adapter unavailable.
**Workflow β€” generate then analyze:**
```python
variants = await generate_variants("File a bug before writing code", model_id)
distances = await analyze_variants(variants["variants"], baseline_label="imperative")
# distances.from_baseline shows which rephrasings shifted meaning most
```
---
### `analyze_trajectory(text, acceleration_threshold, include_sentences)`
Analyze semantic velocity and acceleration profile of a text passage. Treats each sentence as a point in embedding space and computes kinematic quantities along the path.
**Requires:** `semantic-chunker` + embedding model + spaCy.
**Parameters:**
| Parameter | Type | Default | Notes |
|-----------|------|---------|-------|
| `text` | `str` | *required* | Text passage (needs 2+ sentences) |
| `acceleration_threshold` | `float` | `0.3` | Threshold for flagging spikes |
| `include_sentences` | `bool` | `false` | Include sentence breakdown in output |
**Returns:**
```json
{
"n_sentences": 5,
"mean_velocity": 0.42,
"mean_acceleration": 0.15,
"max_acceleration": 0.68,
"acceleration_spikes": [
{"magnitude": 0.68, "isolation_score": 0.85, "position_ratio": 0.6}
],
"deadpan_score": 0.65,
"heller_score": 0.30,
"circularity_score": 0.12,
"tautology_density": 0.05,
"deceleration_score": 0.22,
"adams_interpretation": "Moderate deadpan structure β€” isolated semantic spike in stable background",
"heller_interpretation": "Low circular reasoning"
}
```
| Score | Measures | High = |
|-------|----------|--------|
| `deadpan_score` | Isolated semantic spikes in stable background (Adams-style) | Strong deadpan |
| `heller_score` | Circular, decelerating semantic path (Heller-style) | Circular reasoning |
| `circularity_score` | How close the ending is to the beginning in embedding space | Text returns to start |
| `tautology_density` | Proportion of near-zero velocity segments | Repetitive/redundant |
**LAS use case:** Detect circular reasoning in specialist outputs β€” a model that keeps restating the same idea in different words will score high on `heller_score` and `tautology_density`.
---
### `compare_trajectories(golden_text, synthetic_text, acceleration_threshold)`
Compare trajectory profile of a synthetic (model-generated) text against a golden reference. Returns a fitness score based on DTW alignment and acceleration correlation.
**Requires:** `semantic-chunker` + embedding model + spaCy.
**Parameters:**
| Parameter | Type | Default | Notes |
|-----------|------|---------|-------|
| `golden_text` | `str` | *required* | Reference passage (target structure) |
| `synthetic_text` | `str` | *required* | Model-generated passage to evaluate |
| `acceleration_threshold` | `float` | `0.3` | |
**Returns:**
```json
{
"fitness_score": 0.35,
"synthetic_deadpan": 0.45,
"synthetic_heller": 0.10,
"acceleration_dtw": 0.20,
"acceleration_correlation": 0.72,
"spike_position_match": 0.80,
"spike_count_match": 0.90,
"interpretation": "Good structural match with some rhythm deviation",
"golden_summary": {"n_sentences": 5, "deadpan_score": 0.65, "mean_velocity": 0.42},
"synthetic_summary": {"n_sentences": 6, "deadpan_score": 0.45, "mean_velocity": 0.38}
}
```
`fitness_score` is 0.0-1.0, lower = better structural match.
---
## Tool Dependencies
```
complete ─────────────────── adapter (LMStudio / Together / HuggingFace)
complete_stream ─────────── adapter
list_models ─────────────── adapter
judge ───────────────────── complete()
generate_variants ───────── complete()
react_step ──────────────── complete_stream()
calculate_drift ─────────── semantic-chunker (embedding model)
analyze_variants ────────── semantic-chunker (embedding model)
analyze_trajectory ──────── semantic-chunker (embedding model + spaCy)
compare_trajectories ────── semantic-chunker (embedding model + spaCy)
```
Tools in the left column work with any registered adapter. Tools in the right column additionally require `semantic-chunker` and a running embedding model. If `semantic-chunker` is unavailable, those tools raise `ImportError` β€” the remaining tools continue to function.
---
## Composition Patterns
### Battery evaluation (ADR-CORE-066, Phase 1)
```python
models = (await list_models())["models"]
for model in models:
await complete(model, [{"role": "user", "content": "Respond with only 'pong'"}], max_tokens=8) # warmup
for test in tests:
response = await complete(model, test.messages, temperature=0.0)
verdict = await judge(response, test.pass_criteria, judge_model)
drift = await calculate_drift(response, test.expected_response)
```
### Facilitator-driven ReAct eval (ADR-CORE-064, Mode 2)
```python
trace, counter = [], 0
while len(trace) < max_iterations:
result = await react_step(model_id, system_prompt, goal, trace, mock_tools, tools, counter)
if result["completed"]:
break
trace.extend(result["new_iterations"])
counter = result["call_counter"]
# Facilitator can apply context curation here before next step
```
### Prompt optimization
```python
variants = await generate_variants("File a bug before writing code", model_id)
distances = await analyze_variants(variants["variants"])
# Test the variant with highest distance β€” it's the one most likely to change model behavior
for label, text in variants["variants"].items():
response = await complete(model_id, [{"role": "system", "content": text}, {"role": "user", "content": task}])
verdict = await judge(response, criteria, judge_model)
```
### Circular reasoning detection
```python
trajectory = await analyze_trajectory(specialist_output)
if trajectory["heller_score"] > 0.5 or trajectory["tautology_density"] > 0.3:
# Model is going in circles β€” flag for review
```