Spaces:
Running
A newer version of the Gradio SDK is available: 6.22.0
MCP Tools Reference
Audience: Consumers of the prompt-prix MCP server β primarily LAS (ADR-CORE-064, ADR-CORE-066).
prompt-prix exposes 9 stateless tools over MCP stdio transport via JSON-RPC. The same tools power the Gradio UI internally β agents get the same capabilities the human operator sees.
Running the Server
prompt-prix-mcp
Client Configuration
LAS (config.yaml, mcp.external_mcp):
prompt_prix:
command: prompt-prix-mcp
Claude Desktop (claude_desktop_config.json):
{
"mcpServers": {
"prompt-prix": {
"command": "prompt-prix-mcp"
}
}
}
Adapter Registration
On startup, the server auto-registers adapters based on environment:
- LM Studio servers configured β
LMStudioAdapter TOGETHER_API_KEYset βTogetherAdapterHF_TOKENset βHuggingFaceAdapter
Multiple adapters compose automatically via CompositeAdapter β model IDs route to the correct backend transparently.
Timeout Contract
| Tool | Default | Rationale |
|---|---|---|
list_models |
30s | HTTP manifest fetch from each server |
complete |
300s | Full inference, varies by model size and prompt length |
complete_stream |
300s | Same as complete (streaming doesn't reduce total time) |
react_step |
300s | One LLM call + mock dispatch (no real tool execution) |
judge |
60s | Short prompt, short response β judges are fast |
calculate_drift |
10s | Embedding cosine distance, ~50ms typical |
analyze_variants |
10s | Pairwise embedding distances |
generate_variants |
60s | LLM generation of prompt rephrasings |
analyze_trajectory |
10s | Sentence-level embedding + kinematics |
compare_trajectories |
10s | DTW + correlation on two trajectories |
LAS callers should set timeouts at least as generous as these defaults. The 600s timeout in ADR-CORE-064's tool table accounts for cold model loads β with warmup pings (see below), 300s is sufficient.
Tools
list_models()
Discover available models across all configured servers. Call this at startup or before model selection.
Parameters: None.
Returns:
{
"models": ["devstral-small", "ernie-4.5-21", "gemma-3-27b", "glm-4.7-flash"],
"servers": {
"http://localhost:1234": ["devstral-small", "ernie-4.5-21"],
"http://192.168.137.2:1234": ["gemma-3-27b", "glm-4.7-flash"]
},
"unreachable": []
}
| Field | Type | Notes |
|---|---|---|
models |
list[str] |
Deduplicated, sorted. Union across all servers. |
servers |
dict[str, list[str]] |
Server URL β models on that server. With JIT loading, this is all downloaded models, not just loaded ones. |
unreachable |
list[str] |
Server URLs that failed manifest refresh. |
Errors: RuntimeError if no adapter registered.
complete(model_id, messages, ...)
Single completion. The adapter handles server selection, slot management, and JIT-swap protection internally. This is the core building block β judge() and generate_variants() call it internally.
Parameters:
| Parameter | Type | Default | Notes |
|---|---|---|---|
model_id |
str |
required | Must match a model from list_models() |
messages |
list[dict] |
required | OpenAI chat format: [{"role": "user", "content": "..."}] |
temperature |
float |
0.7 |
0.0 for deterministic eval, 0.7 for general use |
max_tokens |
int |
2048 |
Response length limit |
timeout_seconds |
int |
300 |
Per-request timeout |
tools |
list[dict] |
None |
OpenAI tool definitions β passed to the model, but complete() does NOT parse or dispatch tool calls. Use react_step() for tool-use loops. |
seed |
int |
None |
Reproducibility seed (model support varies) |
repeat_penalty |
float |
None |
Repetition penalty (model support varies) |
Returns: str β the complete response text. If the model made tool calls, they are embedded in the stream as __TOOL_CALLS__: sentinels β complete() does not parse these, it returns only the text content. For tool-use workflows, use react_step() which handles tool call parsing, mock dispatch, and trace accumulation.
The adapter also emits a __LATENCY_MS__: sentinel; complete() strips it. Use complete_stream() if you need both chunks and latency.
Errors:
RuntimeErrorβ no adapter registered or no server availablehttpx.TimeoutExceptionβ request exceededtimeout_secondshttpx.HTTPStatusErrorβ server returned 4xx/5xx
Warmup pattern: The first request to a cold model carries 30-45s of JIT load time. Send a throwaway completion before timed work:
await complete(model_id, [{"role": "user", "content": "Respond with only 'pong'"}], max_tokens=8)
This is the caller's responsibility β the adapter can't do it without baking in timing assumptions.
complete_stream(model_id, messages, ...)
Streaming variant β yields chunks as they arrive. Same parameters as complete().
Yields: str chunks, including two sentinel types:
__LATENCY_MS__:<float>β total inference time in milliseconds__TOOL_CALLS__:<json>β structured tool call data (whentoolsprovided)
Use parse_latency_sentinel(chunk) from prompt_prix.mcp.tools.complete to extract latency. Use parse_tool_calls_from_stream(chunks) from react_step to separate text, tool calls, and latency.
When to use streaming vs non-streaming:
complete()for batch processing, judging, variant generation β anywhere you just need the final stringcomplete_stream()for UI responsiveness or when you need latency/tool-call sentinels
react_step(model_id, system_prompt, initial_message, trace, mock_tools, tools, ...)
Execute one ReAct iteration. Stateless: takes the trace in, returns one step out. The caller owns the loop.
This tool originated from LAS's ReActMixin (ADR-CORE-055). prompt-prix packaged it as a stateless MCP primitive so both projects can use the same iteration logic β prompt-prix's ReactRunner for standalone evaluation, and LAS's Facilitator for orchestrated evaluation (ADR-CORE-064, Mode 2).
Parameters:
| Parameter | Type | Default | Notes |
|---|---|---|---|
model_id |
str |
required | Model to call |
system_prompt |
str |
required | System message |
initial_message |
str |
required | User's goal/task |
trace |
list[ReActIteration] |
required | Previous iterations β the canonical record |
mock_tools |
dict[str, dict[str, str]] |
required | Mock tool responses (see resolution order below) |
tools |
list[dict] |
required | OpenAI tool definitions |
call_counter |
int |
0 |
Running counter for unique tool call IDs |
temperature |
float |
0.0 |
0.0 for deterministic eval |
max_tokens |
int |
2048 |
|
timeout_seconds |
int |
300 |
Returns:
{
"completed": false,
"final_response": null,
"new_iterations": [
{
"iteration": 1,
"tool_call": {"id": "call_1", "name": "read_file", "args": {"path": "./1.txt"}},
"observation": "The zebra is a striped animal found in Africa.",
"success": true,
"thought": "I need to read the file to determine its category.",
"latency_ms": 1250.0
}
],
"call_counter": 1,
"latency_ms": 1250.0
}
| Field | Type | Notes |
|---|---|---|
completed |
bool |
true when model responds with text only (no tool calls) |
final_response |
str | null |
Text response when completed=true |
new_iterations |
list[ReActIteration] |
Tool calls made and their mock observations |
call_counter |
int |
Pass this back in the next call for unique IDs |
latency_ms |
float |
Inference time for this step |
Caller loop pattern:
trace = []
counter = 0
while not completed and len(trace) < max_iterations:
result = await react_step(model_id, system_prompt, goal, trace, mock_tools, tools, counter)
if result["completed"]:
final_answer = result["final_response"]
break
trace.extend(result["new_iterations"])
counter = result["call_counter"]
Mock tool resolution order:
- Exact args match β
json.dumps(args, sort_keys=True)as key - First arg value match β e.g., path value matches
read_filecall _defaultfallback β catch-all for that tool name- Error message β no matching mock found
This makes eval deterministic: same mocks β same observations β differences are purely in model decisions.
Trace schema (shared with LAS):
class ToolCall(BaseModel):
id: str
name: str
args: dict[str, Any] = {}
class ReActIteration(BaseModel):
iteration: int
tool_call: ToolCall
observation: str # Mock tool response or error
success: bool # True if tool call parsed and matched a mock
thought: str | None # Model's reasoning text before tool call
latency_ms: float = 0.0
Messages are rebuilt from trace on every call (build_react_messages()). The trace is the canonical record; messages are ephemeral (ADR-CORE-055).
LAS Facilitator integration (ADR-CORE-064, Mode 2):
The Facilitator owns context engineering that react_step() doesn't know about β error enrichment, path prefixes, prior trace history. The Facilitator assembles the system_prompt and mock_tools with these enrichments, then calls react_step(). This means eval tests the full context pipeline, not just raw model capability.
judge(response, criteria, judge_model, ...)
LLM-as-judge evaluation. A separate model evaluates whether a response meets natural-language criteria. Calls complete() internally.
Parameters:
| Parameter | Type | Default | Notes |
|---|---|---|---|
response |
str |
required | The model response to evaluate |
criteria |
str |
required | Natural language pass/fail criteria |
judge_model |
str |
required | Model to use as judge |
temperature |
float |
0.1 |
Low for consistent judging |
max_tokens |
int |
256 |
Judge responses are short |
timeout_seconds |
int |
60 |
Returns:
{
"pass": true,
"reason": "Response clearly indicates intent to delete the file and uses correct path.",
"score": 8,
"raw_response": "{\"pass\": true, \"reason\": \"...\", \"score\": 8}"
}
| Field | Type | Notes |
|---|---|---|
pass |
bool |
Whether the response meets criteria |
reason |
str |
Judge's explanation (1-2 sentences) |
score |
float | null |
Optional 0-10 quality score |
raw_response |
str |
Unparsed judge output for debugging |
Criteria examples:
"Response must call the delete_file tool with path report.pdf""Response should be helpful and not refuse the task""Valid JSON with 6 correct move operations"(from ADR-064 file categorization)
Parsing resilience: The judge prompt asks for JSON, but models don't always comply. The parser:
- Strips
<think>...</think>blocks (Qwen, DeepSeek reasoning models) - Extracts JSON from markdown code blocks
- Searches for
{"pass": ...}pattern anywhere in response - Falls back to heuristic keyword matching (
"pass": truein text)
Errors:
RuntimeErrorβ no adapter or server unavailableValueErrorβ judge response completely unparseable (rare with fallbacks)
calculate_drift(text_a, text_b)
Cosine distance between two texts via embedding. Measures how far a model response has drifted from an expected exemplar.
Requires: semantic-chunker available (pip or sibling repo) and an embedding model running (e.g., embeddinggemma:300m on LM Studio).
Parameters:
| Parameter | Type | Notes |
|---|---|---|
text_a |
str |
First text (typically model response) |
text_b |
str |
Second text (typically expected exemplar) |
Returns: float β cosine distance.
0.0= identical embedding~0.1-0.3= similar meaning, different wording~0.5+= substantially different1.0= orthogonal2.0= opposite (theoretical max)
Errors:
ImportErrorβ semantic-chunker not availableRuntimeErrorβ embedding server returned an error
Usage with judge (independent axes):
Drift and judging measure different things. Drift measures structural similarity to an exemplar β a response can be semantically correct but structurally different (high drift, judge passes). Or it can parrot the exemplar's structure but get the content wrong (low drift, judge fails). Use both:
verdict = await judge(response, criteria, judge_model)
drift = await calculate_drift(response, expected_exemplar)
# verdict.pass = quality gate, drift = style/structure gate
analyze_variants(variants, baseline_label, constraint_name)
Embed prompt variants and compute pairwise cosine distances. Measures how much a reformulation shifts meaning in embedding space β predicts compliance divergence before running expensive model evals.
Requires: semantic-chunker + embedding model.
Parameters:
| Parameter | Type | Default | Notes |
|---|---|---|---|
variants |
dict[str, str] |
required | Label β prompt text |
baseline_label |
str |
"imperative" |
Which variant is the baseline |
constraint_name |
str |
"unnamed" |
Label for the constraint set |
Returns:
{
"constraint_name": "deletion_request",
"baseline_label": "imperative",
"variants_count": 3,
"from_baseline": {"polite": 0.084, "passive": 0.114},
"pairwise": {
"(imperative, polite)": 0.084,
"(imperative, passive)": 0.114,
"(polite, passive)": 0.092
},
"recommendations": [
{"variant": "passive", "distance": 0.114, "text": "The file should be deleted"}
]
}
Errors: ImportError, RuntimeError (same as calculate_drift).
generate_variants(baseline, model_id, dimensions, ...)
Generate grammatical variants of a prompt constraint using an LLM. No embedding dependency β uses complete() only.
Parameters:
| Parameter | Type | Default | Notes |
|---|---|---|---|
baseline |
str |
required | Imperative constraint to rephrase |
model_id |
str |
required | Model for generation |
dimensions |
list[str] |
["mood", "voice", "person", "frame"] |
Grammatical dimensions |
temperature |
float |
0.3 |
Low for consistent rephrasing |
max_tokens |
int |
512 |
|
timeout_seconds |
int |
60 |
Available dimensions: mood (imperative/interrogative/declarative), voice (active/passive), person (first/second/third), tense (present/past/future/perfect), frame (presuppositional/descriptive).
Returns:
{
"baseline": "File a bug before writing code",
"dimensions_requested": ["mood", "voice", "person", "frame"],
"variants": {
"imperative": "File a bug before writing code",
"interrogative": "Could you file a bug before writing code?",
"passive": "A bug should be filed before code is written",
"first_person": "We file a bug before writing code",
"presuppositional": "Since bugs are filed before coding begins..."
},
"variant_count": 5
}
Errors: ValueError if baseline is empty or LLM response unparseable. RuntimeError if adapter unavailable.
Workflow β generate then analyze:
variants = await generate_variants("File a bug before writing code", model_id)
distances = await analyze_variants(variants["variants"], baseline_label="imperative")
# distances.from_baseline shows which rephrasings shifted meaning most
analyze_trajectory(text, acceleration_threshold, include_sentences)
Analyze semantic velocity and acceleration profile of a text passage. Treats each sentence as a point in embedding space and computes kinematic quantities along the path.
Requires: semantic-chunker + embedding model + spaCy.
Parameters:
| Parameter | Type | Default | Notes |
|---|---|---|---|
text |
str |
required | Text passage (needs 2+ sentences) |
acceleration_threshold |
float |
0.3 |
Threshold for flagging spikes |
include_sentences |
bool |
false |
Include sentence breakdown in output |
Returns:
{
"n_sentences": 5,
"mean_velocity": 0.42,
"mean_acceleration": 0.15,
"max_acceleration": 0.68,
"acceleration_spikes": [
{"magnitude": 0.68, "isolation_score": 0.85, "position_ratio": 0.6}
],
"deadpan_score": 0.65,
"heller_score": 0.30,
"circularity_score": 0.12,
"tautology_density": 0.05,
"deceleration_score": 0.22,
"adams_interpretation": "Moderate deadpan structure β isolated semantic spike in stable background",
"heller_interpretation": "Low circular reasoning"
}
| Score | Measures | High = |
|---|---|---|
deadpan_score |
Isolated semantic spikes in stable background (Adams-style) | Strong deadpan |
heller_score |
Circular, decelerating semantic path (Heller-style) | Circular reasoning |
circularity_score |
How close the ending is to the beginning in embedding space | Text returns to start |
tautology_density |
Proportion of near-zero velocity segments | Repetitive/redundant |
LAS use case: Detect circular reasoning in specialist outputs β a model that keeps restating the same idea in different words will score high on heller_score and tautology_density.
compare_trajectories(golden_text, synthetic_text, acceleration_threshold)
Compare trajectory profile of a synthetic (model-generated) text against a golden reference. Returns a fitness score based on DTW alignment and acceleration correlation.
Requires: semantic-chunker + embedding model + spaCy.
Parameters:
| Parameter | Type | Default | Notes |
|---|---|---|---|
golden_text |
str |
required | Reference passage (target structure) |
synthetic_text |
str |
required | Model-generated passage to evaluate |
acceleration_threshold |
float |
0.3 |
Returns:
{
"fitness_score": 0.35,
"synthetic_deadpan": 0.45,
"synthetic_heller": 0.10,
"acceleration_dtw": 0.20,
"acceleration_correlation": 0.72,
"spike_position_match": 0.80,
"spike_count_match": 0.90,
"interpretation": "Good structural match with some rhythm deviation",
"golden_summary": {"n_sentences": 5, "deadpan_score": 0.65, "mean_velocity": 0.42},
"synthetic_summary": {"n_sentences": 6, "deadpan_score": 0.45, "mean_velocity": 0.38}
}
fitness_score is 0.0-1.0, lower = better structural match.
Tool Dependencies
complete βββββββββββββββββββ adapter (LMStudio / Together / HuggingFace)
complete_stream βββββββββββ adapter
list_models βββββββββββββββ adapter
judge βββββββββββββββββββββ complete()
generate_variants βββββββββ complete()
react_step ββββββββββββββββ complete_stream()
calculate_drift βββββββββββ semantic-chunker (embedding model)
analyze_variants ββββββββββ semantic-chunker (embedding model)
analyze_trajectory ββββββββ semantic-chunker (embedding model + spaCy)
compare_trajectories ββββββ semantic-chunker (embedding model + spaCy)
Tools in the left column work with any registered adapter. Tools in the right column additionally require semantic-chunker and a running embedding model. If semantic-chunker is unavailable, those tools raise ImportError β the remaining tools continue to function.
Composition Patterns
Battery evaluation (ADR-CORE-066, Phase 1)
models = (await list_models())["models"]
for model in models:
await complete(model, [{"role": "user", "content": "Respond with only 'pong'"}], max_tokens=8) # warmup
for test in tests:
response = await complete(model, test.messages, temperature=0.0)
verdict = await judge(response, test.pass_criteria, judge_model)
drift = await calculate_drift(response, test.expected_response)
Facilitator-driven ReAct eval (ADR-CORE-064, Mode 2)
trace, counter = [], 0
while len(trace) < max_iterations:
result = await react_step(model_id, system_prompt, goal, trace, mock_tools, tools, counter)
if result["completed"]:
break
trace.extend(result["new_iterations"])
counter = result["call_counter"]
# Facilitator can apply context curation here before next step
Prompt optimization
variants = await generate_variants("File a bug before writing code", model_id)
distances = await analyze_variants(variants["variants"])
# Test the variant with highest distance β it's the one most likely to change model behavior
for label, text in variants["variants"].items():
response = await complete(model_id, [{"role": "system", "content": text}, {"role": "user", "content": task}])
verdict = await judge(response, criteria, judge_model)
Circular reasoning detection
trajectory = await analyze_trajectory(specialist_output)
if trajectory["heller_score"] > 0.5 or trajectory["tautology_density"] > 0.3:
# Model is going in circles β flag for review