agent-harness / docs /RUNTIME_BASELINE.md
cuber12's picture
Publish agent harness research code and paper artifacts
d61821a verified
|
Raw
History Blame Contribute Delete
2.26 kB

Observed local runtime baseline

This is an initial environment observation, not an experiment result. The runner must repeat discovery and store the returned metadata in every recorded run.

Observed on 2026-07-17 through LM Studio at http://127.0.0.1:1234:

Property Observed value
Canonical study model Qwen3.6-35B-A3B
OpenAI-compatible inference key qwen/qwen3.6-35b-a3b
Selected variant qwen/qwen3.6-35b-a3b@4bit
Runtime format MLX
Quantization 4-bit
Architecture reported by LM Studio qwen3_5_moe
Loaded context length 262,144 tokens
Loaded parallelism 4
Maximum context reported 262,144 tokens
Reasoning capability/default off or on / on
Tool-use capability true
Model size reported 20,429,364,306 bytes
LM Studio CLI commit efce996

The same inventory exposed three embedding models. EMB001 selects the following available runtime:

Property Observed value
Display name Qwen3 Embedding 0.6B
LM Studio key text-embedding-qwen3-embedding-0.6b
Format and quantization GGUF Q8_0
Parameter label 0.6B
Loaded context 8,192 tokens
Maximum context 32,768 tokens
Model size 639,150,592 bytes
Vector dimension 1,024

An endpoint probe embedded two different code snippets. It returned two different 1,024-dimensional vectors with L2 norms 0.99999996 and 0.99999997. LM Studio reported zero prompt and total tokens for these embedding calls, so embedding token counts must be measured independently rather than interpreting that API usage field as real zero cost.

Both /v1/models and /api/v1/models returned without error. A semantic chat-completions probe returned visible MODEL_OK with finish_reason=stop. That request used 20 prompt tokens and 152 completion tokens, of which 147 were reported as reasoning tokens. This overhead is evidence that reasoning tokens must be logged separately rather than treated as free or ignored.

The current study config pins the inference key, variant, format, quantization, loaded context length, and reasoning default. The exact LM Studio application version and host/hardware inventory still need to be captured before the pilot is frozen.