File size: 2,856 Bytes
6cb324a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
# Inference API

## Python SDK

The `qev` distribution contains both the SDK and the English playground. With Python
3.12, install the release wheel or the published PyPI distribution, then:

```python
import qev
from pathlib import Path

model = qev.load()  # Download missing model files; select CUDA when available.
request = qev.DecisionRequest.from_json(open("request.json", encoding="utf-8").read())
result = model.predict(request, image_root=Path("."))
print(result["answers"])
```

The first load retrieves about 4.6 GB of pinned model files. Later loads reuse the cache.
`qev.load(device="cpu", cache_dir="/path/to/hub", offline=True)` selects CPU and requires
complete local files. `QEV_HOME` sets the checkpoint/cache root; `QEV_CACHE_DIR` overrides
the Hub cache. The lower-level `qev.QEV.load(...)` remains available for explicit checkpoint
management and retains the original inference contract.

## Commands

```sh
qev playground
qev download
qev predict --request request.json
qev serve --image-root .
```

All inference commands prepare missing model files automatically. Use `--offline` to
require a complete cache. `predict` resolves image paths relative to the request JSON;
`serve` resolves them under `--image-root`. The API server defaults to
`http://127.0.0.1:8000`; the playground defaults to `http://127.0.0.1:7860`.
See [playground installation, uv and GPU instructions](PLAYGROUND.md).

## Request and response

The public request schema is `qev.DecisionRequest`. Text and one local image can be
combined. Image paths are resolved under the configured image root; remote URLs are not
fetched by the inference API.

```json
{
  "state": {"text": "The customer was charged twice and asks for a refund."},
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which department should handle the request?",
      "criteria": {"billing": "Payments and refunds", "technical": "Software faults"}
    }
  }
}
```

Send JSON to `POST /v1/systemone`. `score.criteria` is an ordered list of descriptions;
the result is an expected zero-based level. `noul.criteria` has `true` and `false` strings.
The maximum is four questions, 16 alternatives per choice/score and one image. Processed
inputs exceeding the configured token budget fail instead of silently truncating.

Inspect `answers`, `probabilities`, abstention fields and `usage` in the returned object.
The model ID remains `Qwen/Qwen3.5-2B` in the validated contract to identify the actual base.
See `src/veyra/schema.py`, `src/veyra/probability.py` and `examples/request.json`.

The playground adds image-upload and resolution endpoints around the same model. A single
resident worker serializes inference. It is a trusted local demo; uploads are shared within
that server session. Local latency measurements do not imply concurrent serving capacity.