|
Download docs/API.md from ken-jo/qev: direct link, hf CLI and curl.
- Browser
- Download file 2.86 kB
-
https://huggingface.co/ken-jo/qev/resolve/main/docs/API.md
- Command line
-
hf download hf://ken-jo/qev/docs/API.md
-
curl -L -o API.md https://huggingface.co/ken-jo/qev/resolve/main/docs/API.md
2.86 kB
| # Inference API | |
| ## Python SDK | |
| The `qev` distribution contains both the SDK and the English playground. With Python | |
| 3.12, install the release wheel or the published PyPI distribution, then: | |
| ```python | |
| import qev | |
| from pathlib import Path | |
| model = qev.load() # Download missing model files; select CUDA when available. | |
| request = qev.DecisionRequest.from_json(open("request.json", encoding="utf-8").read()) | |
| result = model.predict(request, image_root=Path(".")) | |
| print(result["answers"]) | |
| ``` | |
| The first load retrieves about 4.6 GB of pinned model files. Later loads reuse the cache. | |
| `qev.load(device="cpu", cache_dir="/path/to/hub", offline=True)` selects CPU and requires | |
| complete local files. `QEV_HOME` sets the checkpoint/cache root; `QEV_CACHE_DIR` overrides | |
| the Hub cache. The lower-level `qev.QEV.load(...)` remains available for explicit checkpoint | |
| management and retains the original inference contract. | |
| ## Commands | |
| ```sh | |
| qev playground | |
| qev download | |
| qev predict --request request.json | |
| qev serve --image-root . | |
| ``` | |
| All inference commands prepare missing model files automatically. Use `--offline` to | |
| require a complete cache. `predict` resolves image paths relative to the request JSON; | |
| `serve` resolves them under `--image-root`. The API server defaults to | |
| `http://127.0.0.1:8000`; the playground defaults to `http://127.0.0.1:7860`. | |
| See [playground installation, uv and GPU instructions](PLAYGROUND.md). | |
| ## Request and response | |
| The public request schema is `qev.DecisionRequest`. Text and one local image can be | |
| combined. Image paths are resolved under the configured image root; remote URLs are not | |
| fetched by the inference API. | |
| ```json | |
| { | |
| "state": {"text": "The customer was charged twice and asks for a refund."}, | |
| "questions": { | |
| "department": { | |
| "type": "choice", | |
| "instructions": "Which department should handle the request?", | |
| "criteria": {"billing": "Payments and refunds", "technical": "Software faults"} | |
| } | |
| } | |
| } | |
| ``` | |
| Send JSON to `POST /v1/systemone`. `score.criteria` is an ordered list of descriptions; | |
| the result is an expected zero-based level. `noul.criteria` has `true` and `false` strings. | |
| The maximum is four questions, 16 alternatives per choice/score and one image. Processed | |
| inputs exceeding the configured token budget fail instead of silently truncating. | |
| Inspect `answers`, `probabilities`, abstention fields and `usage` in the returned object. | |
| The model ID remains `Qwen/Qwen3.5-2B` in the validated contract to identify the actual base. | |
| See `src/veyra/schema.py`, `src/veyra/probability.py` and `examples/request.json`. | |
| The playground adds image-upload and resolution endpoints around the same model. A single | |
| resident worker serializes inference. It is a trusted local demo; uploads are shared within | |
| that server session. Local latency measurements do not imply concurrent serving capacity. | |