Instructions to use Harry19081/Wald-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Harry19081/Wald-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Harry19081/Wald-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Harry19081/Wald-4B") model = AutoModelForCausalLM.from_pretrained("Harry19081/Wald-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Harry19081/Wald-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Harry19081/Wald-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Harry19081/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Harry19081/Wald-4B
- SGLang
How to use Harry19081/Wald-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Harry19081/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Harry19081/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Harry19081/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Harry19081/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Harry19081/Wald-4B with Docker Model Runner:
docker model run hf.co/Harry19081/Wald-4B
Download docs/api.md from Harry19081/Wald-4B: direct link, hf CLI and curl.
- Browser
- Download file 6.09 kB
-
https://huggingface.co/Harry19081/Wald-4B/resolve/main/docs/api.md
- Command line
-
hf download hf://Harry19081/Wald-4B/docs/api.md
-
curl -L -o api.md https://huggingface.co/Harry19081/Wald-4B/resolve/main/docs/api.md
Wald-Q4B decision API
Wald-Q4B's server, wald-serve, exposes one decision endpoint, POST /v1/systemone. Its request and answer shapes follow TypeSafe's /v1/systemone format, so clients written for Jev can point at a self-hosted Wald server. Wald is independent and is not affiliated with TypeSafe AI.
Start the server with ./run.sh "$PWD" in the downloaded model directory (model card, RUNBOOK.md). It listens on port 8000 and does not check API keys.
Endpoints
| Method and path | Purpose |
|---|---|
POST /v1/systemone (also POST /) |
Answer one or more typed questions about a state |
GET /health |
Effective policy: effort, gate, thought budget, prompt format, context limit |
GET /v1/models |
The served model name |
Request
| Field | Type | Meaning |
|---|---|---|
state |
string, object, array or null | What the decision is about: a message, a conversation, a document, an agent trace. Objects and arrays are flattened to text with their field names kept. |
questions |
object, at least one entry | Question id → question. Every question is answered about the same state. |
effort |
string, optional | none, low, medium, high or high-k2 … high-k8. Overrides the server default for this request. |
model |
string, optional | Accepted for client compatibility; the server answers with the model it serves. |
Each question has a type, optional instructions (any JSON, usually a sentence) and criteria:
type |
criteria |
Answer |
|---|---|---|
choice |
Object of 1–255 options: key → description (description may be null) | choice (the most probable key), probabilities (key → probability), confidence |
noul |
Optional object with true and/or false descriptions |
noul: the probability of yes |
score |
Array of 1–255 ordered levels | score (expected level index), probabilities (level index → probability), confidence |
Every answer also carries mode: A for a one-pass read, B for a read after a thought, K for grouped (knockout) reading of wide option sets.
Prompts longer than the context limit (131,072 tokens by default) are rejected with HTTP 422, never truncated. A malformed request or unknown effort returns HTTP 400.
Example: route a tool call
curl http://localhost:8000/v1/systemone \
-H 'Content-Type: application/json' \
-d '{
"state": {
"user": "What will the weather be in Lisbon tomorrow afternoon?",
"tools_available": ["web_search", "weather_api", "calendar"]
},
"effort": "none",
"questions": {
"tool": {
"type": "choice",
"instructions": "Which tool should the agent call next?",
"criteria": {
"web_search": "General web search",
"weather_api": "Forecast for a city and time",
"calendar": "Read or create calendar events",
"none": "Answer directly without a tool"
}
}
}
}'
Response shape (the numbers here are illustrative, not a measured output):
{
"model": "wald-4b",
"answers": {
"tool": {
"type": "choice",
"choice": "weather_api",
"probabilities": {"web_search": 0.04, "weather_api": 0.94, "calendar": 0.0, "none": 0.02},
"confidence": 0.92,
"mode": "A"
}
},
"usage": {"input_tokens": 142, "output_tokens": 0},
"latency_s": 0.03
}
For choice, confidence rescales the top probability so that 0 means uniform and 1 means certain: (p_top − 1/K) / (1 − 1/K) for K options.
Example: decide whether to ask the user
Ask several questions about one state in a single request. Each question is read separately.
{
"state": "User: book me a table for Friday",
"effort": "medium",
"questions": {
"specific_enough": {
"type": "noul",
"instructions": "Is the request specific enough to act on without asking a follow-up question?",
"criteria": {"true": "Enough detail to act", "false": "Needs clarification first"}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this request?",
"criteria": ["low", "medium", "high"]
}
}
}
Response shape (illustrative numbers):
{
"model": "wald-4b",
"answers": {
"specific_enough": {"type": "noul", "noul": 0.08, "mode": "B"},
"urgency": {"type": "score", "score": 1.1, "probabilities": {"0": 0.15, "1": 0.6, "2": 0.25}, "confidence": 0.6, "mode": "A"}
},
"usage": {"input_tokens": 310, "output_tokens": 96},
"latency_s": 0.6
}
A typical policy: act when noul is above a high threshold, ask the user when it is below a low one, and escalate the cases in between. Set both thresholds on validation data from your own task.
Choosing an effort
| Effort | Use it for |
|---|---|
none |
Lowest latency. One pass, no generated tokens. |
low / medium |
Think only when the top initial probability is below 0.5 / 0.7. |
high |
Think on every eligible question. The published Decision Index score uses this setting. |
high-k2 … high-k8 |
Several thoughts, averaged. Slowest. |
Thinking applies to questions with 2–26 options when the context has room; otherwise the one-pass answer is returned.
Python client
import requests
r = requests.post("http://localhost:8000/v1/systemone", json={
"state": "Refund request for order 1182; the parcel arrived damaged.",
"effort": "none",
"questions": {"route": {"type": "choice", "criteria": {
"refunds": "Refunds and returns", "shipping": "Delivery problems", "other": "Anything else"}}},
}, timeout=30)
answer = r.json()["answers"]["route"]
print(answer["choice"], answer["probabilities"])
Source: server/src/wald_serve/wire.py (request shapes) and server.py (endpoints).