Guardrails and observability are real, and I think they answer a different question than the one I asked.
Observability tells you what the agent did, after the fact. A human can audit the trace. Calibration is whether the agent's own signal tracks its own correctness, in the moment, with nobody reading anything. You can have excellent observability sitting on top of an agent that is confidently wrong, and then what you have built is a very good recording of a failure that nobody got warned about.
The reason I keep poking at this: self-reported confidence is not the internal state. There is a result from last week (2607.08046) where activation probes calibrate better than what the model actually says, and ablating the evidence flips the forecast while the reasoning trace comes out unchanged. The trace looked fine. The trace was not the mechanism.
Which is also why "come and test it" does not quite settle it. On OSWorld I can grade the agent myself, so I do not need it to tell me. It is the ungradeable tasks where a could-not-verify signal is the entire product, and those are exactly the ones I cannot check by testing.
So, concretely: on a task where it failed, does Holo3 say so? Is there a confidence number in the API response, and has it been scored against outcomes?