112
followers ·
413 following AI & ML interests Building the AI-native stack. Agents as infrastructure, safety as architecture, performance as plumbing. I publish the receipts: papers, datasets, demos.
Recent Activity reacted to Alihassaanmughal 's post with 🧠 5 minutes ago LLM-Based Test Oracles: Source-of-Authority Taxonomy—A Systematic Literature Review
Abstract:
Large language models (LLMs) increasingly decide whether software behaves correctly, either by writing a test oracle or by acting as one. Yet two oracles can look identical and rest on different ground: one assertion encodes a written specification, another only what the model learned in training. Prior secondary studies sort oracles by form or by technique, rarely by the property that governs how far a verdict can be trusted: where its authority comes from. This systematic literature review, reported under the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines, screens 2,436 records to 54 included studies, extended by citation searching (snowballing) to 83 in total. We read the corpus along three axes: the source of an oracle’s authority, the form it takes, and the mechanism that adjudicates it. Just over half of the corpus reaches a verdict with no specification at all. That is what lets these oracles work on code with no specification to consult, and what leaves a challenged verdict with less to fall back on. Source and mechanism cross-cut rather than coincide, so a label such as LLM-as-a-judge names how a verdict is produced, not why it should be trusted. Oracle quality is most often judged by resemblance to a known oracle rather than by whether injected faults are caught. The first question to ask of any LLM oracle is therefore what one would point to in defending its verdict. The protocol, search query, and per-study coding sheet are released.
Published in IEEE Access.
Published in: IEEE Access ( Volume: 14)
Page(s): 136142 - 136161
Date of Publication: 02 September 2026
Electronic ISSN: 2169-3536
DOI: 10.1109/ACCESS.2026.3729738
Publisher: IEEE reacted to SoulInPsyAbstract 's post with 🧠 about 1 hour ago We opened FLUX 3 Action's code expecting our own pipeline. It closes 1.5 of 5 floors.
Black Forest Labs' FLUX 3 Action is a "world action model" for the SO-101 robot arm: one diffusion process jointly denoises the next chunk of actions and the next chunk of video frames. Read as a headline, that sounded like exactly the causal-chain-first architecture our own safety pipeline argues for — action and outcome tied together in one step, not an action head bolted onto a frozen representation. So we rented an L40S on Brev and read the code, not just the model card.
Our pipeline is five floors, each depending on the one below it: Causal chain → Probability → Risk/Impact → Decision theory → Markov/Game theory. Here's what FLUX 3 Action actually has.
01
Causal chain
Present, and prioritized: video_loss_weight: 1.0 outweighs action_loss_weight: 0.5. The model is trained to get the outcome right more than the action itself — this is the real thing, not a gesture at it.
present
02
Probability
Technically present, never surfaced. It's a diffusion model — it samples from a distribution by construction. Nothing reads that distribution back out as an uncertainty number a decision could use. The probability exists inside the math and dies there.
hollow
03
Risk / impact
Absent. The model card says so itself: "nothing bounds joint velocity, force, workspace." Not hidden — just not built.
absent
04
Decision theory
Absent. No gate. The model executes 32 actions per chunk; there is no threshold at which it would stop.
absent
05
Markov / game theory
Not applicable at this scope — a single robot arm with no adversary or multi-round state.
n/a
The closure isn't "their floors 1–2 are weaker than ours." They're not — floor 1 here is arguably cleaner than most causal-chain implementations we've seen, because the loss weighting makes the priority explicit in the training objective itself, not just in a README.
A model with two good, real, working floors behaves identically to a model View all activity Organizations