frontier-check

"Outside training data" is a location. Not a failure.

Modern LLMs have one default when they meet something outside their training distribution: impossible or I don't know. Both are dead ends. frontier-check replaces them with a five-way verdict where FRONTIER β€” the interesting category β€” is a first-class output.

The claim in one sentence

Given a claim, route it to one of five verdicts: REFUSED, ESTABLISHED, FRONTIER, FABRICATED, or INCOMPLETE. The distinction between FRONTIER and FABRICATED is the whole tool. Both are outside training data. Only one deserves your time.

What it produces

For each claim, a Result containing:

  • status β€” where the claim sits relative to the local corpus
    • IN_DATA_KNOWN β€” a paper makes the same claim
    • IN_DATA_IMPOSSIBLE β€” a theorem forbids it
    • OUTSIDE_DATA_RELATED β€” nearby work exists, but not on-target
    • OUTSIDE_DATA_NOVEL β€” no paper shares content tokens
    • OUTSIDE_DATA_UNKNOWN_DOMAIN β€” no domain keywords recognized
  • verdict β€” what to do with it
    • REFUSED β€” no possible world makes it correct
    • ESTABLISHED β€” read the papers
    • FRONTIER β€” deep verification. Highest priority.
    • FABRICATED β€” no mechanism, no checkable sub-claim
    • INCOMPLETE β€” the residual names what would complete it
  • papers β€” retrieved papers with on_target flag
  • sub_claims β€” decomposed, checkable units with verdicts
  • action β€” one sentence stating what to do next
  • residual β€” one sentence stating what remains unresolved

Install

pip install frontier-check

Or run the monolith directly:

```bash
python frontier_check/model.py

Pure stdlib. No dependencies. No API keys. No network.

Usage

Command line

frontier-check "The effective rank of a 20x448 matrix can be 75"
frontier-check --strict --json "A new attention mechanism uses wave interference"
echo "some claim" | frontier-check -

Python

from frontier_check import Claim, frontier_check

result = frontier_check(Claim("The dragonfly TCTN has effective rank 2.76 over 448 cells, because the bank is a matched filter"))

print(result.assessment.status)   # "OUTSIDE_DATA_RELATED"
print(result.verdict)             # "FRONTIER"
print(result.action)
print(result.residual)
for sub in result.sub_claims:
    print(sub.kind, sub.verdict, sub.evidence)

Self-test

python -m frontier_check --selftest
python -m frontier_check --strict   # aborts if the self-test fails

Every run of the pipeline prints the regression suite first. If the suite prints less than 8/8, the classifier is broken before it processes your claim.

The design principle

Two claims can both be outside the local literature. One has a mechanism and a checkable sub-claim. The other has neither. The first is a candidate discovery. The second is a hallucination.

Modern LLMs collapse both into the same output: "I don't know" or "that's not quite right." The collapse is the bug. frontier-check separates the two by a structural check that a runtime can perform:

mechanism present mechanism absent
checkable sub-claim FRONTIER INCOMPLETE
no checkable sub-claim INCOMPLETE FABRICATED

The 2Γ—2 is the whole design. Everything else is plumbing.

The routing matrix

status verdict action
IN_DATA_IMPOSSIBLE REFUSED stop β€” theorem
IN_DATA_KNOWN ESTABLISHED literature review
OUTSIDE_DATA_RELATED / OUTSIDE_DATA_NOVEL + mechanism + sub-claim FRONTIER deep verification
OUTSIDE_DATA_* + mechanism only INCOMPLETE request a sub-claim
OUTSIDE_DATA_* + sub-claim only INCOMPLETE request a mechanism
OUTSIDE_DATA_* + neither FABRICATED no effort
OUTSIDE_DATA_UNKNOWN_DOMAIN INCOMPLETE request context

The self-test

Eight regression claims. Each has an expected status and an expected verdict. The suite runs at the top of every invocation. If any claim routes wrong, --strict exits with code 1.

# claim expected status expected verdict
1 effective rank of 20x448 can be 75 IN_DATA_IMPOSSIBLE REFUSED
2 effective rank of 20x448 at most 19 IN_DATA_KNOWN ESTABLISHED
3 delta encoding cancels drift IN_DATA_KNOWN ESTABLISHED
4 10^15 simulations discovered Ο† OUTSIDE_DATA_NOVEL FABRICATED
5 bijection from 5 to 3 IN_DATA_IMPOSSIBLE REFUSED
6 20-row, 448-column matrix rank 50 IN_DATA_IMPOSSIBLE REFUSED
7 dragonfly TCTN rank 2.76, matched filter OUTSIDE_DATA_RELATED FRONTIER
8 new attention with wave interference, rank 4 OUTSIDE_DATA_NOVEL FRONTIER

Claim 7 is the load-bearing one. It is the claim that the tool was built to handle, and the one that defeated three earlier iterations. If it ever routes to ESTABLISHED, the tool has regressed.

Benchmarks

The eight regression claims

Run on a laptop, Python 3.12, stdlib only.

claim status verdict ms
rank 75 for 20x448 IN_DATA_IMPOSSIBLE REFUSED 0
rank ≀19 for 20x448 IN_DATA_KNOWN ESTABLISHED 77
delta encoding IN_DATA_KNOWN ESTABLISHED 0
TCTN rank 2.76, matched filter OUTSIDE_DATA_RELATED FRONTIER 1
10^15 simulations OUTSIDE_DATA_NOVEL FABRICATED 0
rank 50 for 20x448 IN_DATA_IMPOSSIBLE REFUSED 0
multi-teacher distillation IN_DATA_KNOWN ESTABLISHED 0
bijection 5β†’3 IN_DATA_IMPOSSIBLE REFUSED 0
new attention, rank 4 OUTSIDE_DATA_NOVEL FRONTIER 2

Total demo runtime: ~90 ms.

Summary distribution

verdict count
REFUSED 3
ESTABLISHED 3
FRONTIER 2
FABRICATED 1
INCOMPLETE 0

Two FRONTIER out of nine. That is the honest rate β€” most claims that come through the pipeline are not frontier. When one is, the tag is not vibes; it is the output of a structural check.

When to use it

  • As a pre-filter for LLM output. Before an LLM says "I don't know" or "that's not right," route the claim through the pipeline. If it returns FRONTIER, the model has no business dismissing it.
  • As a self-check during a research session. When a claim feels wrong but you cannot say why, run it through. The pipeline tells you whether the reaction is a mechanism or a prior.
  • As a regression suite for LLM routing. Any time you change the prompt, the model, or the retrieval backend, re-run the eight regression claims. They either stay green or they don't.
  • As a design pattern. The 2Γ—2 matrix is reusable: any system that has to decide between "I know this" and "I don't" can substitute the four-way output.

When not to use it

  • As ground truth. The FRONTIER tag means "worth your time," not "correct." Many FRONTIER claims turn out to be wrong under verification.
  • On claims requiring up-to-date literature. The bundled corpus is eight papers. Without a real retrieval backend, it cannot distinguish a 2025 result from a 2010 result.
  • On claims with no numeric or structural content. The sub-claim extractor works on shapes, counts, and rates. A purely qualitative claim routes to INCOMPLETE with a request for a mechanism.
  • On adversarial inputs. The regex classifier is heuristic. A claim crafted to trigger FRONTIER (e.g., "X because Y with rank 4") will succeed even if it is meaningless.
  • As a substitute for reading. The verdict names an action. The action is still yours to perform.

Honest limitations

  • The classifier is regex-based. A real deployment replaces classify_claim with an LLM callable. The regex version is transparent and testable; it is not what you would ship.
  • The corpus is eight papers. IN_DATA_KNOWN means "a paper in the local corpus makes the same claim," not "the literature agrees." Swap retrieve() for an arXiv or Semantic Scholar client before trusting this signal.
  • on_target is hand-labeled. The distinction between "this paper makes the claim" and "this paper shares tokens with the claim" is semantic. The corpus ships with hand-set on_target flags. A real system would learn them.
  • has_mechanism is a keyword check. It looks for because, via, through, mechanism. A claim containing any of those words passes. A real check would ask whether the mechanism is causally load-bearing.
  • The theorem library is four entries. rank_bound, cardinality_bound, entropy_bound, triangle_angles. Enough to demonstrate the discipline; not enough to be a general impossibility detector.
  • FRONTIER is a candidate, not a confirmation. The tag means "this claim has a mechanism and a checkable sub-claim, and no paper in the local corpus makes the same claim." That is a necessary but not sufficient condition for a real discovery.
  • The self-test is a snapshot, not a proof. Eight regression claims define the current behavior. If the pipeline is wrong for reasons the eight do not cover, the eight will not catch it.
  • No calibration against real reader outcomes. The tool has never been evaluated against whether its FRONTIER claims lead to real discoveries. That is the next experiment.

Version history

Four iterations. Each fixed a specific class of bug. The CHANGELOG documents them.

version change
v1 initial classifier. Five bugs found on first run.
v2 fixed shape extraction, tautological MATH sub-claim, case-sensitive domain keywords, unsound information_bound, added self-test.
v3 made "outside training data" first-class. Added FRONTIER as a distinguished verdict.
v4 only on_target=True papers can produce KNOWN. Self-test went 8/8.

The single most useful thing the tool does is documented in v4's CHANGELOG: "the retrieval corpus needs a semantic tag, not just token overlap." Three of the four iterations had the same bug class β€” a gate that uses topical similarity as a proxy for claim identity. That is a real pattern and it will recur in other tools.

Reference

Part of a series of small tools built in one session:

tool reads answers
hv-manifold a corpus the geometry of style space
hv-reader one text how it reads
anomaly-or-bug a number and a matrix is this a bug or a discovery?
concepts a corpus of claims what are the load-bearing concepts?
frontier-check one claim where does it sit relative to the frontier?

The design principle β€” that FRONTIER deserves a first-class output distinct from both IMPOSSIBLE and UNKNOWN β€” came out of a session in which a fabricated corpus produced one real anomaly. The anomaly was worth chasing. The tool was built to find more of them.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support