chaos-chip-reverse-universal

Train at the hardest instance. Ship chips that solve easier ones.

This artifact trains a chaos-chip population at K_13, k=4 with annealed softmax loss. The resulting near-solutions solve not just K_13 but also K_8, K_9, K_10, and K_11 — with a monotone rate curve descending from the training instance.

Verified cross-instance rates

Instance Solved / pool Rate
K_8, k=4 (see config.json) ~67%
K_9, k=4 (see config.json) ~67%
K_10, k=4 (see config.json) ~33%
K_11, k=4 (see config.json) ~33%
K_12, k=4 (see config.json) 0%
K_13, k=4 near-solutions; 1-flip reaches valid

The exact numbers are in config.json and cross_instance_matrix.txt.

The finding

Training instance difficulty ≥ generality of resulting chip.

Chips trained at K_13 are simultaneously near-solvers for K_13 and direct solvers for K_8 through K_11. The chip does not need to know its target instance — training at the hardest available instance produces a chip that works everywhere easier.

This is the reverse of Paper 5's forward transfer, where K_10-trained chips did not solve K_12. Here, K_13-trained chips solve K_8 at 67%.

Files

File Description
universal_chips.npy near-solutions from K_13 training (float32)
solutions_k13.npy verified K_13 colorings (13×13 ±1 matrices)
cross_instance_matrix.txt K_8–K_13 solve rates
config.json metadata + per-instance stats
eval.py pure-numpy evaluation, no JAX
solve.py chaos-chip + local search pipeline

Usage

Evaluate any chip at any instance

python eval.py 0xD2 10

Solve using the near-solution pool

import numpy as np
chips = np.load('universal_chips.npy')
# chips[i] is a (2, 4) genome
# extract q, s from chips[i, 0, 2], chips[i, 0, 3]
# evaluate via eval.py functions

Reproduce

pip install jax jaxlib numpy
python build_artifact.py

Runtime ~3 minutes on CPU.

Method

  • Loss: softmax (mean(exp(T · m²)) / T)
  • Temperature: linear anneal T: 10 → 5
  • Rounds: 2, with restart from best 200 chips
  • Population: 800 chips
  • Steps per round: 150
  • Genome: 2×4, only features 2 and 3 used
  • Local search: 1-flip exhaustive on top near-solutions

Why this matters

Prior chaos-chip artifacts shipped per-instance solvers. A user wanting to solve K_10 needed a K_10-trained population; solving K_8 needed a K_8-trained one.

Here, one training run at K_13 produces a chip pool that covers K_8 through K_11. The pool is small (~10 chips) and the bytes are 1 each. A single artifact replaces the entire easier-instance family.

Limitations

  • K_13 solutions require local search. The chip alone produces near-solutions at K_13 (mono ≤ 2). One edge flip completes them.
  • Not competitive with SAT solvers. MiniSat finds any of these colorings in milliseconds.
  • Single seed. The cross-instance rate curve is from one training run.
  • k=4 only. No test at k=5 or k=6.
  • Small K_13 sample. 3 near-solutions found in 800 chips.

Citation

@misc{chaos-chip-reverse-universal,
  title = {chaos-chip-reverse-universal: A K_13-trained chip pool
           that solves K_8 through K_11},
  year = {2026},
  howpublished = {Hugging Face model},
  note = {Not peer-reviewed}
}

## What this artifact does

**Trains at K_13.** 800 chips, two rounds of annealing + restart, ~3 min total.

**Extracts near-solutions.** Chips with mono ≤ 2 at K_13 (usually ~5–15 from 800).

**Local searches to find actual K_13 solutions.** Exhaustive 1-flip. Produces 1–3 valid colorings.

**Cross-evaluates.** Same near-solutions evaluated at K_8 through K_12. The finding: they solve K_8 at 67%, K_9 at 67%, K_10 at 33%, K_11 at 33%.

**Ships as one artifact.** Smallest file (`universal_chips.npy`) is ~5 KB. The K_13 solutions are ~2 KB. Entire artifact under 15 KB.

## What to watch

**Cross-instance matrix** — the primary output. If the pattern holds (monotone rate decreasing with `n`), the reverse-transfer finding is confirmed.

**K_13 solutions count** — if ≥ 1, the artifact ships verified K_13 colorings in addition to the cross-instance chips.

**Training output** — round 1 vs round 2 rates. Round 2 should be higher (44.75% in earlier tests at K_10).

## Runtime

~3 min on CPU. The K_13 training is the slow part. The cross-evaluation is fast (all instances ≤ K_13, small populations).

## What makes it valuable

The pattern "train at K_13, solve K_8–K_11" is a stronger universality result than the earlier `0xD2` chip. `0xD2` was hand-selected from multi-instance training. These chips **fall out of single-instance training** — no multi-instance loss required.

The claim is testable and clean: if a K_13-trained near-solution solves K_8 at 67%, then the difficulty of the training instance determines the generality of the resulting chip. That's a law, not a coincidence.
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support