STOK-SpatialNet-v3: A Reward-Free Spatial-Attention Amortizer

STOK-SpatialNet-v3 is a compact, multi-task spatial neural network for reward-free reinforcement learning and simultaneous-turn planning.

Instead of regressing a single scalar reward or value function, it computes the quantities of the State-Time Option Kernel (STOK) framework (Ringstrom & Schrater, 2025):

  • Feasibility ($\hat{\kappa}$): Probability of reaching a target goal without constraint violation.
  • Time Urgency ($\hat{\eta}$): Arrival time distribution $1 / (1 + \mathbb{E}[t_f])$.
  • Valence ($\Delta\hat{E}$): Signed change in action-space empowerment (preserving future degrees of freedom).
  • Survivability ($\hat{\text{lose}}$): Predicted adversarial lose rate.
  • Theory of Mind: Softmax distributions predicting the opponent's active tactical concept and goal family.

Architecture: spatial-attn2

The network operates on a 2D spatial field with relational goal conditioning:

Board Cells (T tiles x 39 ch) ──► Embed (39 β†’ 32) + ReLU
                                         β”‚
                                         β–Ό
                         Tile ⟷ Tile Self-Attention (Residual)
                                         β”‚
                                         β–Ό
Relational Goal (8 dims) ───────► Goal-Query Cross-Attention
                                         β”‚
                                         β–Ό
Context Vector (32) ─────┐
High-level Features (80) ┼──────► Dense Trunk (138 β†’ 96) + ReLU
Intent History (18) β”€β”€β”€β”€β”€β”˜               β”‚
                                         β–Ό
                               37 Output Heads (96 β†’ 37)

Input Channels (39 per tile)

  1. Geometry: Normalized absolute coordinates, viewer- and opponent-relative offsets, Chebyshev distances.
  2. Occupancy & Stats: Self/ally/enemy presence, normalized HP.
  3. Terrain & Content: Substrate, walls, rocks, hazardous elements, blocking flags.
  4. Path-Aware Reachability: BFS arrival steps for both agents under dynamic obstacle rules.
  5. Plan Grounding: Footprints of candidate movement, strikes, and terrain alterations.

37 Multi-Task Output Heads

  • Heads 0–3 (The Core Quartet): Feasibility $\hat{\kappa}$, urgency $\hat{\eta}$, signed valence $\Delta\hat{E}$, survivability $\hat{\text{lose}}$.
  • Heads 4–8 (Goal-Setting Feasibilities): Predicted feasibility for 5 goal families: Strike, Engage, Evade, Turtle, Zone.
  • Heads 9–11 (Arrival ETA Histogram): Softmax distribution over $[\text{kill@t1}, \text{kill@t2}, \text{not in window}]$.
  • Heads 12–16 (Opponent Goal Prior): Predicted distribution over which goal the opponent will execute this turn.
  • Heads 17–18 (Decomposed Valence): Own option-capacity growth vs denial of opponent option-capacity.
  • Heads 19–26 (Tactical Concept Tagging): Distribution over 8 human-annotated tactical concepts (trap, spacing, punish, bait, coverage, conversion, denial, tempo).
  • Heads 27–34 (Theory of Mind): Predicted tactical concept the adversary is currently playing.
  • Heads 35–36 (Trajectory Valence): Long-horizon realized empowerment flow ($V_{\text{traj}}$, $V_{\text{past}}$).

Quickstart (Zero External Dependencies)

Requires only standard numpy. No PyTorch or C++ runtime required.

import numpy as np
from modeling_stok import STOKSpatialAmortizer

# Load model from local files or Hugging Face Hub
model = STOKSpatialAmortizer.from_pretrained(".")

# Prepare inputs for a 10x10 grid (100 tiles x 39 channels)
spatial_field = np.zeros((100, 39), dtype=np.float32)
goal_vector = np.zeros(8, dtype=np.float32)      # 5-way one-hot + (dx, dz, has_target)
goal_vector[0] = 1.0                             # Strike goal
flat_features = np.zeros(106, dtype=np.float32)  # High-level features + goal + history

# Forward inference
pred = model.forward(spatial_field, goal_vector, flat_features)

print(f"Goal Feasibility (kappa):   {pred.kappa:.4f}")
print(f"Empowerment Valence (dE):   {pred.valence:.4f}")
print(f"Expected time-to-goal:      {pred.expected_time():.2f} turns")
print(f"Opponent Intent (Argmax):   {max(pred.opp_concept.items(), key=lambda kv: kv[1])}")

License

This model, weights, and associated code are licensed under the Hyphaeic Public License (HPL). See LICENSE for the full terms.


Citation & References

@article{ringstrom2025stok,
  title   = {A Unified Theory of Compositionality, Modularity, and Interpretability in Markov Decision Processes},
  author  = {Ringstrom, Tyler and Schrater, Paul},
  journal = {arXiv preprint arXiv:2506.09499},
  year    = {2025}
}
Downloads last month
-
Safetensors
Model size
23.8k params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for hyphaeic/stok-spatialnet-v3