A Unified Theory of Compositionality, Modularity, and Interpretability in Markov Decision Processes
Paper β’ 2506.09499 β’ Published
STOK-SpatialNet-v3 is a compact, multi-task spatial neural network for reward-free reinforcement learning and simultaneous-turn planning.
Instead of regressing a single scalar reward or value function, it computes the quantities of the State-Time Option Kernel (STOK) framework (Ringstrom & Schrater, 2025):
spatial-attn2
The network operates on a 2D spatial field with relational goal conditioning:
Board Cells (T tiles x 39 ch) βββΊ Embed (39 β 32) + ReLU
β
βΌ
Tile β· Tile Self-Attention (Residual)
β
βΌ
Relational Goal (8 dims) ββββββββΊ Goal-Query Cross-Attention
β
βΌ
Context Vector (32) ββββββ
High-level Features (80) βΌβββββββΊ Dense Trunk (138 β 96) + ReLU
Intent History (18) ββββββ β
βΌ
37 Output Heads (96 β 37)
Strike, Engage, Evade, Turtle, Zone.trap, spacing, punish, bait, coverage, conversion, denial, tempo).Requires only standard numpy. No PyTorch or C++ runtime required.
import numpy as np
from modeling_stok import STOKSpatialAmortizer
# Load model from local files or Hugging Face Hub
model = STOKSpatialAmortizer.from_pretrained(".")
# Prepare inputs for a 10x10 grid (100 tiles x 39 channels)
spatial_field = np.zeros((100, 39), dtype=np.float32)
goal_vector = np.zeros(8, dtype=np.float32) # 5-way one-hot + (dx, dz, has_target)
goal_vector[0] = 1.0 # Strike goal
flat_features = np.zeros(106, dtype=np.float32) # High-level features + goal + history
# Forward inference
pred = model.forward(spatial_field, goal_vector, flat_features)
print(f"Goal Feasibility (kappa): {pred.kappa:.4f}")
print(f"Empowerment Valence (dE): {pred.valence:.4f}")
print(f"Expected time-to-goal: {pred.expected_time():.2f} turns")
print(f"Opponent Intent (Argmax): {max(pred.opp_concept.items(), key=lambda kv: kv[1])}")
This model, weights, and associated code are licensed under the Hyphaeic Public License (HPL). See LICENSE for the full terms.
@article{ringstrom2025stok,
title = {A Unified Theory of Compositionality, Modularity, and Interpretability in Markov Decision Processes},
author = {Ringstrom, Tyler and Schrater, Paul},
journal = {arXiv preprint arXiv:2506.09499},
year = {2025}
}