- Fock Attention with MLP V_theta (Direct Token-to-Token Exchange Force Language Model)
Fock Attention with MLP V_theta (Direct Token-to-Token Exchange Force Language Model)
The Fock Attention with MLP model (previously presented on this page simply as "Fock Attention") implements the Section 5.1 Feynman diagram as a literal non-conservative force: each token emits a virtual photon carrying a key and payload, and each token absorbs with a query. The exchange coupling and the force . This is the (instantaneous exchange) limit of the Fock mechanism -- no registers, no persistence, no creation/destruction gates.
This "Route 2" across the Conservative Obstruction is the counterpart to the Fock register pool's "Route 1". It achieves 9.42 PPL on TinyStories, honest and leak-free. After correcting an earlier comparison that used the leaky-checkpoint Fock-PARFLM v2.1 number (9.30), this model is actually 0.28 PPL better than the honest, leak-fixed register-based Fock-PARFLM v2.1 (9.70) -- see Evaluation Results for the full, corrected table.
The card name now spells out "with MLP " on purpose: as the V_theta vs. Exchange Mechanism section below shows, on TinyStories the single-body potential's shape -- not the choice between register-mediated and direct-exchange Fock mechanisms -- is the dominant factor separating this model's 9.42 PPL from the family-best 9.04 (depth-conditioned anisotropic-Gaussian ).
Part of the Semantic Simulation framework.
Table of Contents
- Model Details
- Architecture
- Why Not a Transformer?
- Geometric Capabilities
- How to Get Started
- Training Details
- Evaluation Results
- SPLM Family Overview
- Bias, Risks, and Limitations
- Citation
- Environmental Impact
Model Details
Model Description
The Fock Attention PARFLM extends the Multi-Xi PARFLM with a direct token-to-token exchange force inspired by Feynman's virtual particle exchange diagram. Unlike the register-based Fock-PARFLM which uses persistent auxiliary state, this variant implements instantaneous exchange:
- Token emits: key , payload
- Token absorbs: query
- Coupling:
- Exchange force on token i:
The force is injected post-Verlet step as a non-conservative addition:
The learned exchange scale converges to approximately -0.32 (repulsive), meaning the exchange force pushes tokens apart in hidden-state space rather than attracting them.
- Developed by: Dimitar P. Gueorguiev (Independent Researcher)
- Model type: Conservative autoregressive LM with direct exchange force
- Language: English
- License: CC-BY-4.0
Model Sources
- Paper: Semantic Simulation: A Prescriptive Lagrangian Framework for Efficient Semantic Inference (Section 17c)
- Repository: github.com/dimitarpg13/semsimula-paper
- Model source code:
notebooks/conservative_arch/parf/model_fock_attention.py
Architecture
Input tokens x_1, ..., x_T
|
Embedding E[x] + positional encoding
|
For each of L=8 integration steps:
|
+-- K-EMA channels: xi^(k)_t = causal_ema(h, alpha_k) [K=4 channels]
|
+-- Single-body: V_theta([xi_1..xi_K, h]) -> R [3-layer MLP]
|
+-- Sparse PARF: V_phi(h_t, h_s) [top-k=8 routing]
|
+-- Conservative force: f = -grad_h(V_theta + V_phi) [autograd]
|
+-- Damped Euler step: v += dt*f/m; v /= (1+dt*gamma); h += dt*v
|
+-- Direct exchange force (Feynman diagram):
| q_i = W_Q * h_t [4 heads, d_k=32]
| k_j = W_K * h_s
| v_j = W_V * h_s
| alpha_ij = softmax_j(q_i . k_j / sqrt(d_k)) [causal mask]
| F_i = sum_j alpha_ij * v_j
| h += (dt^2/m) * tanh(s_ex) * F_i [non-conservative]
|
+-- LayerNorm(h)
|
Logits = h @ E^T [tied embeddings]
| Parameter | Value |
|---|---|
| Hidden dim (d) | 256 |
| Layers (L) | 8 |
| hidden / depth | 1024 / 3 |
| Xi channels (K) | 4 |
| kind | structural_competitive |
| hidden (H) | 128 |
| Sparse routing top_k | 8 |
| Exchange heads | 4 |
| Exchange d_k | 32 |
| Learned exchange scale | ~ -0.32 (repulsive) |
| Mass model | logfreq |
| Damping | 0.30 nominal, ~0.033 effective (LayerNorm prevents compounding; see note below) |
| Total parameters | 16,714,708 |
Effective damping. The nominal overstates the true dissipation. The LayerNorm applied after each integration step rescales the hidden state, absorbing most of the velocity decay. The dynamics are therefore heavily underdamped even at this nominal value. Gamma-sweep experiments on the OpenWebText-scale Fock-PARFLM variant confirm that the effective damping γeff is much smaller than the nominal coefficient.
Key Design Properties
- Two routes across the obstruction: This model takes "Route 2" (direct exchange, ) while the Fock-PARFLM takes "Route 1" (register-mediated, ). Both achieve comparable, honest PPL (9.42 vs. 9.70).
- Repulsive exchange: The learned scale is negative, meaning the exchange force diversifies token representations rather than collapsing them.
- Minimal overhead: Only ~131K parameters for the exchange mechanism (Q/K/V projections + scale scalar).
- Otherwise conservative: The core dynamics remain fully conservative; the exchange force is the only non-conservative component.
V_theta vs. Exchange Mechanism: What Actually Drives PPL on TinyStories
This model's single-body potential is a plain 3-layer MLP -- the same unbounded, black-box shape used in the original Fock-PARFLM. Its most visible design feature is the direct exchange force above, so it is tempting to credit that Feynman-diagram mechanism for most of the PPL gain over the no-Fock Multi-Xi PARFLM baseline (12.06 -> 9.42). The family's own ablations, holding one factor fixed at a time, point the other way:
| shape | Exchange / Fock mechanism | Honest PPL |
|---|---|---|
| MLP | none (Multi-Xi PARFLM) | 12.06 |
| MLP | register-based Fock v2.1 (Fock-PARFLM v2.1) | 9.70 |
| MLP | direct exchange (this model) | 9.42 |
| Isotropic Gaussian | register-based Fock v2.1 (sibling) | 16.33 |
| Anisotropic Gaussian | register-based Fock v2.1 (sibling) | 9.04 |
| Anisotropic Gaussian | register-based Fock v2.1 + first-order integrator (Fock-G1) | 8.95 |
Holding the exchange mechanism fixed at "register-based Fock v2.1" and varying only 's shape swings PPL by more than 7 points (16.33 down to 8.95). Holding fixed at MLP and varying only the exchange mechanism (none -> registers -> direct exchange) swings PPL by about 2.6 points -- but the choice between the two Fock-mechanism flavours (registers vs. direct exchange) is worth only 0.3 points, and direct exchange (this model) actually wins that comparison.
On TinyStories, 's profile is doing most of the work; the specific exchange mechanism is a second-order effect. Whether the anisotropic-Gaussian-beats-MLP ordering seen here survives at OpenWebText scale is still an open question: only the anisotropic-Gaussian has been scaled up to OpenWebText so far (see the gamma-sweep family), so there is currently no MLP-\(V_\theta\) OpenWebText-scale checkpoint to compare against directly.
Why Not a Transformer?
The Fock Attention PARFLM is not based on the Transformer architecture. There are no Transformer-style FFN towers. The conservative core is driven by two small scalar-potential MLPs — (≈3.4M params) and (≈19K params). The direct exchange force adds only ≈131K parameters (Q/K/V projections).
Key structural differences from Transformers:
| Property | Transformer (GPT-2 small) | Fock Attention (this model) |
|---|---|---|
| Architecture | Self-attention + FFN blocks | Scalar-potential gradient flow + direct exchange |
| Core computation | 50.3M (MLP) + 28.3M (attention) | 3.4M + 19K + 131K (exchange) |
| Runtime state per token | — KV-cache grows linearly | — exchange force is |
| Total parameters | 124M | 16.7M |
| Pairwise token interaction | dense attention | direct exchange + sparse PARF |
Note: Because this model uses a direct token-to-token exchange force (the "Route 2" across the Conservative Obstruction), its runtime cost is like attention. For fully inference, see the register-based Fock-PARFLM v2.1 ("Route 1"), which achieves an honest 9.70 PPL vs. this model's 9.42 PPL -- i.e. inference here comes at a small PPL cost relative to this model's direct exchange, not the other way around (see the causal-leak notice on that card for why an earlier version of this comparison had the direction reversed).
Geometric Capabilities
Note: This model uses a direct exchange force that is non-conservative, breaking the full Riemannian guarantee. The full Riemannian geometry — Jacobi metric, computable geodesics, curvature-based hallucination detection, native chain-of-thought — is available only in the purely conservative variants: Multi-Xi SPLM, Multi-Xi PARFLM, and Fock-PARFLM v2.1. See Section 18d and Section 23 of the paper for details.
Damped Riemannian Geometry (June 2026 update)
A Riemannian Geometry Diagnostic Battery run on all three SPLM-family checkpoints confirmed that while the force field is conservative, the full dynamics are dominated by the damping term in the integrator. Key findings:
- Metric validity (Arm 1): The layer-dependent conformal factor at 100% of positions — the Riemannian metric is well-defined everywhere.
- Geodesic compliance (Arm 2): Undamped geodesics predict the wrong direction (compliance ); the damped geodesic equation (with friction ) is required.
- Energy dissipation (Arm 4): Energy decays monotonically across layers (SPLM: 172% drift), consistent with the designed damping — not a numerical artefact.
- Asymmetry (Arm 5): The layer map is strongly asymmetric (\(R^2_{\text{sym}} \ll 0\)), expected for damped dynamics + LayerNorm.
The theoretical framework has been updated from undamped Maupertuis-Jacobi to damped Riemannian geometry with a contact Hamiltonian interpretation. See the companion note for full details.
How to Get Started
# Clone the companion repository for full source code
# git clone https://github.com/dimitarpg13/semsimula-paper.git
# cd semsimula-paper/notebooks/conservative_arch
import torch
import sys
sys.path.insert(0, "parf")
sys.path.insert(0, "multixi")
sys.path.insert(0, "energetic_minima")
sys.path.insert(0, "sarf_mass_variant")
from parf.model_fock_attention import FockAttentionPARFLM, FockAttentionConfig
config = FockAttentionConfig(
vocab_size=50257,
d=256,
n_layers=8,
v_hidden=1024,
v_depth=3,
max_len=1024,
block_size=512,
gamma=0.30,
xi_channels=4,
v_phi_kind="structural_competitive",
v_phi_hidden=128,
top_k=8,
exchange_n_heads=4,
exchange_d_k=32,
)
model = FockAttentionPARFLM(config)
print(f"Parameters: {sum(p.numel() for p in model.parameters()):,}")
# Forward pass
x = torch.randint(0, 50257, (1, 64))
logits, loss = model(x, targets=x)
Available Checkpoint
A trained checkpoint (PPL 9.42, 16k steps) is included in this repository:
| File | Description |
|---|---|
checkpoint/model.pt |
Full model state dict (64 MB) |
training_log.jsonl |
Per-step training metrics |
loss_curve.png |
Training/validation loss plot |
training_summary.md |
Hyperparameters and final metrics |
To load the checkpoint:
from huggingface_hub import hf_hub_download
import torch
# Download checkpoint
ckpt_path = hf_hub_download(
repo_id="dimitarpg13/semsimula-fock-attention",
filename="checkpoint/model.pt",
)
# Load into model (after creating model as above)
state = torch.load(ckpt_path, map_location="cpu")
model.load_state_dict(state["model_state_dict"])
model.eval()
Training Details
Training Data
TinyStories -- GPT-2 BPE tokenization. Training cap: 5M tokens.
Training Procedure
| Hyperparameter | Value |
|---|---|
| Optimizer | AdamW |
| Learning rate | 5e-4 (cosine decay) |
| Warmup steps | 400 |
| Weight decay | 0.01 |
| Gradient clipping | 1.0 |
| Batch size | 16 |
| Block size | 512 |
| Training steps | 16,000 |
| Memory optimisation | Level-2 grad checkpoint |
| Hardware | A100/H100 (Google Colab) |
Training Script
notebooks/conservative_arch/scaleup/train_fock_attention_scaleup.py
Colab Notebook
notebooks/conservative_arch/scaleup/colab_fock_attention_h128.ipynb — 4-arm sweep over head count and schedule with live progress display, saves results to Google Drive.
Training Results
notebooks/conservative_arch/scaleup/results/semsimula_fock_attention_h128/ — training logs, loss curves, and experiment report.
Head / d_k Sweep
| Arm | Heads | d_k | Steps | PPL |
|---|---|---|---|---|
| direct_K4_h1_8k | 1 | 64 | 8k | 11.48 |
| direct_K4_h4_8k | 4 | 32 | 8k | 10.93 |
| direct_K8_h4_8k | 8 | 32 | 8k | 11.65 |
| direct_K4_h4_16k | 4 | 32 | 16k | 9.42 |
Evaluation Results
TinyStories Validation Perplexity
| Model | PPL | Params | Gap vs Attention |
|---|---|---|---|
| Matched Attention (baseline) | 7.81 | 19.5M | -- |
| Hybrid SPLM+Attn | 8.50 | ~19.0M | +0.69 |
| Fock Attention (MLP V_theta) (this model) | 9.42 | 16.7M | +1.61 |
| Fock-PARFLM v2.1 | 9.70 | 17.4M | +1.89 |
| Multi-Xi SPLM | 11.51 | 16.5M | +3.70 |
| Multi-Xi PARFLM | 12.06 | 17.6M | +4.25 |
Corrected comparison: earlier versions of this card compared against Fock-PARFLM v2.1's leaky-checkpoint PPL of 9.30, concluding that "register persistence provides a small advantage." That checkpoint has since been re-trained with the causal-leak fix; its honest PPL is 9.70, which means Fock Attention (this model) is actually 0.28 PPL better, not worse (see the causal-leak notice on the Fock-PARFLM v2.1 card for the full account). Fock Attention also has the smallest parameter count in the family (16.7M) and the simplest Fock mechanism (no creation/destruction gates, no stack discipline). As the V_theta vs. Exchange Mechanism section above argues, the 0.62 PPL gap between this model and the family-best anisotropic-Gaussian variant (9.04) is better explained by 's shape than by the exchange mechanism.
SPLM Family Overview
This model is part of the Semantic Simulation SPLM family:
| Model | Design | Corpus | PPL | HuggingFace |
|---|---|---|---|---|
| Multi-Xi SPLM (MLP) | Pure scalar potential | TinyStories | 11.51 | semsimula-splm-multixi |
| Multi-Xi SPLM (SQ3) | Structured scalar potential | TinyStories | 13.33 | semsimula-splm-multixi-structured-vtheta |
| Multi-Xi PARFLM (MLP) | Scalar + pairwise forces | TinyStories | 12.06 | semsimula-parflm-multixi |
| Multi-Xi PARFLM (SQ3) | Structured scalar + pairwise | TinyStories | 12.27 | semsimula-parflm-multixi-structured-vtheta |
| Fock-PARFLM v2.1 (MLP) | PARFLM + Fock registers | TinyStories | 9.70 | semsimula-fock-parflm |
| Fock-PARFLM v2.1 (SQ3) | Structured + pairwise + Fock | TinyStories | 10.90 | semsimula-fock-parflm-structured-vtheta |
| Fock-PARFLM v2.1 (depth-cond. isotropic Gaussian) | Bounded multi-context + pairwise + Fock | TinyStories | 16.33 | semsimula-fock-parflm-depthcond-vtheta |
| Fock-PARFLM v2.1 (depth-cond. anisotropic Gaussian + fock-reg) | Bounded, ellipsoidal multi-context + pairwise + Fock | TinyStories | 9.04 | semsimula-fock-parflm-anisogaussian-vtheta |
| Fock-PARFLM v2.1 (aniso-Gaussian + fock-reg, first-order/Fock-G1) | Same as above, gradient-flow integrator | TinyStories | 8.95 | semsimula-fock-parflm-anisogaussian-vtheta-fock-g1 |
| Fock Attention (MLP V_theta) | Fock + attention | TinyStories | 9.42 | this model |
| Hybrid SPLM+Attn | Attention + SPLM refinement | TinyStories | 8.50 | semsimula-hybrid-splm |
| Fock-PARFLM v2.1 (aniso-Gaussian + fock-reg, gamma sweep, d=384) | Bounded, ellipsoidal multi-context + pairwise + Fock — 8-way gamma sweep + geodesic analysis | OpenWebText | 278.27 (best of 8, 3K-step sweep) | semsimula-fock-parflm-anisogaussian-vtheta-owt-d384-gammasweep |
| Fock-PARFLM v2.1 (aniso-Gaussian + fock-reg, gamma sweep, d=768) | Bounded, ellipsoidal multi-context + pairwise + Fock — 8-way gamma sweep + geodesic analysis | OpenWebText | 326.97 (best of 8, 3K-step sweep) | semsimula-fock-parflm-anisogaussian-vtheta-owt-d768-gammasweep |
| Fock-PARFLM v2.1 (aniso-Gaussian + fock-reg, gamma sweep, d=1024) | Bounded, ellipsoidal multi-context + pairwise + Fock — 8-way gamma sweep + geodesic analysis | OpenWebText | 244.23 (best of 8, 3K-step sweep) | semsimula-fock-parflm-anisogaussian-vtheta-owt-d1024-gammasweep |
| Fock-PARFLM v2.1 (aniso-Gaussian + fock-reg, Verlet instability, d=384) | Same architecture, two full-run attempts — SCAF stiffness audit identifies structural Verlet instability | OpenWebText | 184.11 / 211.63 (both runs stalled, not final) | semsimula-fock-parflm-anisogaussian-vtheta-owt-d384-verlet-instability |
Collection: Semantic Simulation SPLM Model Family
Bias, Risks, and Limitations
- Research checkpoint only. Proof-of-concept for direct exchange force as a non-conservative complement to conservative dynamics.
- TinyStories only. Trained exclusively on synthetic children's stories (~5M tokens).
- English only. No multilingual capability.
- Small scale. 16.7M parameters, 256-dim hidden states.
- No safety training. No RLHF, DPO, or safety filtering has been applied.
Citation
@misc{Gueorguiev2026SemSim,
author = {Gueorguiev, Dimitar P.},
title = {Semantic Simulation: A Prescriptive Lagrangian Framework
for Efficient Semantic Inference --- A Conservative-by-
Construction Language Model and the Shared-Potential
Separator, with a Correspondence to Joint Embedding
Predictive Architectures},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.19712427},
url = {https://doi.org/10.5281/zenodo.19712427},
note = {Version v15 (Jun 7, 2026).
Companion code repository (DOI 10.5281/zenodo.20579561):
\url{https://github.com/dimitarpg13/semsimula-paper}}
}
Environmental Impact
- Hardware: NVIDIA A100/H100 (Google Colab)
- Training time: ~8 hours (16,000 steps with gradient checkpointing)
- Carbon footprint: Estimated < 3 kg CO2
- Downloads last month
- 172
Dataset used to train dimitarpg13/semsimula-fock-attention
Collection including dimitarpg13/semsimula-fock-attention
Evaluation results
- Validation Perplexity on TinyStoriesvalidation set self-reported9.420
