Walker

Walker is a decoder-only language model that adds trajectory modeling to the Transformer.

The main idea is simple:

Attention looks at other tokens. Walker looks at how representations move.

How does attention work?

A Transformer represents each token as a vector. For every token, self-attention creates:

  • Query (Q): what the token is looking for
  • Key (K): what each token contains
  • Value (V): the information that can be retrieved

The query is compared with other keys to determine which tokens are relevant:

Attention(Q,K,V)=softmax(QKTd)V \text{Attention}(Q,K,V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d}}\right)V

In simple terms:

Attention asks: "What should I look at?"

A different approach: trajectory

Walker keeps attention, but asks another question:

"Where is my representation going?"

Consider the hidden states:

x1, x2, x3, x4,… x_1,\ x_2,\ x_3,\ x_4,\ldots

Instead of only using the states themselves, Walker measures their changes:

dt=xtβˆ’xtβˆ’1 d_t = x_t-x_{t-1}

This is a discrete derivative of the hidden representation.

Now the model can observe a trajectory rather than just a sequence of points:

x₁ β†’ xβ‚‚ β†’ x₃ β†’ xβ‚„ β†’ ...

     ↓
 differences
     ↓
 trajectory

Walker uses multiple scales of these differences to capture both short- and longer-range motion, along with changes in direction.

Why "Walker"?

Imagine the hidden states forming a path through a high-dimensional space.

A Transformer can look around the path using attention.

Walker also tries to walk along the path.

Transformer
    β”‚
    └── Attention
          └── "What should I look at?"

Walker
    β”‚
    └── Trajectory
          └── "Where am I going?"

The two mechanisms are complementary. Walker is not designed to replace attention; it adds trajectory features as a residual pathway alongside the Transformer.

Architecture

                 Input Tokens
                      β”‚
                 Token Embedding
                      β”‚
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β”‚                       β”‚
      Attention                Walker
          β”‚                       β”‚
   Token relationships        Trajectory
          β”‚                       β”‚
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                      β”‚
                 Residual Update
                      β”‚
                      β–Ό
                 Next Block

The Walker pathway is initialized as an exact no-op, allowing the model to begin as a normal Transformer and gradually learn how much trajectory information to use.

Summary

Walker=Attention+Trajectory Modeling \boxed{ \text{Walker} = \text{Attention} + \text{Trajectory Modeling} }

Attention: What information is relevant?

Walker: How is the representation moving?

Walker is a Transformer that not only looks at the sequence, but also walks through the trajectory of its own representations.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train Papatpong/Walker

Space using Papatpong/Walker 1