Decision Transformer Pocket

Decision Transformer Pocket trains an offline return-conditioned policy on 4,000 mixed-quality corridor trajectories. A nearby claim pays 0.4; the larger reward requires first moving away from it, retrieving a key, crossing a door, and claiming treasure.

The behavior-cloning control sees identical states and actions but no requested return. Evaluation asks each policy to realize both low- and high-return targets from three starting positions.

Verified local result

The 18,371-parameter Decision Transformer selected the nearby reward in 100% of 300 target-0.4 episodes and the key-door treasure in 100% of 300 target-1.0 episodes. The 595-parameter behavior-cloning policy could not switch intention, reaching only 33.3% and 66.7% desired-terminal rates.

uv run python projects/decision-transformer-pocket/train.py
uv run pytest tests/test_decision_transformer_pocket.py
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support