kpshinnik's picture
|
download
raw
1.03 kB
# Paper's numerical experiment setup (Appendix D)
MDP: S = {0,1,2}; N = 6 state-action pairs, order:
(0,stay),(0,leave),(1,stay),(1,leave),(2,left),(2,right)
Reward r = (0.001, 0.5, 0.001, 0.5, 1, 1)
True transition P (rows = state-action, cols = next state 0,1,2):
(0,stay): [1, 0, 0 ]
(0,leave): [0.4, 0, 0.6]
(1,stay): [0, 1, 0 ]
(1,leave): [0, 0.4, 0.6]
(2,left): [0.4, 0.4, 0.2]
(2,right): [0.2, 0.2, 0.6]
Settings: γ = 0.5, ε = 0.05, initial dist q = (0.4, 0.4, 0.2). Vary samples T from 100 to 16000.
Inaccurate prediction P̂ (OpPMD-NAC):
[0,1,0],[0,1,0],[1,0,0],[1,0,0],[1,0,0],[0,1,0] ("Dist(P,P̂) ~ 1" per paper)
Three algorithms compared:
1. OpPMD-AC: Algorithm 1 with accurate prediction P̂ = P
2. OpPMD-NAC: Algorithm 1 with inaccurate P̂ above
3. SMD-DMDP-JINSID: Jin & Sidford [2020] Algorithm 1 (no prediction) — baseline
Metrics: Figure 1 = Duality Gap GAP(μ̄, v̄); Figure 2 = value function.
Optimal policy: leave (state 0), leave (state 1), right (state 2).

Xet Storage Details

Size:
1.03 kB
·
Xet hash:
069fa4a92d3cf89b037f5d52076f030d497c4b8d84475ab9c21355b6fa727e53

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.