Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Spaces:
Crusadersk
/
icml26-delayed-obs-rl-repro
like
1
Running
App
Files
Files
Community
main
icml26-delayed-obs-rl-repro
463 kB
Ctrl+K
Ctrl+K
1 contributor
History:
7 commits
Crusadersk
Add full-MDP factor sweep (real |S|<=49 MDP): sqrt-D_max/S/K + linear-H reproduce the theorem exponents; sqrt-A instance-dependent (0.37) but O(sqrt A) upper bound holds; honest, not tuned
b774650
verified
about 1 month ago
artifacts
Publish portable ICML 2026 reproduction logbook
about 1 month ago
evidence-package
Add full-MDP factor sweep (real |S|<=49 MDP): sqrt-D_max/S/K + linear-H reproduce the theorem exponents; sqrt-A instance-dependent (0.37) but O(sqrt A) upper bound holds; honest, not tuned
about 1 month ago
pages
Add full-MDP factor sweep (real |S|<=49 MDP): sqrt-D_max/S/K + linear-H reproduce the theorem exponents; sqrt-A instance-dependent (0.37) but O(sqrt A) upper bound holds; honest, not tuned
about 1 month ago
.gitattributes
Safe
1.52 kB
initial commit
about 1 month ago
README.md
Safe
476 Bytes
Update logbook: Repro - Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
about 1 month ago
bucket-icon.svg
Safe
413 Bytes
Update logbook: Repro - Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
about 1 month ago
index.html
Safe
1.93 kB
Update logbook: Repro - Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
about 1 month ago
logbook.css
Safe
29.6 kB
Update logbook: Repro - Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
about 1 month ago
logbook.js
Safe
77.3 kB
Update logbook: Repro - Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
about 1 month ago
logbook.json
Safe
1.95 kB
Add full-MDP factor sweep (real |S|<=49 MDP): sqrt-D_max/S/K + linear-H reproduce the theorem exponents; sqrt-A instance-dependent (0.37) but O(sqrt A) upper bound holds; honest, not tuned
about 1 month ago
trackio-logo-light.png
Safe
30 kB
Update logbook: Repro - Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
about 1 month ago
trackio-logo.png
Safe
55.6 kB
Update logbook: Repro - Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
about 1 month ago
trackio-wordmark-dark.png
Safe
89.8 kB
Update logbook: Repro - Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
about 1 month ago