3D Interactive Physics Simulator
Agent Ready (Inference)
LEARNING TIMELINE CHECKPOINTS (νμλ¨Έμ μ μ±
μ ν)
Current: Step 50k
Step 0
Untrained
무μμ νμ & ν곡 νλλ¦
Step 10k
Novice
ν½ μ κ·Ό μλ & λΉλ§μΆ€
Step 25k
Intermediate
νκ² λ°©ν₯ μ‘°μ€ & μ€λ²μ
Step 50k
Master
μ λ° κΆ€μ & λͺ©ν μ§μ μμ°©
Live Training Telemetry
Current Reward
-1.42
β² +82.5 vs initial
Success Rate
96.4%
Target radius < 0.25m
Distance to Goal
4.2 cm
Puck to Goal
EPISODE REWARD & 100-MA CONVERGENCE
POLICY & VALUE LOSS DYNAMICS