Mean Reward: 7.56 Std Reward: 2.71
Environment: Taxi-v4
Trained using a custom Q-Learning implementation.
-