WM/W2R trajectories of WebShop TM-WM (step60/80/92) with Qwen3-8B and Qwen3-32B agents at 40k context.
YOULING HUANG
Ricardo-H
AI & ML interests
None yet
Recent Activity
liked a model 1 day ago
Kwaipilot/KAT-Coder-V2.5-Dev updated a collection 1 day ago
KAT upvoted a collection 1 day ago
KATOrganizations
TW WM-TM (LLaMA3.1-8B) Step171 TextWorld WM/W2R Trajectories
WM/W2R trajectories of Ricardo-H/tw-wm-token-match-llama-step171 on TextWorld test split.
TW WM-TM Step170 TextWorld WM/W2R Trajectories
WM/W2R trajectories of Ricardo-H/tw-wm-tm-0501-step-170 on TextWorld test split, evaluated by various agents.
Step92 WebShop WM/W2R Trajectories
Step92 WebShop world-model trajectories and W2R replays for Qwen and GPT agents.
OCAR · Surprise Agent-RL (Archived)
Archived checkpoints from the terminated OCAR/surprise-as-credit line. See post-mortem in the verl-agent GitHub repo.
alfworld-dual-token-0416
grpo-alfworld-0410
rlvr-f1-llama-textworld-f1
rlvr-f1
RLVR-World style Token F1 reward ablation models for BehR-WM rebuttal experiments
ws-wm-f1-0314
ws-wm-0224
Exp3-FactR: FactR-Only GRPO ablation. Exponential(a=1), behavior=0, facts=1. 240 steps, Qwen2.5-7B.
BehR-WM (LLaMA3.1-8B) TextWorld WM/W2R Trajectories
WM/W2R trajectories of Ricardo-H/BehR-WorldModel-Textworld-Llama3.1-8B on TextWorld test split.
WebShop TM-WM Checkpoint Sweep - Qwen3-32B Agent (32k, TP=4)
WM/W2R trajectories of WebShop TM-WM checkpoints (step60/80/...) evaluated with Qwen3-32B agent at 32k context.
tw-wm-tm-0501
ws-llama-webshop-token-match-0429
BehR: Behavior-Consistent World Models
Models and datasets from the BehR paper. Code: https://github.com/Ricardo-H/behr-wm
ws-wm-0410ministral
ws-wm-crossjudge-llama-0406
rlvr-f1-llama-webshop-f1
ws-wm-0314
ws-wm-llama-0227
WebShop World Model - LLaMA3.1-8B BehR-Only GRPO checkpoints (2026-02-27)
WS TM-WM Sweep - Qwen3 Agents (40k)
WM/W2R trajectories of WebShop TM-WM (step60/80/92) with Qwen3-8B and Qwen3-32B agents at 40k context.
BehR-WM (LLaMA3.1-8B) TextWorld WM/W2R Trajectories
WM/W2R trajectories of Ricardo-H/BehR-WorldModel-Textworld-Llama3.1-8B on TextWorld test split.
TW WM-TM (LLaMA3.1-8B) Step171 TextWorld WM/W2R Trajectories
WM/W2R trajectories of Ricardo-H/tw-wm-token-match-llama-step171 on TextWorld test split.
WebShop TM-WM Checkpoint Sweep - Qwen3-32B Agent (32k, TP=4)
WM/W2R trajectories of WebShop TM-WM checkpoints (step60/80/...) evaluated with Qwen3-32B agent at 32k context.
TW WM-TM Step170 TextWorld WM/W2R Trajectories
WM/W2R trajectories of Ricardo-H/tw-wm-tm-0501-step-170 on TextWorld test split, evaluated by various agents.
tw-wm-tm-0501
Step92 WebShop WM/W2R Trajectories
Step92 WebShop world-model trajectories and W2R replays for Qwen and GPT agents.
ws-llama-webshop-token-match-0429
OCAR · Surprise Agent-RL (Archived)
Archived checkpoints from the terminated OCAR/surprise-as-credit line. See post-mortem in the verl-agent GitHub repo.
BehR: Behavior-Consistent World Models
Models and datasets from the BehR paper. Code: https://github.com/Ricardo-H/behr-wm
alfworld-dual-token-0416
ws-wm-0410ministral
grpo-alfworld-0410
ws-wm-crossjudge-llama-0406
rlvr-f1-llama-textworld-f1
rlvr-f1-llama-webshop-f1
rlvr-f1
RLVR-World style Token F1 reward ablation models for BehR-WM rebuttal experiments
ws-wm-0314
ws-wm-f1-0314
ws-wm-llama-0227
WebShop World Model - LLaMA3.1-8B BehR-Only GRPO checkpoints (2026-02-27)
ws-wm-0224
Exp3-FactR: FactR-Only GRPO ablation. Exponential(a=1), behavior=0, facts=1. 240 steps, Qwen2.5-7B.