Checkpoint and deployment contract
Files
model_10479.pt (v0.2.0 experimental)
Full Stage 65b RSL-RL continuation state with terrain-height observations. The actor has 162 inputs, the privileged critic has 176 inputs, and the policy still produces the same 14 joint-position actions at 50 Hz. This is the recommended checkpoint for continued obstacle training, not the currently validated deployment candidate.
terrain_scan_branch_parent_model_10280.pt
Exact Stage 65 terrain branch parent. Its 104 appended height-scan columns and normalizer entries were initialized so actor output matched the 58-input source policy before terrain learning. Keep this checkpoint to reproduce or ablate the obstacle continuation.
model_8745.pt
Full RSL-RL continuation state. It contains actor and critic weights, observation normalizers, policy distribution state, optimizer state, and iteration metadata. Use this file for Isaac Lab playback, deterministic evaluation, or further PPO training.
policy.pt
TorchScript actor exported by RSL-RL. This is the smallest convenient PyTorch deployment artifact and does not include the critic or optimizer.
policy.onnx and policy.onnx.data
ONNX actor graph plus its external tensor-data sidecar. Keep the two files in
the same directory. Moving only policy.onnx produces an incomplete model.
warm_start_model_8666.pt
The Stage 45 parent used to reproduce the Stage 46 continuation. It is not the recommended deployment checkpoint.
Observation contract
The v0.2.0 actor appends a 104-value yaw-aligned terrain-height grid to the v0.1.0 contract:
| Term | Width |
|---|---|
| measured base linear velocity | 3 |
| base angular velocity, scaled by 0.2 | 3 |
| projected gravity | 3 |
commanded (vx, vy, yaw_rate) |
3 |
| relative position of 14 policy joints | 14 |
| relative velocity of 14 policy joints, scaled by 0.05 | 14 |
| previous 14 actions | 14 |
gait phase (sin, cos) |
2 |
| COM height / vertical-velocity reference error | 2 |
| local terrain-height scan | 104 |
| Total | 162 |
The v0.2 privileged critic is 176 values because it additionally observes 14
joint efforts. The resolved definition is in
configs/release_v0.2.0/env.yaml.
The original v0.1.0 observation contract follows.
The v0.1.0 actor input has 58 values in this concatenation order:
| Term | Width |
|---|---|
| measured base linear velocity | 3 |
| base angular velocity, scaled by 0.2 | 3 |
| projected gravity | 3 |
commanded (vx, vy, yaw_rate) |
3 |
| relative position of 14 policy joints | 14 |
| relative velocity of 14 policy joints, scaled by 0.05 | 14 |
| previous 14 actions | 14 |
gait phase (sin, cos) |
2 |
| COM height / vertical-velocity reference error | 2 |
The actor's empirical normalizer is part of the exported policy. A custom runtime must still reproduce the same raw term definitions, order, scaling, gait clock, joint order, and command frame.
The privileged critic input has 72 values. It replaces policy noise with clean signals and adds 14 joint-effort values; it is unnecessary for actor-only deployment.
Action contract
The actor outputs 14 normalized joint-position offsets at 50 Hz. Isaac Lab's action term converts them to targets using the stored default offsets, action scale, joint ordering, softened position bounds, and configured actuator groups. See:
dropbear_walk/velocity_env_cfg.py::ActionsCfgdropbear_walk/isaaclab_asset/dropbear.pyconfigs/release_v0.1.0/env.yaml
Do not send the raw 14 values directly to physical actuators.
Integrity
SHA256SUMS records every released model/export and the USD. Verify from the
repository root:
sha256sum -c SHA256SUMS
Known validation boundary
The frozen policy passed zero-fall deterministic trials at ±0.20 m/s cardinal commands on a plane without pushes. This does not establish equivalent safety on rough terrain, with external disturbances, at ±0.50 m/s, during crouched locomotion, or on hardware.
That statement applies to model_8745.pt. model_10479.pt has training
telemetry on randomized low obstacles but has not yet completed an equivalent
deterministic direction-by-direction obstacle evaluation.