twanghcmut's picture
|
download
raw
5.3 kB

Cosmos-Transfer2.5-2B input bundle

186 frames, 1280x720, 16 FPS = 11.625 s. 186 = 2 x 93, the frame multiple the model card reports as best.

What to feed the model

File Role
prompt.txt text prompt (151 words; the card requires < 300)
input_video.mp4 the video being transferred: real background plate + CG robot/brick
depth.mp4 depth control -- relative inverse depth, 8-bit, near = bright
fg_mask.mp4 binary robot+brick mask, for a control's mask_path
seg.mp4 segmentation control (stable palette, see export_meta.json)
edge.mp4 Canny edge control
vis.mp4 a pre-blurred copy -- normally unused, see below
first_frame.png real photograph of this scene at the trajectory's t=0
spec_*.json ready-to-run parameter files (paths relative to the spec)

Run it

python examples/inference.py -i <this dir>/spec_depth_vis.json -o outputs/run1
# multi-GPU:
torchrun --nproc_per_node=4 examples/inference.py -i ... -o ...
Spec What it does
spec_depth_vis.json start here. depth 0.9 + vis 0.4
spec_depth_masked_vis.json same, but depth applies only to the robot/brick
spec_depth_only.json depth alone at 1.0, matching the shipped example
spec_multicontrol.json all four branches

Why vis has no control_path

The repo computes the blur itself from video_path when you give vis a preset_blur_strength instead of a control video -- and that preset is, by construction, the blur the branch was trained on. vis.mp4 here was made with a hand-picked Gaussian sigma (18 px), which is a guess at that distribution. The specs therefore use the preset. vis.mp4 is kept only so you can A/B it.

Do not raise the vis weight much: it is the strongest appearance control, and pushing it up makes the output a de-blurred CG render rather than a photorealistic video. depth is where our real information is.

On masks

mask_path in this repo is binary -- white means "use this control here", black means "don't". It is not a per-pixel weight map; per-control strength is the scalar control_weight. spec_depth_masked_vis.json uses fg_mask.mp4 to confine depth to the robot and brick, which is worth trying because foreground depth is exact mesh z-buffer at full resolution while background depth is upsampled from a 320x180 capture and partly interpolated.

Two things that are easy to get wrong

Depth is not metric here. It is relative inverse depth, matching what DepthAnything (Cosmos's own extractor) produces, normalised once for the whole clip over 0.231 .. 1.688 m. Per-frame normalisation makes the control flicker and the model turns that into brightness pumping. The metric 16-bit millimetre PNGs remain in outputs/conditioning_cosmos_720p and are the source of truth; regenerate from those, never from depth.mp4.

The background is a real photograph. The capture camera is static to machine precision, so a per-pixel temporal median over the real episode footage removes the moving arm and leaves background_plate.png, pixel-aligned with the render by construction. Only the robot and the brick are CG. This is why edge.mp4 is usable at all -- the background carries real texture edges rather than point-cloud speckle.

Spec schema

Taken from the repo's own shipped examples (assets/robot_example/*/*_spec.json), not guessed: name, prompt_path, video_path, guidance, and per-control {control_path, control_weight, mask_path}. Paths are relative to the spec file.

Known limitations inherited from the render

  • no_dynamics: Kinematic replay only: no mass, friction, or inertia is modelled. A grasp is a rigid attachment of the object to the gripper frame, not a contact/force simulation. Commanded speed changes only the timing of the motion, never its path.
  • single_viewpoint_capture: Applies to the DEPTH channel only in this bundle. Background depth is still a single-viewpoint 2.5D capture, so surfaces the camera never saw have no points and any view far from the capture pose exposes real holes. Background appearance no longer has this limitation: rgb/vis/edge take their background from a real photograph (temporal median of the episode), not from the cloud.
  • no_collision_checking: No collision checking was performed between the robot, the manipulated object(s), or the static scene, at any stage of producing this trajectory or this render.
  • no_settling_at_release: SUPERSEDED for this bundle: the released object WAS settled under gravity (MuJoCo 3.11.0). Its tilt at release was 37.26 deg and 0.09 deg after settling, reached in 0.389 s of simulated time. Mass is not load-bearing: a rigid body's fall-and-settle trajectory is mass-independent, which was verified by re-running at 10x mass. Everything else in no_dynamics still holds -- the transport motion itself is kinematic.
  • background depth resolution: Background depth is unprojected from a single ~320x180 capture, so it is smooth/blocky relative to the 1280x720 output. Robot and object depth comes from the mesh renderer at full output resolution and is sharp. Holes were filled by nearest valid pixel; mean 18.00%, max 20.68% of pixels per frame.

Xet Storage Details

Size:
5.3 kB
·
Xet hash:
8018e51351a41632bde5fdcc6fd964dccd0e84cd5d5421b2c1239691f830e4f2

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.