twanghcmut's picture
|
download
raw
5.3 kB
# Cosmos-Transfer2.5-2B input bundle
186 frames, 1280x720, 16 FPS = 11.625 s.
186 = 2 x 93, the frame multiple the model card reports as best.
## What to feed the model
| File | Role |
|---|---|
| `prompt.txt` | text prompt (151 words; the card requires < 300) |
| `input_video.mp4` | the video being transferred: real background plate + CG robot/brick |
| `depth.mp4` | **depth control** -- relative *inverse* depth, 8-bit, near = bright |
| `fg_mask.mp4` | binary robot+brick mask, for a control's `mask_path` |
| `seg.mp4` | segmentation control (stable palette, see `export_meta.json`) |
| `edge.mp4` | Canny edge control |
| `vis.mp4` | a pre-blurred copy -- **normally unused**, see below |
| `first_frame.png` | **real photograph** of this scene at the trajectory's t=0 |
| `spec_*.json` | ready-to-run parameter files (paths relative to the spec) |
## Run it
python examples/inference.py -i <this dir>/spec_depth_vis.json -o outputs/run1
# multi-GPU:
torchrun --nproc_per_node=4 examples/inference.py -i ... -o ...
| Spec | What it does |
|---|---|
| `spec_depth_vis.json` | **start here.** depth 0.9 + vis 0.4 |
| `spec_depth_masked_vis.json` | same, but depth applies only to the robot/brick |
| `spec_depth_only.json` | depth alone at 1.0, matching the shipped example |
| `spec_multicontrol.json` | all four branches |
## Why `vis` has no `control_path`
The repo computes the blur itself from `video_path` when you give `vis` a
`preset_blur_strength` instead of a control video -- and that preset is, by
construction, the blur the branch was trained on. `vis.mp4` here was made with a
hand-picked Gaussian sigma (18 px), which is a *guess* at that
distribution. The specs therefore use the preset. `vis.mp4` is kept only so you can
A/B it.
Do not raise the `vis` weight much: it is the strongest appearance control, and
pushing it up makes the output a de-blurred CG render rather than a photorealistic
video. `depth` is where our real information is.
## On masks
`mask_path` in this repo is **binary** -- white means "use this control here", black
means "don't". It is not a per-pixel weight map; per-control strength is the scalar
`control_weight`. `spec_depth_masked_vis.json` uses `fg_mask.mp4` to confine depth to
the robot and brick, which is worth trying because foreground depth is exact mesh
z-buffer at full resolution while background depth is upsampled from a 320x180
capture and partly interpolated.
## Two things that are easy to get wrong
**Depth is not metric here.** It is relative inverse depth, matching what
DepthAnything (Cosmos's own extractor) produces, normalised **once for the whole clip**
over `0.231 .. 1.688 m`. Per-frame
normalisation makes the control flicker and the model turns that into brightness
pumping. The metric 16-bit millimetre PNGs remain in `outputs/conditioning_cosmos_720p`
and are the source of truth; regenerate from those, never from `depth.mp4`.
**The background is a real photograph.** The capture camera is static to machine
precision, so a per-pixel temporal median over the real episode footage removes the
moving arm and leaves `background_plate.png`, pixel-aligned with the render by
construction. Only the robot and the brick are CG. This is why `edge.mp4` is usable at
all -- the background carries real texture edges rather than point-cloud speckle.
## Spec schema
Taken from the repo's own shipped examples
(`assets/robot_example/*/*_spec.json`), not guessed: `name`, `prompt_path`,
`video_path`, `guidance`, and per-control `{control_path, control_weight, mask_path}`.
Paths are relative to the spec file.
## Known limitations inherited from the render
- **no_dynamics**: Kinematic replay only: no mass, friction, or inertia is modelled. A grasp is a rigid attachment of the object to the gripper frame, not a contact/force simulation. Commanded speed changes only the timing of the motion, never its path.
- **single_viewpoint_capture**: Applies to the DEPTH channel only in this bundle. Background depth is still a single-viewpoint 2.5D capture, so surfaces the camera never saw have no points and any view far from the capture pose exposes real holes. Background *appearance* no longer has this limitation: rgb/vis/edge take their background from a real photograph (temporal median of the episode), not from the cloud.
- **no_collision_checking**: No collision checking was performed between the robot, the manipulated object(s), or the static scene, at any stage of producing this trajectory or this render.
- **no_settling_at_release**: SUPERSEDED for this bundle: the released object WAS settled under gravity (MuJoCo 3.11.0). Its tilt at release was 37.26 deg and 0.09 deg after settling, reached in 0.389 s of simulated time. Mass is not load-bearing: a rigid body's fall-and-settle trajectory is mass-independent, which was verified by re-running at 10x mass. Everything else in no_dynamics still holds -- the transport motion itself is kinematic.
- **background depth resolution**: Background depth is unprojected from a single ~320x180 capture, so it is smooth/blocky relative to the 1280x720 output. Robot and object depth comes from the mesh renderer at full output resolution and is sharp.
Holes were filled by nearest valid pixel; mean 18.00%,
max 20.68% of pixels per frame.

Xet Storage Details

Size:
5.3 kB
·
Xet hash:
8018e51351a41632bde5fdcc6fd964dccd0e84cd5d5421b2c1239691f830e4f2

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.