cone-distance / src /depth /__init__.py
Aryan Sethi
Claude Opus 5 (1M context)
Document src/depth for transcription
699e481
Raw History Blame Contribute Delete
3.43 kB
"""Phase 2: distance to an object from one RGB image, by geometry.
No depth sensor and no learned depth model. Everything here works from focal
length, principal point, camera height and the pixel coordinates of a box. The
LiDAR-derived `gt_distance_m` in the manifest is used to score the result and
never appears in the estimation path -- delete `evaluate.py` and the rest still
runs.
Transcription order
-------------------
Each file only depends on the ones above it, so read and write them in this
order and you can run what you have at every step.
1. `ground_plane.py` The idea. A pixel is a ray; a ray that points below the
horizon hits the road; that intersection is a distance.
`ground_depth` is the whole thing in a dozen lines and
everything else is scaffolding around it. Start here.
2. `known_size.py` The other idea, and the shorter one. If you know how big
something really is, how big it looks tells you how far
away it is. Three functions, all one-liners.
3. `estimate.py` Where the two meet a real detection box. Which estimator
a class gets, how to sample a box rather than trusting
one pixel, and what to return when neither works.
4. `evaluate.py` Scoring against LiDAR truth. Nothing conceptual, but the
slicing in `report` is what made the per-scene error
visible, and that finding mattered more than any of the
estimator code.
Four things worth having in mind before you start
-------------------------------------------------
* **Distance means forward distance.** The Z component in the camera frame,
not Euclidean range. It has to match `manifest.gt_distance_m` or the scores
are meaningless. An object 30 m ahead and 10 m left is at 30, not 31.6.
* **Camera height means height above the road.** nuScenes puts its ego origin
on the ground so its translation's z component is the height; Argoverse 2
puts it at the rear axle, ~0.26 m up, and using the translation there makes
every distance read ~18% short. `Calibration.cam_height_m` is corrected.
* **A pixel is not a pinhole pixel until you undistort it.** Argoverse 2
cameras carry k1 near -0.24 on unrectified images. nuScenes ships rectified
frames and passes zeros, so the two paths only agree once distortion is in
the model.
* **The bottom edge of the box is the whole estimate.** Everything above it is
in the air, and a ray through a point in the air lands on the road somewhere
well beyond the object. The most common way to get a plausible-looking wrong
answer is to sample the middle of a box.
What the numbers ended up saying
--------------------------------
Sub-metre under 10 m, and error growing as Z^2 after that -- which is inherent,
since sensitivity is `Z^2 / (fy * h)` and no estimator recovers information the
pixel does not contain. But the *dominant* error term turned out not to be range
at all: it is a per-scene offset, running -35% to +35% between scenes and close
to constant within one, which a single constant row shift per scene halves. That
is the flat-road assumption meeting a gradient. Per-frame ground-plane estimation
is worth more than any refinement of the maths in these four files.
"""