"""Phase 2: distance to an object from one RGB image, by geometry. No depth sensor and no learned depth model. Everything here works from focal length, principal point, camera height and the pixel coordinates of a box. The LiDAR-derived `gt_distance_m` in the manifest is used to score the result and never appears in the estimation path -- delete `evaluate.py` and the rest still runs. Transcription order ------------------- Each file only depends on the ones above it, so read and write them in this order and you can run what you have at every step. 1. `ground_plane.py` The idea. A pixel is a ray; a ray that points below the horizon hits the road; that intersection is a distance. `ground_depth` is the whole thing in a dozen lines and everything else is scaffolding around it. Start here. 2. `known_size.py` The other idea, and the shorter one. If you know how big something really is, how big it looks tells you how far away it is. Three functions, all one-liners. 3. `estimate.py` Where the two meet a real detection box. Which estimator a class gets, how to sample a box rather than trusting one pixel, and what to return when neither works. 4. `evaluate.py` Scoring against LiDAR truth. Nothing conceptual, but the slicing in `report` is what made the per-scene error visible, and that finding mattered more than any of the estimator code. Four things worth having in mind before you start ------------------------------------------------- * **Distance means forward distance.** The Z component in the camera frame, not Euclidean range. It has to match `manifest.gt_distance_m` or the scores are meaningless. An object 30 m ahead and 10 m left is at 30, not 31.6. * **Camera height means height above the road.** nuScenes puts its ego origin on the ground so its translation's z component is the height; Argoverse 2 puts it at the rear axle, ~0.26 m up, and using the translation there makes every distance read ~18% short. `Calibration.cam_height_m` is corrected. * **A pixel is not a pinhole pixel until you undistort it.** Argoverse 2 cameras carry k1 near -0.24 on unrectified images. nuScenes ships rectified frames and passes zeros, so the two paths only agree once distortion is in the model. * **The bottom edge of the box is the whole estimate.** Everything above it is in the air, and a ray through a point in the air lands on the road somewhere well beyond the object. The most common way to get a plausible-looking wrong answer is to sample the middle of a box. What the numbers ended up saying -------------------------------- Sub-metre under 10 m, and error growing as Z^2 after that -- which is inherent, since sensitivity is `Z^2 / (fy * h)` and no estimator recovers information the pixel does not contain. But the *dominant* error term turned out not to be range at all: it is a per-scene offset, running -35% to +35% between scenes and close to constant within one, which a single constant row shift per scene halves. That is the flat-road assumption meeting a gradient. Per-frame ground-plane estimation is worth more than any refinement of the maths in these four files. """