StreamPETR1 / AGENTS.md
Kimhi's picture
Upload StreamPETR EVA-02 source without weights
6a176cb verified
|
Raw
History Blame Contribute Delete
3.12 kB
# AGENTS.md
This repository implements temporal memory updating for efficient camera-only 3D detection.
The project should support both StreamPETR-style and RepDETR3D-style detectors. Do not hard-code the implementation to one detector unless the task explicitly asks for a detector-specific patch.
## Context policy
Do not read every markdown file by default. Start with this file and `docs/PROJECT_BRIEF.md`. Then read only the design note relevant to the requested change:
- Architecture or module design: `docs/ARCHITECTURE.md`
- 3D position embedding, ego-motion, or coordinate-frame issues: `docs/PE_AND_EGOMOTION.md`
- Training losses, teacher features, or freezing strategy: `docs/DISTILLATION.md`
- StreamPETR vs RepDETR3D compatibility: `docs/DETECTOR_ADAPTERS.md`
- Experiment configs and ablations: `docs/EXPERIMENTS.md`
- Repo integration and coding rules: `docs/IMPLEMENTATION_NOTES.md`
- Literature framing: `docs/RELATED_WORK_NOTES.md`
When context is limited, prefer reading the most specific file rather than loading all docs.
## Non-negotiable goals
1. Preserve baseline behavior when temporal memory is disabled.
2. Implement the memory updater behind a detector-agnostic adapter interface.
3. Separate appearance memory from 3D positional encoding whenever possible.
4. Include a naive reuse baseline before implementing complicated update modules.
5. Include full-encoder teacher distillation for non-keyframe memory updates.
6. Measure actual wall-clock latency, GPU memory, mAP, and NDS. FLOPs alone are insufficient.
## Definitions
- Full memory: final image-token/image-feature representation produced by the full backbone/neck path on a keyframe.
- Partial feature: intermediate ViT/block feature from the current frame, computed with less depth or cheaper processing.
- Updated memory: feature passed to the detector head/decoder on non-keyframes after combining previous memory with current partial features.
- Keyframe: frame where the full encoder path is run.
- Non-keyframe: frame where the full final memory is approximated by the updater.
## Expected implementation style
Prefer small, testable modules:
- `TemporalMemoryController`: decides keyframe vs non-keyframe and owns cached memory.
- `DetectorAdapter`: extracts/injects features for StreamPETR or RepDETR3D.
- `MemoryUpdater`: combines previous memory and current partial features.
- `PEHandler`: handles 3D PE recomputation or ego-motion-aware transformation.
- `DistillationLoss`: optional teacher loss between updated memory and full current memory.
Keep existing configs intact. Add new configs instead of mutating baseline configs.
## Minimum tests before training
- With memory disabled, outputs should match baseline within numerical tolerance.
- With `keyframe_interval=1`, outputs should match the full path.
- With naive reuse, no crash across a sequence.
- With memory update enabled, tensor shapes must match the detector input exactly.
- Ego-motion transformation code must have an identity-transform test.
## Naming
Use `tm_` or `temporal_memory_` prefixes for new modules, configs, and flags.