File size: 3,119 Bytes
6a176cb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
# AGENTS.md

This repository implements temporal memory updating for efficient camera-only 3D detection.

The project should support both StreamPETR-style and RepDETR3D-style detectors. Do not hard-code the implementation to one detector unless the task explicitly asks for a detector-specific patch.

## Context policy

Do not read every markdown file by default. Start with this file and `docs/PROJECT_BRIEF.md`. Then read only the design note relevant to the requested change:

- Architecture or module design: `docs/ARCHITECTURE.md`
- 3D position embedding, ego-motion, or coordinate-frame issues: `docs/PE_AND_EGOMOTION.md`
- Training losses, teacher features, or freezing strategy: `docs/DISTILLATION.md`
- StreamPETR vs RepDETR3D compatibility: `docs/DETECTOR_ADAPTERS.md`
- Experiment configs and ablations: `docs/EXPERIMENTS.md`
- Repo integration and coding rules: `docs/IMPLEMENTATION_NOTES.md`
- Literature framing: `docs/RELATED_WORK_NOTES.md`

When context is limited, prefer reading the most specific file rather than loading all docs.

## Non-negotiable goals

1. Preserve baseline behavior when temporal memory is disabled.
2. Implement the memory updater behind a detector-agnostic adapter interface.
3. Separate appearance memory from 3D positional encoding whenever possible.
4. Include a naive reuse baseline before implementing complicated update modules.
5. Include full-encoder teacher distillation for non-keyframe memory updates.
6. Measure actual wall-clock latency, GPU memory, mAP, and NDS. FLOPs alone are insufficient.

## Definitions

- Full memory: final image-token/image-feature representation produced by the full backbone/neck path on a keyframe.
- Partial feature: intermediate ViT/block feature from the current frame, computed with less depth or cheaper processing.
- Updated memory: feature passed to the detector head/decoder on non-keyframes after combining previous memory with current partial features.
- Keyframe: frame where the full encoder path is run.
- Non-keyframe: frame where the full final memory is approximated by the updater.

## Expected implementation style

Prefer small, testable modules:

- `TemporalMemoryController`: decides keyframe vs non-keyframe and owns cached memory.
- `DetectorAdapter`: extracts/injects features for StreamPETR or RepDETR3D.
- `MemoryUpdater`: combines previous memory and current partial features.
- `PEHandler`: handles 3D PE recomputation or ego-motion-aware transformation.
- `DistillationLoss`: optional teacher loss between updated memory and full current memory.

Keep existing configs intact. Add new configs instead of mutating baseline configs.

## Minimum tests before training

- With memory disabled, outputs should match baseline within numerical tolerance.
- With `keyframe_interval=1`, outputs should match the full path.
- With naive reuse, no crash across a sequence.
- With memory update enabled, tensor shapes must match the detector input exactly.
- Ego-motion transformation code must have an identity-transform test.

## Naming

Use `tm_` or `temporal_memory_` prefixes for new modules, configs, and flags.