Why is the context length 15?
Hi Comma Team,
Based on the model config
model.config
input_size (15, 16, 32)
patch_size (1, 2, 2)
in_channels 32
out_channels 32
pose_size 6
time_factor 1.0
The context length of the DiT is 15. This doesn't check out with what it's said in the blog. The context (conditioning) is supposed to contain 1s future anchor (5 frames) + 2s past history (10 frames). Along with the query noise for diffusion, the DiT should at least have a context length of 16. Am i missing something here?
Thanks for clarifying
Yes that's correct. The history (10 frames) includes the query frame, it's just the last frame of the context. We will write a README for this model soon!
Thx for your response! Just to be clear, so there will be 5 future frame + 9 clean history + 1 (noise -> clean) in the context at each step of denoising?
nvm i found everything i need in https://github.com/YassineYousfi/openpilot.distill . Thanks for open-sourcing, amazing work!!
Updated the README with links and a short description of the model. Feel free to share what you have been using the model for if you would like to.
I'm exploring the possibility of incorporating reprojection heuristics into WM generation. Instead of generating from gauss noises, we can generate from reprojective candidates. The majority of content (especially in FOV) should already be good enough in the candidates, the WM can focus on correcting small visual artifacts, i.e. moving vehicles, skewed pixels. Some examples:
Yes I have explored things like this in the past.
For simple reprojective transformations, the diffusion model will not correct very obvious artifacts like angled vertical poles etc.
If we want the model to correct those, we would need to add that in training.
The model learns those reprojection heuristics pretty easily anyways.
The main failure is autoregressive drift with very long rollouts.


