Comment by in-silico
How would it do what it does without those things?
How would it do what it does without those things?
Yes, it reproduces what it is given by modelling the rules of physics, geometry, etc.
For example, image generators like stable diffusion carry strong representations of depth and geometry, such that performant depth estimation models can be built out of them with minimal retraining. This continues to be true for video generation models.
Early work on the subject: https://arxiv.org/pdf/2409.09144
Like all these models work, by simple interpolation.