None defined yet.
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
Scaling Properties of Text Conditioning in Visual Generation