How Vision-Language Models Learned to Reason About Space (10 Papers, One Thread)

A guided tour of spatial reasoning in VLMs — from BLINK's perception gap to GR3D — read through one lens: the representational mismatch between language tokens and continuous 3D geometry.

Read Original

Related