We examine positional encodings for multi-view transformers that course of tokens from a set of posed enter photos, and search a mechanism that encodes patches uniquely, permits SE(3)-invariant consideration with multi-frequency similarity, and may be adaptive to the geometry of the underlying scene. We discover that prior (absolute or relative) encoding schemes for multi-view consideration don’t meet the above desiderata, and current RayRoPE to deal with this hole. RayRoPE represents patch positions primarily based on related rays however leverages a predicted level alongside the ray as a substitute of the path for a geometry-aware encoding. To realize SE(3) invariance, RayRoPE computes query-frame projective coordinates for computing multi-frequency similarity. Lastly, because the ‘predicted’ 3D level alongside a ray will not be exact, RayRoPE presents a mechanism to analytically compute the anticipated place encoding below uncertainty. We validate RayRoPE on the duties of novel-view synthesis and stereo depth estimation and present that it persistently improves over alternate place encoding schemes (e.g. 15% relative enchancment on LPIPS in CO3D). We additionally present that RayRoPE can seamlessly incorporate RGB-D enter, leading to even bigger positive factors over options that can’t positionally encode this info.
- †Carnegie Mellon College
