Every example for both comparisons below — pick one from each dropdown to switch scenes. This page is separate from the main project page so the front page loads instantly.
From the visualizations above, we observe that point clouds obtained by projecting the depth-diffusion baselines' predicted depth maps (WM-DD and VGGT-DD) are considerably noisier, with scattered, uneven surfaces. Their Z3D counterparts, by contrast, produce smooth, coherent point clouds.
Drag to rotate — all point clouds stay in sync. Pick an example above to switch scenes.
Source Views
Target Views (4 target viewpoints)
Combined Point Clouds (all target views fused)
From the visualizations above, we observe that point clouds obtained by projecting Z3D's predicted depth maps are smooth and coherent.
Drag to rotate — all three point clouds stay in sync. Pick an example above to switch scenes.