GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting

Lei Yao

Yi Wang

Yi Zhang

Moyun Liu

Lap-Pui Chau

The significance of informative and robust point representations has been widely acknowledged for 3D scene understanding. Despite existing self-supervised pre-training counterparts demonstrating promising performance, the model collapse and structural information deficiency remain prevalent due to insufficient point discrimination difficulty, yielding unreliable expressions and suboptimal performance. In this paper, we present GaussianCross, a novel cross-modal self-supervised 3D representation learning architecture integrating feed-forward 3D Gaussian Splatting (3DGS) techniques to address current challenges. GaussianCross seamlessly converts scale-inconsistent 3D point clouds into a unified cuboid-normalized Gaussian representation without missing details, enabling stable and generalizable pre-training. Subsequently, a tri-attribute adaptive distillation splatting module is incorporated to construct a 3D feature field, facilitating synergetic feature capturing of appearance, geometry, and semantic cues to maintain cross-modal consistency. To validate GaussianCross, we perform extensive evaluations on various benchmarks, including ScanNet, ScanNet200, and S3DIS. In particular, GaussianCross shows a prominent parameter and data efficiency, achieving superior performance through linear probing (<0.1% parameters) and limited data training (1% of scenes) compared to state-of-the-art methods. Furthermore, GaussianCross demonstrates strong generalization capabilities, improving the full fine-tuning accuracy by 9.3% mIoU and 6.1% AP$_{50}$ on ScanNet200 semantic and instance segmentation tasks, respectively, supporting the effectiveness of our approach. The code, weights, and visualizations are publicly available at \href{https://rayyoh.github.io/GaussianCross/}{https://rayyoh.github.io/GaussianCross/}.

PDF URL

Featured

Platforms

SplatCapture 3.0.0 Adds Spline Capture

SplatCapture 3.0.0 adds a spline capture mode, Unreal Engine 5.8 support, and square-aspect camera rigs that give Gaussian splatting trainers clean intrinsics.

Michael Rubloff

Jul 17, 2026

Platforms

Marble x Nuke Loads World Labs Marble Splats Straight Into Nuke's 3D Viewport

Marble x Nuke is a free Nuke 17 toolset that turns World Labs Marble text or image prompts into Gaussian splats loaded straight into Nuke's native 3D viewport.

Michael Rubloff

Jul 15, 2026

Platforms

Houdini 22 Ships, Making Native Gaussian Splats Generally Available

Houdini 22 is now generally available, making SideFX's native Gaussian splatting pipeline, PDG training and Copernicus compositing, production ready today.

Michael Rubloff

Jul 15, 2026

360 Gaussian v1.4.5 Adds Per-Clip Extraction Settings

360 Gaussian v1.4.5 updates the 360 video to 3DGS tool with per-clip extraction settings and fixes swapped GPS coordinates for accurate scale.

Michael Rubloff

Jul 15, 2026