Abstract
Compare to traditional Scene Completion (SC), Semantic Scene Completion (SSC) is a challenging task that aims to generate complete and semantically consistent 3D scene from partial and sparse input data, which is fundamental to fully understanding the scene and being able to interact with it. Consequently, the SSC task has received much attention in recent years. Most of the methods are voxel-based approaches, but they have high computational and memory requirements. A few works based on point cloud do not sufficiently exploit the correlation between semantic segmentation and geometric completion subtasks, while focusing too much on point cloud shape features and ignoring the rich texture information that RGB images can provide. In this paper, we present SSCCPC-Net (Semantic Scene Completion with CLIP on Point Cloud-Net), a novel network architecture for point cloud semantic scene completion using a combination of 2D and 3D features. Inspired by recent works of large pretrained vision-language models in semantic segmentation, we explore to accomplish SSC task with the help of Contrastive Language-Image Pre-Training (CLIP) model. Specifically, we use the CLIP features for guidance to fuse the 2D features extracted from the RGB image and the 3D features extracted from the point cloud. The fused features are then fed into our designed Semantic-Completion Decoder for per-point semantic prediction and semantic labeling-assisted point cloud completion. Finally, we obtain the complete semantically point cloud. Numerous experiments have demonstrated that our method has higher effectiveness and generalizability compared to state-of-the-art methods.
Access this chapter
Tax calculation will be finalised at checkout
Purchases are for personal use only
Similar content being viewed by others
References
Chen, R., Wu, J., Luo, Y., Xu, G.: PointMM: point cloud semantic segmentation CNN under multi-spatial feature encoding and multi-head attention pooling. Remote Sens. 16(7) (2024). https://doi.org/10.3390/rs16071246
Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., Niessner, M.: ScanNet: richly-annotated 3D reconstructions of indoor scenes. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017). https://doi.org/10.1109/cvpr.2017.261
Ding, R., Yang, J., Xue, C., Zhang, W., Bai, S., Qi, X.: PLA: language-driven open-vocabulary 3d scene understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7010–7019 (2023)
Dong, H., et al.: CVSformer: cross-view synthesis transformer for semantic scene completion. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 8874–8883 (2023)
Garbade, M., Chen, Y.T., Sawatzky, J., Gall, J.: Two stream 3D semantic scene completion. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (2019)
Guo, Y., Tong, X.: View-volume network for semantic scene completion from a single depth image. In: Proceedings of the 27th International Joint Conference on Artificial Intelligence, pp. 726–732. IJCAI 2018, AAAI Press (2018)
Hu, W., Zhao, H., Jiang, L., Jia, J., Wong, T.T.: Bidirectional projection network for cross dimension scene understanding. In: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021). https://doi.org/10.1109/cvpr46437.2021.01414
Jaritz, M., Gu, J., Su, H.: Multi-view PointNet for 3D scene understanding. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops (2019)
Jatavallabhula, K., et al.: ConceptFusion: open-set multimodal 3D mapping. Sci. Syst. (RSS), Robot. (2023)
Kolodiazhnyi, M., Vorontsova, A., Konushin, A., Rukhovich, D.: OneFormer3D: one transformer for unified point cloud segmentation (2023)
Li, B., Weinberger, K.Q., Belongie, S., Koltun, V., Ranftl, R.: Language-driven semantic segmentation. In: International Conference on Learning Representations (2022)
Li, J., et al.: RGBD based dimensional decomposition residual network for 3D semantic scene completion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019)
Li, J., Song, Q., Yan, X., Chen, Y., Huang, R.: From front to rear: 3D semantic scene completion through planar convolution and attention-based network. IEEE Transactions on Multimedia, p. 1–14 (2023). https://doi.org/10.1109/tmm.2023.3234441
Liang, Y., Chen, B., Song, S.: SSCNav: confidence-aware semantic scene completion for visual semantic navigation. In: 2021 IEEE International Conference on Robotics and Automation (ICRA) (2021)
Lin, D., Dong, H., Ma, E., Wang, L., Li, P.: Multi-head multi-scale feature fusion network for semantic scene completion. In: 2023 International Conference on Artificial Intelligence and Education (ICAIE) (2023)
Michele, B., Boulch, A., Puy, G., Bucher, M., Marlet, R.: Generative zero-shot learning for semantic segmentation of 3D point clouds. In: 2021 International Conference on 3D Vision (3DV), pp. 992–1002 (2021)
Peng, S., Genova, K., Jiang, C., Tagliasacchi, A., Pollefeys, M., Funkhouser, T.: OpenScene: 3D scene understanding with open vocabularies. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 815–824 (2023)
Qi, C.R., Su, H., Mo, K., Guibas, L.J.: PointNet: deep learning on point sets for 3D classification and segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017)
Qi, C.R., Yi, L., Su, H., Guibas, L.J.: PointNet++: deep hierarchical feature learning on point sets in a metric space. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Advances in Neural Information Processing Systems, vol. 30. Curran Associates, Inc. (2017)
Radford, A., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748–8763. PMLR (2021)
Song, S., Yu, F., Zeng, A., Chang, A.X., Savva, M., Funkhouser, T.: Semantic scene completion from a single depth image. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 190–198 (2017). https://doi.org/10.1109/CVPR.2017.28
Wang, F., Zhang, D., Zhang, H., Tang, J., Sun, Q.: Semantic scene completion with cleaner self. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 867–877 (2023)
Wang, Y., Wang, J., Qu, Y., Qi, Y.: RIP-NeRF: learning rotation-invariant point-based neural radiance field for fine-grained editing and compositing. In: Proceedings of the 2023 ACM International Conference on Multimedia Retrieval, pp. 125–134 (2023)
Wu, Y., Han, X.F., Xiao, G.: Language-driven open-vocabulary 3D semantic segmentation with knowledge distillation. In: ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3320–3324 (2024). https://doi.org/10.1109/ICASSP48485.2024.10448295
Xia, Z., et al.: SCPNet: semantic scene completion on point cloud. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 17642–17651 (2023)
Xu, J., et al: CasFusionNet: A cascaded network for point cloud semantic scene completion by dense feature fusion. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 3, pp. 3018–3026 (2023)
Yao, J., et al: NDC-scene: boost monocular 3D semantic scene completion in normalized device coordinates space. In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9421–9431. IEEE Computer Society (2023)
Zhang, J., Zhao, H., Yao, A., Chen, Y., Zhang, L., Liao, H.: Efficient semantic scene completion network with spatial group convolution, pp. 749–765 (2018). https://doi.org/10.1007/978-3-030-01258-8_45
Zhang, P., Liu, W., Lei, Y., Lu, H., Yang, X.: Cascaded context pyramid for full-resolution 3D semantic scene completion. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (2019)
Zhang, S., Li, S., Hao, A., Qin, H.: Point cloud semantic scene completion from RGB-D images. In: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 3385–3393 (2022)
Zhang, Z., Han, X., Dong, B., Li, T., Yin, B., Yang, X.: Point cloud scene completion with joint color and semantic estimation from single RGB-D image. IEEE Trans. Pattern Anal. Mach. Intell. 45(9), 11079–11095 (2023). https://doi.org/10.1109/TPAMI.2023.3264449
Zhong, M., Zeng, G.: Semantic point completion network for 3D semantic scene completion. European Conference on Artificial Intelligence,European Conference on Artificial Intelligence (2020)
Acknowledgments
This paper is supported by National Natural Science Foundation of China (No. 62072020)and the Leading Talents in Innovation and Entrepreneurship of Qingdao, China (19-3-2-21-zhc).
Author information
Authors and Affiliations
Corresponding author
Editor information
Editors and Affiliations
Ethics declarations
Disclosure of Interests
The authors have no competing interests.
Rights and permissions
Copyright information
© 2025 The Author(s), under exclusive license to Springer Nature Switzerland AG
About this paper
Cite this paper
Duan, W., Bao, Y., Qi, Y. (2025). SSCCPC-Net: Simultaneously Learning 2D and 3D Features with CLIP for Semantic Scene Completion on Point Cloud. In: Magnenat-Thalmann, N., Kim, J., Sheng, B., Deng, Z., Thalmann, D., Li, P. (eds) Advances in Computer Graphics. CGI 2024. Lecture Notes in Computer Science, vol 15340. Springer, Cham. https://doi.org/10.1007/978-3-031-82024-3_2
Download citation
DOI: https://doi.org/10.1007/978-3-031-82024-3_2
Published:
Publisher Name: Springer, Cham
Print ISBN: 978-3-031-82023-6
Online ISBN: 978-3-031-82024-3
eBook Packages: Computer ScienceComputer Science (R0)Springer Nature Proceedings Computer Science
