close
Skip to main content

SSCCPC-Net: Simultaneously Learning 2D and 3D Features with CLIP for Semantic Scene Completion on Point Cloud

  • Conference paper
  • First Online:
Advances in Computer Graphics (CGI 2024)

Part of the book series: Lecture Notes in Computer Science ((LNCS,volume 15340))

Included in the following conference series:

  • 461 Accesses

  • 1 Citation

Abstract

Compare to traditional Scene Completion (SC), Semantic Scene Completion (SSC) is a challenging task that aims to generate complete and semantically consistent 3D scene from partial and sparse input data, which is fundamental to fully understanding the scene and being able to interact with it. Consequently, the SSC task has received much attention in recent years. Most of the methods are voxel-based approaches, but they have high computational and memory requirements. A few works based on point cloud do not sufficiently exploit the correlation between semantic segmentation and geometric completion subtasks, while focusing too much on point cloud shape features and ignoring the rich texture information that RGB images can provide. In this paper, we present SSCCPC-Net (Semantic Scene Completion with CLIP on Point Cloud-Net), a novel network architecture for point cloud semantic scene completion using a combination of 2D and 3D features. Inspired by recent works of large pretrained vision-language models in semantic segmentation, we explore to accomplish SSC task with the help of Contrastive Language-Image Pre-Training (CLIP) model. Specifically, we use the CLIP features for guidance to fuse the 2D features extracted from the RGB image and the 3D features extracted from the point cloud. The fused features are then fed into our designed Semantic-Completion Decoder for per-point semantic prediction and semantic labeling-assisted point cloud completion. Finally, we obtain the complete semantically point cloud. Numerous experiments have demonstrated that our method has higher effectiveness and generalizability compared to state-of-the-art methods.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Subscribe and save

Springer+
from $39.99 /Month
  • Starting from 10 chapters or articles per month
  • Access and download chapters and articles from more than 300k books and 2,500 journals
  • Cancel anytime
View plans

Buy Now

Chapter
USD 29.95
Price excludes VAT (USA)
  • Available as PDF
  • Read on any device
  • Instant download
  • Own it forever
eBook
USD 64.99
Price excludes VAT (USA)
  • Available as EPUB and PDF
  • Read on any device
  • Instant download
  • Own it forever
Softcover Book
USD 84.99
Price excludes VAT (USA)
  • Compact, lightweight edition
  • Free shipping worldwide - view details

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Similar content being viewed by others

References

  1. Chen, R., Wu, J., Luo, Y., Xu, G.: PointMM: point cloud semantic segmentation CNN under multi-spatial feature encoding and multi-head attention pooling. Remote Sens. 16(7) (2024). https://doi.org/10.3390/rs16071246

  2. Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., Niessner, M.: ScanNet: richly-annotated 3D reconstructions of indoor scenes. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017). https://doi.org/10.1109/cvpr.2017.261

  3. Ding, R., Yang, J., Xue, C., Zhang, W., Bai, S., Qi, X.: PLA: language-driven open-vocabulary 3d scene understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7010–7019 (2023)

    Google Scholar 

  4. Dong, H., et al.: CVSformer: cross-view synthesis transformer for semantic scene completion. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 8874–8883 (2023)

    Google Scholar 

  5. Garbade, M., Chen, Y.T., Sawatzky, J., Gall, J.: Two stream 3D semantic scene completion. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (2019)

    Google Scholar 

  6. Guo, Y., Tong, X.: View-volume network for semantic scene completion from a single depth image. In: Proceedings of the 27th International Joint Conference on Artificial Intelligence, pp. 726–732. IJCAI 2018, AAAI Press (2018)

    Google Scholar 

  7. Hu, W., Zhao, H., Jiang, L., Jia, J., Wong, T.T.: Bidirectional projection network for cross dimension scene understanding. In: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021). https://doi.org/10.1109/cvpr46437.2021.01414

  8. Jaritz, M., Gu, J., Su, H.: Multi-view PointNet for 3D scene understanding. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops (2019)

    Google Scholar 

  9. Jatavallabhula, K., et al.: ConceptFusion: open-set multimodal 3D mapping. Sci. Syst. (RSS), Robot. (2023)

    Google Scholar 

  10. Kolodiazhnyi, M., Vorontsova, A., Konushin, A., Rukhovich, D.: OneFormer3D: one transformer for unified point cloud segmentation (2023)

    Google Scholar 

  11. Li, B., Weinberger, K.Q., Belongie, S., Koltun, V., Ranftl, R.: Language-driven semantic segmentation. In: International Conference on Learning Representations (2022)

    Google Scholar 

  12. Li, J., et al.: RGBD based dimensional decomposition residual network for 3D semantic scene completion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019)

    Google Scholar 

  13. Li, J., Song, Q., Yan, X., Chen, Y., Huang, R.: From front to rear: 3D semantic scene completion through planar convolution and attention-based network. IEEE Transactions on Multimedia, p. 1–14 (2023). https://doi.org/10.1109/tmm.2023.3234441

  14. Liang, Y., Chen, B., Song, S.: SSCNav: confidence-aware semantic scene completion for visual semantic navigation. In: 2021 IEEE International Conference on Robotics and Automation (ICRA) (2021)

    Google Scholar 

  15. Lin, D., Dong, H., Ma, E., Wang, L., Li, P.: Multi-head multi-scale feature fusion network for semantic scene completion. In: 2023 International Conference on Artificial Intelligence and Education (ICAIE) (2023)

    Google Scholar 

  16. Michele, B., Boulch, A., Puy, G., Bucher, M., Marlet, R.: Generative zero-shot learning for semantic segmentation of 3D point clouds. In: 2021 International Conference on 3D Vision (3DV), pp. 992–1002 (2021)

    Google Scholar 

  17. Peng, S., Genova, K., Jiang, C., Tagliasacchi, A., Pollefeys, M., Funkhouser, T.: OpenScene: 3D scene understanding with open vocabularies. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 815–824 (2023)

    Google Scholar 

  18. Qi, C.R., Su, H., Mo, K., Guibas, L.J.: PointNet: deep learning on point sets for 3D classification and segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017)

    Google Scholar 

  19. Qi, C.R., Yi, L., Su, H., Guibas, L.J.: PointNet++: deep hierarchical feature learning on point sets in a metric space. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Advances in Neural Information Processing Systems, vol. 30. Curran Associates, Inc. (2017)

    Google Scholar 

  20. Radford, A., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748–8763. PMLR (2021)

    Google Scholar 

  21. Song, S., Yu, F., Zeng, A., Chang, A.X., Savva, M., Funkhouser, T.: Semantic scene completion from a single depth image. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 190–198 (2017). https://doi.org/10.1109/CVPR.2017.28

  22. Wang, F., Zhang, D., Zhang, H., Tang, J., Sun, Q.: Semantic scene completion with cleaner self. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 867–877 (2023)

    Google Scholar 

  23. Wang, Y., Wang, J., Qu, Y., Qi, Y.: RIP-NeRF: learning rotation-invariant point-based neural radiance field for fine-grained editing and compositing. In: Proceedings of the 2023 ACM International Conference on Multimedia Retrieval, pp. 125–134 (2023)

    Google Scholar 

  24. Wu, Y., Han, X.F., Xiao, G.: Language-driven open-vocabulary 3D semantic segmentation with knowledge distillation. In: ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3320–3324 (2024). https://doi.org/10.1109/ICASSP48485.2024.10448295

  25. Xia, Z., et al.: SCPNet: semantic scene completion on point cloud. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 17642–17651 (2023)

    Google Scholar 

  26. Xu, J., et al: CasFusionNet: A cascaded network for point cloud semantic scene completion by dense feature fusion. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 3, pp. 3018–3026 (2023)

    Google Scholar 

  27. Yao, J., et al: NDC-scene: boost monocular 3D semantic scene completion in normalized device coordinates space. In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9421–9431. IEEE Computer Society (2023)

    Google Scholar 

  28. Zhang, J., Zhao, H., Yao, A., Chen, Y., Zhang, L., Liao, H.: Efficient semantic scene completion network with spatial group convolution, pp. 749–765 (2018). https://doi.org/10.1007/978-3-030-01258-8_45

  29. Zhang, P., Liu, W., Lei, Y., Lu, H., Yang, X.: Cascaded context pyramid for full-resolution 3D semantic scene completion. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (2019)

    Google Scholar 

  30. Zhang, S., Li, S., Hao, A., Qin, H.: Point cloud semantic scene completion from RGB-D images. In: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 3385–3393 (2022)

    Google Scholar 

  31. Zhang, Z., Han, X., Dong, B., Li, T., Yin, B., Yang, X.: Point cloud scene completion with joint color and semantic estimation from single RGB-D image. IEEE Trans. Pattern Anal. Mach. Intell. 45(9), 11079–11095 (2023). https://doi.org/10.1109/TPAMI.2023.3264449

    Article  Google Scholar 

  32. Zhong, M., Zeng, G.: Semantic point completion network for 3D semantic scene completion. European Conference on Artificial Intelligence,European Conference on Artificial Intelligence (2020)

    Google Scholar 

Download references

Acknowledgments

This paper is supported by National Natural Science Foundation of China (No. 62072020)and the Leading Talents in Innovation and Entrepreneurship of Qingdao, China (19-3-2-21-zhc).

Author information

Authors and Affiliations

Authors

Corresponding author

Correspondence to Yue Qi.

Editor information

Editors and Affiliations

Ethics declarations

Disclosure of Interests

The authors have no competing interests.

Rights and permissions

Reprints and permissions

Copyright information

© 2025 The Author(s), under exclusive license to Springer Nature Switzerland AG

About this paper

Check for updates. Verify currency and authenticity via CrossMark

Cite this paper

Duan, W., Bao, Y., Qi, Y. (2025). SSCCPC-Net: Simultaneously Learning 2D and 3D Features with CLIP for Semantic Scene Completion on Point Cloud. In: Magnenat-Thalmann, N., Kim, J., Sheng, B., Deng, Z., Thalmann, D., Li, P. (eds) Advances in Computer Graphics. CGI 2024. Lecture Notes in Computer Science, vol 15340. Springer, Cham. https://doi.org/10.1007/978-3-031-82024-3_2

Download citation

Keywords

Publish with us

Policies and ethics