close
Skip to main content

FS-SAM2: Adapting Segment Anything Model 2 for Few-Shot Semantic Segmentation via Low-Rank Adaptation

  • Conference paper
  • First Online:
Image Analysis and Processing – ICIAP 2025 (ICIAP 2025)

Abstract

Few-shot semantic segmentation has recently attracted great attention. The goal is to develop a model capable of segmenting unseen classes using only a few annotated samples. Most existing approaches adapt a pre-trained model by training from scratch an additional module. Achieving optimal performance with these approaches requires extensive training on large-scale datasets. The Segment Anything Model 2 (SAM2) is a foundational model for zero-shot image and video segmentation with a modular design. In this paper, we propose a Few-Shot segmentation method based on SAM2 (FS-SAM2), where SAM2’s video capabilities are directly repurposed for the few-shot task. Moreover, we apply a Low-Rank Adaptation (LoRA) to the original modules in order to handle the diverse images typically found in standard datasets, unlike the temporally connected frames used in SAM2’s pre-training. With this approach, only a small number of parameters is meta-trained, which effectively adapts SAM2 while benefiting from its impressive segmentation performance. Our method supports any K-shot configuration. We evaluate FS-SAM2 on the PASCAL-5\(^i\), COCO-20\(^i\) and FSS-1000 datasets, achieving remarkable results and demonstrating excellent computational efficiency during inference. Code is available at https://github.com/fornib/FS-SAM2

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Subscribe and save

Springer+
from $39.99 /Month
  • Starting from 10 chapters or articles per month
  • Access and download chapters and articles from more than 300k books and 2,500 journals
  • Cancel anytime
View plans

Buy Now

Chapter
USD 29.95
Price excludes VAT (USA)
  • Available as PDF
  • Read on any device
  • Instant download
  • Own it forever
eBook
USD 79.99
Price excludes VAT (USA)
  • Available as EPUB and PDF
  • Read on any device
  • Instant download
  • Own it forever
Softcover Book
USD 99.99
Price excludes VAT (USA)
  • Compact, lightweight edition
  • Free shipping worldwide - view details

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Similar content being viewed by others

References

  1. Bai, Y., Yu, Q., Yun, B., Jin, D., Xia, Y., Wang, Y.: Revsam2: prompt sam2 for medical image segmentation via reverse-propagation without fine-tuning. arXiv:2409.04298 (2024)

  2. Bensaid, R., Gripon, V., Leduc-Primeau, F., Mauch, L., Hacene, G.B., Cardinaux, F.: A novel benchmark for few-shot semantic segmentation in the era of foundation models. arXiv:2401.11311 (2024)

  3. Csurka, G., Volpi, R., Chidlovskii, B.: Semantic image segmentation: two decades of research. Found. Trends Comput. Graph. Vis. 14(1–2), 1–162 (2022)

    Article  Google Scholar 

  4. Dosovitskiy, A., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. ICLR (2021)

    Google Scholar 

  5. Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The pascal visual object classes (voc) challenge. IJCV 88, 303–338 (2010)

    Article  Google Scholar 

  6. Fateh, A., Mohammadi, M.R., Motlagh, M.R.J.: Msdnet: Multi-scale decoder for few-shot semantic segmentation via transformer-guided prototyping. arXiv:2409.11316 (2024)

  7. Fei-Fei, L., Fergus, R., Perona, P.: One-shot learning of object categories. IEEE TPAMI 28(4), 594–611 (2006)

    Article  Google Scholar 

  8. Hariharan, B., Arbeláez, P., Bourdev, L., Maji, S., Malik, J.: Semantic contours from inverse detectors. In: ICCV, pp. 991–998. IEEE (2011)

    Google Scholar 

  9. Hong, S., Cho, S., Nam, J., Lin, S., Kim, S.: Cost aggregation with 4d convolutional swin transformer for few-shot segmentation. In: ECCV. Springer (2022)

    Google Scholar 

  10. Hu, E.J., et al.: Lora: Low-rank adaptation of large language models. ICLR (2022)

    Google Scholar 

  11. Kirillov, A., et al.: Segment anything. In: ICCV, pp. 4015–4026 (2023)

    Google Scholar 

  12. Li, X., Wei, T., Chen, Y.P., Tai, Y.W., Tang, C.K.: Fss-1000: a 1000-class dataset for few-shot segmentation. In: CVPR, pp. 2869–2878 (2020)

    Google Scholar 

  13. Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: ECCV, pp. 740–755. Springer (2014)

    Google Scholar 

  14. Liu, Y., Zhu, M., Li, H., Chen, H., Wang, X., Shen, C.: Matcher: Segment anything with one shot using all-purpose feature matching. arXiv:2305.13310 (2023)

  15. Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. ICLR (2018)

    Google Scholar 

  16. Min, J., Kang, D., Cho, M.: Hypercorrelation squeeze for few-shot segmentation. In: ICCV, pp. 6941–6952 (2021)

    Google Scholar 

  17. Nguyen, K., Todorovic, S.: Feature weighting and boosting for few-shot segmentation. In: ICCV. pp. 622–631 (2019)

    Google Scholar 

  18. Oquab, M., et al.: Dinov2: Learning robust visual features without supervision. arXiv:2304.07193 (2023)

  19. Peng, B., Tian, Z., Wu, X., Wang, C., Liu, S., Su, J., Jia, J.: Hierarchical dense correlation distillation for few-shot segmentation. In: CVPR. pp. 23641–23651 (2023)

    Google Scholar 

  20. Ravi, N., Gabeur, V., et al.: Sam 2: Segment anything in images and videos. arXiv:2408.00714 (2024)

  21. Ryali, C., et al.: Hiera: a hierarchical vision transformer without the bells-and-whistles. In: ICML, pp. 29441–29454 (2023)

    Google Scholar 

  22. Shaban, A., Bansal, S., Liu, Z., Essa, I., Boots, B.: One-shot learning for semantic segmentation. arXiv:1709.03410 (2017)

  23. Shi, X., Wei, D., Zhang, Y., Lu, D., Ning, M., Chen, J., Ma, K., Zheng, Y.: Dense cross-query-and-support attention weighted mask aggregation for few-shot segmentation. In: ECCV, pp. 151–168. Springer (2022)

    Google Scholar 

  24. Sun, Y., Chen, J., Zhang, S., Zhang, X., Chen, Q., Zhang, G., Ding, E., Wang, J., Li, Z.: Vrp-sam: Sam with visual reference prompt. In: CVPR, pp. 23565–23574 (2024)

    Google Scholar 

  25. Sun, Y., Chen, Q., He, X., Wang, J., Feng, H., Han, J., Ding, E., Cheng, J., Li, Z., Wang, J.: Singular value fine-tuning: few-shot segmentation requires few-parameters fine-tuning. NeurIPS 35, 37484–37496 (2022)

    Google Scholar 

  26. Tian, Z., Zhao, H., Shu, M., Yang, Z., Li, R., Jia, J.: Prior guided feature enrichment network for few-shot segmentation. IEEE TPAMI 44(2), 1050–1065 (2020)

    Article  Google Scholar 

  27. Wang, K., Liew, J.H., Zou, Y., Zhou, D., Feng, J.: Panet: Few-shot image semantic segmentation with prototype alignment. In: ICCV, pp. 9197–9206 (2019)

    Google Scholar 

  28. Xu, Q., Zhao, W., Lin, G., Long, C.: Self-calibrated cross attention network for few-shot segmentation. In: ICCV, pp. 655–665 (2023)

    Google Scholar 

  29. Zanella, M., Ben Ayed, I.: Low-rank few-shot adaptation of vision-language models. In: CVPR, pp. 1593–1603 (2024)

    Google Scholar 

  30. Zhang, A., Gao, G., Jiao, J., Liu, C.H., Wei, Y.: Bridge the points: Graph-based few-shot segment anything semantically. arXiv:2410.06964 (2024)

  31. Zhang, G., Kang, G., Yang, Y., Wei, Y.: Few-shot segmentation via cycle-consistent transformer. NeurIPS 34, 21984–21996 (2021)

    Google Scholar 

  32. Zhang, J.W., Sun, Y., Yang, Y., Chen, W.: Feature-proxy transformer for few-shot segmentation. NeurIPS 35, 6575–6588 (2022)

    Google Scholar 

  33. Zhu, J., Qi, Y., Wu, J.: Medical sam 2: Segment medical images as video via segment anything model 2. arXiv:2408.00874 (2024)

Download references

Acknowledgements

We acknowledge ISCRA for awarding this project access to the LEONARDO supercomputer, owned by the EuroHPC Joint Undertaking, hosted by CINECA (Italy).

Author information

Authors and Affiliations

Authors

Corresponding author

Correspondence to Bernardo Forni.

Editor information

Editors and Affiliations

Rights and permissions

Reprints and permissions

Copyright information

© 2026 The Author(s), under exclusive license to Springer Nature Switzerland AG

About this paper

Check for updates. Verify currency and authenticity via CrossMark

Cite this paper

Forni, B., Lombardi, G., Pozzi, F., Planamente, M. (2026). FS-SAM2: Adapting Segment Anything Model 2 for Few-Shot Semantic Segmentation via Low-Rank Adaptation. In: Rodolà, E., Galasso, F., Masi, I. (eds) Image Analysis and Processing – ICIAP 2025. ICIAP 2025. Lecture Notes in Computer Science, vol 16167. Springer, Cham. https://doi.org/10.1007/978-3-032-10185-3_6

Download citation

Keywords

Publish with us

Policies and ethics

Profiles

  1. Mirco Planamente