Abstract
Few-shot semantic segmentation has recently attracted great attention. The goal is to develop a model capable of segmenting unseen classes using only a few annotated samples. Most existing approaches adapt a pre-trained model by training from scratch an additional module. Achieving optimal performance with these approaches requires extensive training on large-scale datasets. The Segment Anything Model 2 (SAM2) is a foundational model for zero-shot image and video segmentation with a modular design. In this paper, we propose a Few-Shot segmentation method based on SAM2 (FS-SAM2), where SAM2’s video capabilities are directly repurposed for the few-shot task. Moreover, we apply a Low-Rank Adaptation (LoRA) to the original modules in order to handle the diverse images typically found in standard datasets, unlike the temporally connected frames used in SAM2’s pre-training. With this approach, only a small number of parameters is meta-trained, which effectively adapts SAM2 while benefiting from its impressive segmentation performance. Our method supports any K-shot configuration. We evaluate FS-SAM2 on the PASCAL-5\(^i\), COCO-20\(^i\) and FSS-1000 datasets, achieving remarkable results and demonstrating excellent computational efficiency during inference. Code is available at https://github.com/fornib/FS-SAM2
Access this chapter
Tax calculation will be finalised at checkout
Purchases are for personal use only
Similar content being viewed by others
References
Bai, Y., Yu, Q., Yun, B., Jin, D., Xia, Y., Wang, Y.: Revsam2: prompt sam2 for medical image segmentation via reverse-propagation without fine-tuning. arXiv:2409.04298 (2024)
Bensaid, R., Gripon, V., Leduc-Primeau, F., Mauch, L., Hacene, G.B., Cardinaux, F.: A novel benchmark for few-shot semantic segmentation in the era of foundation models. arXiv:2401.11311 (2024)
Csurka, G., Volpi, R., Chidlovskii, B.: Semantic image segmentation: two decades of research. Found. Trends Comput. Graph. Vis. 14(1–2), 1–162 (2022)
Dosovitskiy, A., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. ICLR (2021)
Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The pascal visual object classes (voc) challenge. IJCV 88, 303–338 (2010)
Fateh, A., Mohammadi, M.R., Motlagh, M.R.J.: Msdnet: Multi-scale decoder for few-shot semantic segmentation via transformer-guided prototyping. arXiv:2409.11316 (2024)
Fei-Fei, L., Fergus, R., Perona, P.: One-shot learning of object categories. IEEE TPAMI 28(4), 594–611 (2006)
Hariharan, B., Arbeláez, P., Bourdev, L., Maji, S., Malik, J.: Semantic contours from inverse detectors. In: ICCV, pp. 991–998. IEEE (2011)
Hong, S., Cho, S., Nam, J., Lin, S., Kim, S.: Cost aggregation with 4d convolutional swin transformer for few-shot segmentation. In: ECCV. Springer (2022)
Hu, E.J., et al.: Lora: Low-rank adaptation of large language models. ICLR (2022)
Kirillov, A., et al.: Segment anything. In: ICCV, pp. 4015–4026 (2023)
Li, X., Wei, T., Chen, Y.P., Tai, Y.W., Tang, C.K.: Fss-1000: a 1000-class dataset for few-shot segmentation. In: CVPR, pp. 2869–2878 (2020)
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: ECCV, pp. 740–755. Springer (2014)
Liu, Y., Zhu, M., Li, H., Chen, H., Wang, X., Shen, C.: Matcher: Segment anything with one shot using all-purpose feature matching. arXiv:2305.13310 (2023)
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. ICLR (2018)
Min, J., Kang, D., Cho, M.: Hypercorrelation squeeze for few-shot segmentation. In: ICCV, pp. 6941–6952 (2021)
Nguyen, K., Todorovic, S.: Feature weighting and boosting for few-shot segmentation. In: ICCV. pp. 622–631 (2019)
Oquab, M., et al.: Dinov2: Learning robust visual features without supervision. arXiv:2304.07193 (2023)
Peng, B., Tian, Z., Wu, X., Wang, C., Liu, S., Su, J., Jia, J.: Hierarchical dense correlation distillation for few-shot segmentation. In: CVPR. pp. 23641–23651 (2023)
Ravi, N., Gabeur, V., et al.: Sam 2: Segment anything in images and videos. arXiv:2408.00714 (2024)
Ryali, C., et al.: Hiera: a hierarchical vision transformer without the bells-and-whistles. In: ICML, pp. 29441–29454 (2023)
Shaban, A., Bansal, S., Liu, Z., Essa, I., Boots, B.: One-shot learning for semantic segmentation. arXiv:1709.03410 (2017)
Shi, X., Wei, D., Zhang, Y., Lu, D., Ning, M., Chen, J., Ma, K., Zheng, Y.: Dense cross-query-and-support attention weighted mask aggregation for few-shot segmentation. In: ECCV, pp. 151–168. Springer (2022)
Sun, Y., Chen, J., Zhang, S., Zhang, X., Chen, Q., Zhang, G., Ding, E., Wang, J., Li, Z.: Vrp-sam: Sam with visual reference prompt. In: CVPR, pp. 23565–23574 (2024)
Sun, Y., Chen, Q., He, X., Wang, J., Feng, H., Han, J., Ding, E., Cheng, J., Li, Z., Wang, J.: Singular value fine-tuning: few-shot segmentation requires few-parameters fine-tuning. NeurIPS 35, 37484–37496 (2022)
Tian, Z., Zhao, H., Shu, M., Yang, Z., Li, R., Jia, J.: Prior guided feature enrichment network for few-shot segmentation. IEEE TPAMI 44(2), 1050–1065 (2020)
Wang, K., Liew, J.H., Zou, Y., Zhou, D., Feng, J.: Panet: Few-shot image semantic segmentation with prototype alignment. In: ICCV, pp. 9197–9206 (2019)
Xu, Q., Zhao, W., Lin, G., Long, C.: Self-calibrated cross attention network for few-shot segmentation. In: ICCV, pp. 655–665 (2023)
Zanella, M., Ben Ayed, I.: Low-rank few-shot adaptation of vision-language models. In: CVPR, pp. 1593–1603 (2024)
Zhang, A., Gao, G., Jiao, J., Liu, C.H., Wei, Y.: Bridge the points: Graph-based few-shot segment anything semantically. arXiv:2410.06964 (2024)
Zhang, G., Kang, G., Yang, Y., Wei, Y.: Few-shot segmentation via cycle-consistent transformer. NeurIPS 34, 21984–21996 (2021)
Zhang, J.W., Sun, Y., Yang, Y., Chen, W.: Feature-proxy transformer for few-shot segmentation. NeurIPS 35, 6575–6588 (2022)
Zhu, J., Qi, Y., Wu, J.: Medical sam 2: Segment medical images as video via segment anything model 2. arXiv:2408.00874 (2024)
Acknowledgements
We acknowledge ISCRA for awarding this project access to the LEONARDO supercomputer, owned by the EuroHPC Joint Undertaking, hosted by CINECA (Italy).
Author information
Authors and Affiliations
Corresponding author
Editor information
Editors and Affiliations
Rights and permissions
Copyright information
© 2026 The Author(s), under exclusive license to Springer Nature Switzerland AG
About this paper
Cite this paper
Forni, B., Lombardi, G., Pozzi, F., Planamente, M. (2026). FS-SAM2: Adapting Segment Anything Model 2 for Few-Shot Semantic Segmentation via Low-Rank Adaptation. In: Rodolà, E., Galasso, F., Masi, I. (eds) Image Analysis and Processing – ICIAP 2025. ICIAP 2025. Lecture Notes in Computer Science, vol 16167. Springer, Cham. https://doi.org/10.1007/978-3-032-10185-3_6
Download citation
DOI: https://doi.org/10.1007/978-3-032-10185-3_6
Published:
Publisher Name: Springer, Cham
Print ISBN: 978-3-032-10184-6
Online ISBN: 978-3-032-10185-3
eBook Packages: Computer ScienceComputer Science (R0)Springer Nature Proceedings Computer Science
