{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,17]],"date-time":"2026-05-17T04:24:30Z","timestamp":1778991870216,"version":"3.51.4"},"reference-count":74,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,1,11]],"date-time":"2024-01-11T00:00:00Z","timestamp":1704931200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,4,30]]},"abstract":"<jats:p>While deep learning has been widely used for video analytics, such as video classification and action detection, dense action detection with fast-moving subjects from sports videos is still challenging. In this work, we release yet another sports video benchmark<jats:bold>P<jats:sup>2<\/jats:sup>ANet<\/jats:bold>for<jats:italic><jats:underline>P<\/jats:underline><\/jats:italic>ing<jats:italic><jats:underline>P<\/jats:underline><\/jats:italic>ong-<jats:italic><jats:underline>A<\/jats:underline><\/jats:italic>ction detection, which consists of 2,721 video clips collected from the broadcasting videos of professional table tennis matches in World Table\u00a0Tennis Championships and Olympiads. We work with a crew of table tennis professionals and referees on a specially designed annotation toolbox to obtain fine-grained action labels (in 14 classes) for every ping-pong action that appeared in the dataset, and formulate two sets of action detection problems\u2014<jats:italic>action localization<\/jats:italic>and<jats:italic>action recognition<\/jats:italic>. We evaluate a number of commonly seen action recognition (e.g., TSM, TSN, Video SwinTransformer, and Slowfast) and action localization models (e.g., BSN, BSN++, BMN, TCANet), using<jats:bold>P<jats:sup>2<\/jats:sup>ANet<\/jats:bold>for both problems, under various settings. These models can only achieve 48% area under the AR-AN curve for localization and 82% top-one accuracy for recognition since the ping-pong actions are dense with fast-moving subjects but broadcasting videos are with only 25 FPS. The results confirm that<jats:bold>P<jats:sup>2<\/jats:sup>ANet<\/jats:bold>is still a challenging task and can be used as a special benchmark for dense action detection from videos. We invite readers to examine our dataset by visiting the following link:<jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/Fred1991\/P2ANET\">https:\/\/github.com\/Fred1991\/P2ANET<\/jats:ext-link>.<\/jats:p>","DOI":"10.1145\/3633516","type":"journal-article","created":{"date-parts":[[2023,11,28]],"date-time":"2023-11-28T12:16:25Z","timestamp":1701173785000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":21,"title":["<b>P<sup>2<\/sup>ANet<\/b>: A Large-Scale Benchmark for Dense Action Detection from Table\u00a0Tennis Match Broadcasting Videos"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6997-1989","authenticated-orcid":false,"given":"Jiang","family":"Bian","sequence":"first","affiliation":[{"name":"Baidu Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2582-8256","authenticated-orcid":false,"given":"Xuhong","family":"Li","sequence":"additional","affiliation":[{"name":"Baidu Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4740-6932","authenticated-orcid":false,"given":"Tao","family":"Wang","sequence":"additional","affiliation":[{"name":"Baidu Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1562-8098","authenticated-orcid":false,"given":"Qingzhong","family":"Wang","sequence":"additional","affiliation":[{"name":"Baidu Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2696-6452","authenticated-orcid":false,"given":"Jun","family":"Huang","sequence":"additional","affiliation":[{"name":"Baidu Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-1671-7427","authenticated-orcid":false,"given":"Chen","family":"Liu","sequence":"additional","affiliation":[{"name":"Baidu Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-0522-8079","authenticated-orcid":false,"given":"Jun","family":"Zhao","sequence":"additional","affiliation":[{"name":"Baidu Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3952-4402","authenticated-orcid":false,"given":"Feixiang","family":"Lu","sequence":"additional","affiliation":[{"name":"Baidu Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7561-1672","authenticated-orcid":false,"given":"Dejing","family":"Dou","sequence":"additional","affiliation":[{"name":"Baidu Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5451-3253","authenticated-orcid":false,"given":"Haoyi","family":"Xiong","sequence":"additional","affiliation":[{"name":"Baidu Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,1,11]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Sami Abu-El-Haija Nisarg Kothari Joonseok Lee Paul Natsev George Toderici Balakrishnan Varadarajan and Sudheendra Vijayanarasimhan. 2016. Youtube-8M: A large-scale video classification benchmark. arXiv:1609.08675. Retrieved from https:\/\/arxiv.org\/abs\/1609.08675"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2020.103655"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.239"},{"key":"e_1_3_2_6_2","first-page":"813","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Bertasius Gedas","year":"2021","unstructured":"Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding?. In Proceedings of the International Conference on Machine Learning. PMLR, 813\u2013824."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3161050"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Desheng Cai Shengsheng Qian Quan Fang Jun Hu Wenkui Ding and Changsheng Xu. 2023. Heterogeneous graph contrastive learning network for personalized micro-video recommendation. IEEE Transactions on Multimedia 25 (2023) 2761\u20132773.","DOI":"10.1109\/TMM.2022.3151026"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1117\/1.JMI.7.4.044503"},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","unstructured":"Ekin D. Cubuk Barret Zoph Dandelion Mane Vijay Vasudevan and Quoc V. Le. 2019. Autoaugment: Learning augmentation strategies from data. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 113\u2013123.","DOI":"10.1109\/CVPR.2019.00020"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.610"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2005.177"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00508"},{"key":"e_1_3_2_16_2","unstructured":"Jianfeng Dong Xirong Li Chaoxi Xu Xun Yang Gang Yang Xun Wang and Meng Wang. 2022. Dual encoding for video retrieval by Text. IEEE Trans. Pattern Anal. Mach. Intell . 44 8 (Aug 2022) 4065\u20134080."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/DICTA.2017.8227494"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00630"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58548-8_13"},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","unstructured":"Raghav Goyal Samira Ebrahimi Kahou Vincent Michalski Joanna Materzynska Susanne Westphal Heuna Kim Valentin Haenel Ingo Fruend Peter Yianilos Moritz Mueller-Freitag Florian Hoppe Christian Thurau Ingo Bax and Roland Memisevic. 2017. The \u201csomething something\u201d video database for learning and evaluating visual common sense. In Proceedings of the IEEE International Conference on Computer Vision . 5842\u20135850.","DOI":"10.1109\/ICCV.2017.622"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","unstructured":"Chunhui Gu Chen Sun David A Ross Carl Vondrick Caroline Pantofaru Yeqing Li Sudheendra Vijayanarasimhan George Toderici Susanna Ricco Rahul Sukthankar Cordelia Schmid and Jitendra Malik. 2018. Ava: A video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6047\u20136056.","DOI":"10.1109\/CVPR.2018.00633"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00685"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.217"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3422844.3423051"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.223"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-92185-9_46"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01576"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00510"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00515"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00718"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00399"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_1"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-017-5002-5"},{"key":"e_1_3_2_34_2","unstructured":"Yang Liu Samuel Albanie Arsha Nagrani and Andrew Zisserman. 2019. Use what you have: Video retrieval using representations from collaborative experts. arXiv:1907.13487. Retrieved from https:\/\/arxiv.org\/abs\/1907.13487"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICMEW.2019.00126"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","unstructured":"Ze Liu Yutong Lin Yue Cao Han Hu Yixuan Wei Zheng Zhang Stephen Lin and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE\/CVF International Conference on Computer Vision . 10012\u201310022.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00817"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2023.105936"},{"issue":"1","key":"e_1_3_2_40_2","first-page":"105","article-title":"PaddlePaddle: An open-source deep learning platform from industrial practice","volume":"1","author":"Ma Yanjun","year":"2019","unstructured":"Yanjun Ma, Dianhai Yu, Tian Wu, and Haifeng Wang. 2019. PaddlePaddle: An open-source deep learning platform from industrial practice. Frontiers of Data and Domputing 1, 1 (2019), 105\u2013115.","journal-title":"Frontiers of Data and Domputing"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CBMI.2018.8516488"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3475722.3482793"},{"key":"e_1_3_2_43_2","unstructured":"Megha Nawhal and Greg Mori. 2021. Activity graph transformer for temporal action localization. arXiv:2101.08540. Retrieved from https:\/\/arxiv.org\/abs\/2101.08540"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-15552-9_29"},{"key":"e_1_3_2_45_2","volume-title":"Referee Biographical Information","year":"2012","unstructured":"Olympedia. 2012. Referee Biographical Information. Retrieved 13 December 2023 from http:\/\/www.olympedia.org\/athletes\/5004924"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00055"},{"key":"e_1_3_2_47_2","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems.","journal-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00269"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.119"},{"key":"e_1_3_2_50_2","article-title":"Two-stream convolutional networks for action recognition in videos","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. Proceedings of the 27th International Conference on Neural Information Processing Systems.","journal-title":"Proceedings of the 27th International Conference on Neural Information Processing Systems"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-09396-3_9"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.23919\/ICMU50196.2021.9638855"},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","unstructured":"Haisheng Su Weihao Gan Wei Wu Yu Qiao and Junjie Yan. 2021. Bsn++: Complementary boundary regressor with scale-balanced relation modeling for temporal action proposal generation. In Proceedings of the AAAI Conference on Artificial Intelligence 35 (2021) 2602\u20132610.","DOI":"10.1609\/aaai.v35i3.16363"},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","unstructured":"Haritha Thilakarathne Aiden Nibali Zhen He and Stuart Morgan. 2022. Pose is all you need: The pose only group activity recognition system (pogars). Machine Vision and Applications 33 6 (2022) 95.","DOI":"10.1007\/s00138-022-01346-2"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW50498.2020.00450"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1080\/21520704.2018.1528324"},{"issue":"10","key":"e_1_3_2_58_2","first-page":"5130","article-title":"Football match intelligent editing system based on deep learning","volume":"13","author":"Wang Bin","year":"2019","unstructured":"Bin Wang, Wei Shen, FanSheng Chen, and Dan Zeng. 2019. Football match intelligent editing system based on deep learning. KSII Transactions on Internet and Information Systems 13, 10 (2019), 5130\u20135143.","journal-title":"KSII Transactions on Internet and Information Systems"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01398"},{"key":"e_1_3_2_60_2","unstructured":"Limin Wang Yu Qiao Xiaoou Tang. 2014. Action recognition and detection by combining motion and appearance features. THUMOS14 Action Recognition Challenge 1 2 (2014) 2."},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2934824"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00813"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00194"},{"key":"e_1_3_2_65_2","doi-asserted-by":"crossref","unstructured":"Chen Wei Haoqi Fan Saining Xie Chao-Yuan Wu Alan Yuille and Christoph Feichtenhofer. 2022. Masked feature prediction for self-supervised visual pre-training. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 14668\u201314678.","DOI":"10.1109\/CVPR52688.2022.01426"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00023"},{"key":"e_1_3_2_67_2","doi-asserted-by":"crossref","unstructured":"Fei Wu Qingzhong Wang Jiang Bian Ning Ding Feixiang Lu Jun Cheng Dejing Dou and Haoyi Xiong. 2023. A survey on video action recognition in sports: Datasets methods and applications. IEEE Transactions on Multimedia . 25 (2023) 7943\u20137966.","DOI":"10.1109\/TMM.2022.3232034"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00713"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2019.2926078"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.293"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.337"},{"key":"e_1_3_2_73_2","first-page":"4694","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Ng Joe Yue-Hei","year":"2015","unstructured":"Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. 2015. Beyond short snippets: Deep networks for video classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4694\u20134702."},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2020.103870"},{"key":"e_1_3_2_75_2","unstructured":"Yi Zhu Xinyu Li Chunhui Liu Mohammadreza Zolfaghari Yuanjun Xiong Chongruo Wu Zhi Zhang Joseph Tighe R. Manmatha and Mu Li. 2020. A comprehensive study of deep video action recognition. arXiv:2012.06567. Retrieved from https:\/\/arxiv.org\/abs\/2012.06567"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3633516","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3633516","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:50:09Z","timestamp":1750287009000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3633516"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,11]]},"references-count":74,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,4,30]]}},"alternative-id":["10.1145\/3633516"],"URL":"https:\/\/doi.org\/10.1145\/3633516","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,11]]},"assertion":[{"value":"2023-03-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}