{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T05:09:38Z","timestamp":1755839378282,"version":"3.41.2"},"reference-count":43,"publisher":"Wiley","issue":"26","license":[{"start":{"date-parts":[[2022,8,28]],"date-time":"2022-08-28T00:00:00Z","timestamp":1661644800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Concurrency and Computation"],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Although more layers and more parameters generally improve the accuracy of the models, such big models generally have high computational complexity and require big memory, which exceed the capacity of small devices for inference and incurs long training time. In addition, it is difficult to afford long training time and inference time of big models even in high performance servers, as well. As an efficient approach to compress a large deep model (a teacher model) to a compact model (a student model), knowledge distillation emerges as a promising approach to deal with the big models. Existing knowledge distillation methods cannot exploit the elastic available computing resources and correspond to low efficiency. In this paper, we propose an Elastic Deep Learning framework for knowledge Distillation, that is, EDL\u2010Dist. The advantages of EDL\u2010Dist are threefold. First, the inference and the training process is separated. Second, elastic available computing resources can be utilized to improve the efficiency. Third, fault\u2010tolerance of the training and inference processes is supported. We take extensive experimentation to show that the throughput of EDL\u2010Dist is up to 3.125 times faster than the baseline method (online knowledge distillation) while the accuracy is similar or higher.<\/jats:p>","DOI":"10.1002\/cpe.7272","type":"journal-article","created":{"date-parts":[[2022,8,29]],"date-time":"2022-08-29T00:48:45Z","timestamp":1661734125000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Large\u2010scale knowledge distillation with elastic heterogeneous computing resources"],"prefix":"10.1002","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4710-5697","authenticated-orcid":false,"given":"Ji","family":"Liu","sequence":"first","affiliation":[{"name":"Baidu Inc.  Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daxiang","family":"Dong","sequence":"additional","affiliation":[{"name":"Baidu Inc.  Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xi","family":"Wang","sequence":"additional","affiliation":[{"name":"Baidu Inc.  Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"An","family":"Qin","sequence":"additional","affiliation":[{"name":"Baidu Inc.  Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xingjian","family":"Li","sequence":"additional","affiliation":[{"name":"Baidu Inc.  Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Patrick","family":"Valduriez","sequence":"additional","affiliation":[{"name":"LIRMM Inria, University of Montpellier, CNRS  Montpellier France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dejing","family":"Dou","sequence":"additional","affiliation":[{"name":"Baidu Inc.  Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dianhai","family":"Yu","sequence":"additional","affiliation":[{"name":"Baidu Inc.  Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2022,8,28]]},"reference":[{"key":"e_1_2_10_2_1","unstructured":"VillegasR YangJ ZouY SohnS LinX LeeH.Learning to generate long\u2010term future via hierarchical prediction. Proceedings of Machine Learning Research; Vol.70 2017:3560\u20103569; PMLR."},{"key":"e_1_2_10_3_1","doi-asserted-by":"crossref","unstructured":"SzegedyC LiuW JiaY et al.Going deeper with convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR);2015:1\u20109; IEEE Computer Society.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_2_10_4_1","unstructured":"DevlinJ ChangMW LeeK ToutanovaK.BERT: pre\u2010training of deep bidirectional transformers for language understanding. Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL\u2010HLT);2019:4171\u20104186."},{"key":"e_1_2_10_5_1","unstructured":"SunY WangS FengS et al.ERNIE 3.0: large\u2010scale knowledge enhanced pre\u2010training for language understanding and generation. arXiv preprint arXiv:2107.02137 2021."},{"key":"e_1_2_10_6_1","doi-asserted-by":"crossref","unstructured":"CaruanaR LouY GehrkeJ KochP SturmM ElhadadN.Intelligible models for healthcare: predicting pneumonia risk and hospital 30\u2010day readmission. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining;2015:1721\u20101730.","DOI":"10.1145\/2783258.2788613"},{"key":"e_1_2_10_7_1","unstructured":"DevlinJ ChangMW LeeK ToutanovaK.BERT: pre\u2010training of deep bidirectional transformers for language understanding. CoRR 2018; abs\/1810.04805."},{"key":"e_1_2_10_8_1","doi-asserted-by":"crossref","unstructured":"BucilaC CaruanaR Niculescu\u2010MizilA.Model compression. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD);2006:535\u2010541.","DOI":"10.1145\/1150402.1150464"},{"key":"e_1_2_10_9_1","unstructured":"HintonG VinyalsO DeanJ.Distilling the knowledge in a neural network. Proceedings of the NeurIPS Deep Learning and Representation Learning Workshop;2015."},{"key":"e_1_2_10_10_1","unstructured":"GouJ YuB MaybankSJ TaoD.Knowledge distillation: a survey. CoRR 2020; abs\/2006.05525."},{"key":"e_1_2_10_11_1","unstructured":"MirzadehSI FarajtabarM LiA GhasemzadehH.Improved knowledge distillation via teacher assistant: bridging the gap between student and teacher. CoRR 2019; abs\/1902.03393."},{"issue":"3","key":"e_1_2_10_12_1","first-page":"1","article-title":"Knowledge distillation with attention for deep transfer learning of convolutional networks","volume":"16","author":"Li X","year":"2021","journal-title":"ACM Trans Knowl Discov Data (TKDD)"},{"key":"e_1_2_10_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-26253-2"},{"key":"e_1_2_10_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2020.3011952"},{"key":"e_1_2_10_15_1","doi-asserted-by":"crossref","unstructured":"DongD LiuJ WangX et al.Elastic deep learning using knowledge distillation with heterogeneous computing resources. Proceedings of the European Conference on Parallel Processing Workshop;2022:116\u2010128.","DOI":"10.1007\/978-3-031-06156-1_10"},{"key":"e_1_2_10_16_1","unstructured":"ZmoraN JacobG ZlotnikL ElhararB NovikG.Neural network distiller: a python package for DNN compression research. CoRR 2019; abs\/1910.12232."},{"key":"e_1_2_10_17_1","doi-asserted-by":"crossref","unstructured":"ZhangY XiangT HospedalesTM LuH.Deep mutual learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR);2018:4320\u20104328.","DOI":"10.1109\/CVPR.2018.00454"},{"key":"e_1_2_10_18_1","doi-asserted-by":"crossref","unstructured":"ChenD MeiJP WangC FengY ChenC.Online knowledge distillation with diverse peers. Proceedings of the Conference on Artificial Intelligence (AAAI);2020:3430\u20103437.","DOI":"10.1609\/aaai.v34i04.5746"},{"key":"e_1_2_10_19_1","unstructured":"ChengY WangD ZhouP ZhangT.A survey of model compression and acceleration for deep neural networks. CoRR 2017; abs\/1710.09282."},{"key":"e_1_2_10_20_1","unstructured":"AnilR PereyraG PassosA Orm\u00e1ndiR DahlGE HintonGE.Large scale distributed neural network training through online distillation. Proceedings of the International Conference on Learning Representations (ICLR);2018."},{"key":"e_1_2_10_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-816718-2.00008-7"},{"key":"e_1_2_10_22_1","doi-asserted-by":"crossref","unstructured":"NarayananD HarlapA PhanishayeeA et al.PipeDream: generalized pipeline parallelism for DNN training. Proceedings of the ACM Symposium on Operating Systems Principles (SOSP);2019:1\u201015.","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_2_10_23_1","first-page":"19","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Li M","year":"2014"},{"key":"e_1_2_10_24_1","unstructured":"LiuJ WuZ YuD et al.Heterps: distributed deep learning with reinforcement learning based scheduling in heterogeneous environments. arXiv preprint arXiv:2111.10635 2021."},{"key":"e_1_2_10_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/79173.79181"},{"key":"e_1_2_10_26_1","unstructured":"ZhuR YangS PfadlerA QianZ ZhouJ.Learning efficient parameter server synchronization policies for distributed SGD. Proceedings of the International Conference on Learning Representations (ICLR); 2020."},{"key":"e_1_2_10_27_1","unstructured":"HoQ CiparJ CuiH et al.More effective distributed ML via a stale synchronous parallel parameter server; 2013:1223\u20101231."},{"key":"e_1_2_10_28_1","first-page":"3049","volume-title":"Machine Learning Research","author":"Lian X","year":"2018"},{"key":"e_1_2_10_29_1","unstructured":"GibianskyA.Bringing HPC techniques to deep learning. Accessed August 12 2017.https:\/\/andrew.gibiansky.com\/blog\/machine\u2010learning\/baidu\u2010allreduce\/"},{"key":"e_1_2_10_30_1","unstructured":"LianX ZhangC ZhangH HsiehCJ ZhangW LiuJ.Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent. Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS); 2017:5330\u20105340."},{"key":"e_1_2_10_31_1","unstructured":"SergeevA BalsoMD.Horovod: fast and easy distributed deep learning in TensorFlow. arXiv preprint arXiv:1802.05799 2018."},{"key":"e_1_2_10_32_1","unstructured":"WuY MaK YanX LiuZ ChengJ.Elastic deep learning in multi\u2010tenant GPU cluster. CoRR. 2019; abs\/1909.11985."},{"key":"e_1_2_10_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2020.06.007"},{"key":"e_1_2_10_34_1","unstructured":"HuntP KonarM JunqueiraFP ReedB.ZooKeeper: wait\u2010free coordination for internet\u2010scale systems. Proceedings of the USENIX Annual Technical Conference; Vol. 11 2010."},{"issue":"1","key":"e_1_2_10_35_1","first-page":"105","article-title":"PaddlePaddle: an open\u2010source deep learning platform from industrial practice","volume":"1","author":"Ma Y","year":"2019","journal-title":"Front Data Comput"},{"key":"e_1_2_10_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2018.2867857"},{"key":"e_1_2_10_37_1","doi-asserted-by":"crossref","unstructured":"DengJ DongW SocherR LiLJ LiK LiFF.ImageNet: a large\u2010scale hierarchical image database. Proceedings IEEE Conference on Computer Vision and Pattern Recognition (CVPR);2009:248\u2010255.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_10_38_1","doi-asserted-by":"crossref","unstructured":"HeK ZhangX RenS SunJ.Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR);2016:770\u2010778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_10_39_1","doi-asserted-by":"crossref","unstructured":"KoonceB.MobileNetV3:125\u2010144;2021; Apress.","DOI":"10.1007\/978-1-4842-6168-2_11"},{"key":"e_1_2_10_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3298981"},{"key":"e_1_2_10_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-022-01664-x"},{"key":"e_1_2_10_42_1","doi-asserted-by":"crossref","unstructured":"ZhouC LiuJ JiaJ et al.Efficient device scheduling with multi\u2010job federated learning. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI):1\u20109;2022. To appear.","DOI":"10.1609\/aaai.v36i9.21235"},{"key":"e_1_2_10_43_1","doi-asserted-by":"crossref","unstructured":"ZhangH LiuJ JiaJ ZhouY DaiH.FedDUAP: federated learning with dynamic update and adaptive pruning using shared data on the server. Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI):1\u20107;2022. To appear.","DOI":"10.24963\/ijcai.2022\/385"},{"key":"e_1_2_10_44_1","doi-asserted-by":"crossref","unstructured":"LiG HuY ZhangM et al.FedHiSyn: a hierarchical synchronous federated learning framework for resource and data heterogeneity. Proceedings of the International Conference on Parallel Processing (ICPP);2022:1\u201010. To appear.","DOI":"10.1145\/3545008.3545065"}],"container-title":["Concurrency and Computation: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.7272","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/cpe.7272","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.7272","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,27]],"date-time":"2023-10-27T01:40:33Z","timestamp":1698370833000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/cpe.7272"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,28]]},"references-count":43,"journal-issue":{"issue":"26","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1002\/cpe.7272"],"URL":"https:\/\/doi.org\/10.1002\/cpe.7272","archive":["Portico"],"relation":{},"ISSN":["1532-0626","1532-0634"],"issn-type":[{"type":"print","value":"1532-0626"},{"type":"electronic","value":"1532-0634"}],"subject":[],"published":{"date-parts":[[2022,8,28]]},"assertion":[{"value":"2021-10-20","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-07-13","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-08-28","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e7272"}}