TY - GEN
T1 - A Dynamic Virtual Memory Management System for LLMs on AI Chips
AU - Wei, Gaolin
AU - Zhang, Zhaorui
AU - Xu, Jiaqi
AU - Zhang, Chen Jason
AU - Yao, Xin
AU - Liu, Benben
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025/12
Y1 - 2025/12
N2 - With the widespread application of large-scale DNN models, efficient training and inference with limited resources has become a popular research area. However, memory fragmentation is a significant barrier to efficient training and inference. In this article, we propose an efficient memory management strategy for execution on the Ascend platform based on the latest NPU virtual memory management APIs and develop efficient memory management logic for these. Indeed, the modification is mainly within the caching allocator of the NPU PyTorch extension, designing related garbage collection, allocation, splitting, reallocation, and fusing small fragmented blocks. In a four-card study on the Ascend 910B platform with training and inference accelerating architecture, MindSpeed, the greatest performance improvement was achieved by reducing the average fragmentation rate from 10% to 6%.
AB - With the widespread application of large-scale DNN models, efficient training and inference with limited resources has become a popular research area. However, memory fragmentation is a significant barrier to efficient training and inference. In this article, we propose an efficient memory management strategy for execution on the Ascend platform based on the latest NPU virtual memory management APIs and develop efficient memory management logic for these. Indeed, the modification is mainly within the caching allocator of the NPU PyTorch extension, designing related garbage collection, allocation, splitting, reallocation, and fusing small fragmented blocks. In a four-card study on the Ascend 910B platform with training and inference accelerating architecture, MindSpeed, the greatest performance improvement was achieved by reducing the average fragmentation rate from 10% to 6%.
KW - AI Chip
KW - LLM Training and Fine-Tuning
KW - Virtual Memory Management System
UR - https://www.scopus.com/pages/publications/105032507073
U2 - 10.1109/ICCD65941.2025.00062
DO - 10.1109/ICCD65941.2025.00062
M3 - Conference article published in proceeding or book
AN - SCOPUS:105032507073
T3 - Proceedings - IEEE International Conference on Computer Design: VLSI in Computers and Processors
SP - 389
EP - 392
BT - Proceedings - 2025 IEEE 43rd International Conference on Computer Design, ICCD 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 43rd International Conference on Computer Design, ICCD 2025
Y2 - 10 November 2025 through 12 November 2025
ER -