VLDB 2026 Research / reviewers in the wild / expert
Mingzhen Li 0001
dblp:230/3070-1
· DBLP profile ↗
25ranked-venue papers
7as first author
23since 2021 · last 2025
0000-0002-4115-9072ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FastSpMM: Leveraging Tensor Cores for Sparse Matrix Multiplication
Mingzhen Li 0001, Weile Jia, Hailong Yang 0002, Guangming Tan |
CF | 2 |
| 2025 | Efficient Long Context Fine-tuning with Chunk FlowabstractLong context fine-tuning of large language models(LLMs) involves training on datasets that are predominantly composed of short sequences and a small proportion of longer sequences. However, existing approaches overlook this long-tail distribution and employ training strategies designed specifically for long sequences. Moreover, these approaches also fail to address the challenges posed by variable sequence lengths during distributed training, such as load imbalance in data parallelism and severe pipeline bubbles in pipeline parallelism. These issues lead to suboptimal training performance and poor GPU resource utilization. To tackle these problems, we propose a chunk-centric training method named ChunkFlow. ChunkFlow reorganizes input sequences into uniformly sized chunks by consolidating short sequences and splitting longer ones. This approach achieves optimal computational efficiency and balance among training inputs. Additionally, ChunkFlow incorporates a state-aware chunk scheduling mechanism to ensure that the peak memory usage during training is primarily determined by the chunk size rather than the maximum sequence length in the dataset. Integrating this scheduling mechanism with existing pipeline scheduling algorithms further enhances the performance of distributed training. Experimental results demonstrate that, compared with Megatron-LM, ChunkFlow can be up to 4.53x faster in the long context fine-tuning of LLMs. Furthermore, we believe that ChunkFlow serves as an effective solution for a broader range of scenarios, such as long context continual pre-training, where datasets contain variable-length sequences. Xiulong Yuan, Hongtao Xu, Wenting Shen, Ang Wang, Xiafei Qiu, Jie Zhang 0135, Yuqiong Liu, Bowen Yu 0002, Junyang Lin, Mingzhen Li 0001, Weile Jia, Yong Li 0045, Wei Lin 0016 |
ICML | 10 |
| 2025 | ESC: Effective Submanifold Convolution using Tensor CoresabstractSubmanifold convolution is an effective method to process 3D point cloud data, playing a significant role in fields such as robotics, autonomous driving, and AR/VR. However, due to the high sparsity and irregularity of point cloud data, it is challenging to accelerate submanifold convolution on modern GPUs, especially using tensor cores. Previous works have proposed implicit GEMM methods to accelerate submanifold convolution on GPU. However, the performance of such methods is limited by massive redundant computation and suboptimal parameter configurations. In this paper, we propose ESC, a new method to leverage GPU tensor cores for accelerating submanifold convolution with improved performance. Firstly, we propose an online similarity-aware reordering method to increase the point cloud data locality and yield more opportunities for eliminating redundancy. Secondly, we propose TC-aware redundancy elimination to reduce the redundant computation at the fine TC-tile granularity. Moreover, we propose an adaptive configuration selector to select the optimal configuration based on offline profiling results and online input data. Experimental results demonstrate that ESC outperforms the state-of-the-art works on representative datasets. Hailong Yang 0002, Xin You 0001, Yufan Xu 0001, Kaige Zhang 0002, Mingzhen Li 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
ICPP | 8 |
| 2025 | Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data SchedulingabstractLong-context supervised fine-tuning (Long-SFT) plays a vital role in enhancing the performance of large language models (LLMs) on long-context tasks. To smoothly adapt LLMs to long-context scenarios, this process typically entails training on mixed datasets containing both long and short sequences. However, this heterogeneous sequence length distribution poses significant challenges for existing training systems, as they fail to simultaneously achieve high training efficiency for both long and short sequences, resulting in sub-optimal end-to-end system performance in Long-SFT.
In this paper, we present a novel perspective on data scheduling to address the challenges posed by the heterogeneous data distributions in Long-SFT. We propose Skrull, a dynamic data scheduler specifically designed for efficient long-SFT. Through dynamic data scheduling, Skrull balances the computation requirements of long and short sequences, improving overall training efficiency. Furthermore, we formulate the scheduling process as a joint optimization problem and thoroughly analyze the trade-offs involved. Based on those analysis, Skrull employs a lightweight scheduling algorithm to achieve near-zero cost online scheduling in Long-SFT. Finally, we implement Skrull upon DeepSpeed, a state-of-the-art distributed training system for LLMs. Experimental results demonstrate that Skrull outperforms DeepSpeed by 3.76x on average (up to 7.54x) in real-world long-SFT scenarios. Hongtao Xu, Wenting Shen, Yuanxin Wei, Ang Wang, Guo Runfan, Tianxing Wang 0006, Yong Li 0045, Mingzhen Li 0001, Weile Jia |
NeurIPS | 8 |
| 2025 | Mario: Near Zero-cost Activation Checkpointing in Pipeline ParallelismabstractLarge language models have to be trained in parallel due to their large number of parameters and significant memory footprint. Among various parallelism techniques, pipeline parallelism is widely adopted in inter-node scenarios with minimal communication overhead. However, state-of-the-art pipeline schemes lead to extra and imbalanced memory footprints, leaving room for further improvement. In this paper, we propose Mario, a pipeline optimizer that automatically tessellates activation checkpointing to existing pipeline schemes, enabling training larger models (or longer sequences) with less and balanced memory footprint across GPUs and improved GPU utilization. First, the activation recomputation can be effectively overlapped in the bubbles by moving it earlier in the execution process, thereby improving overall efficiency. With eliminated memory footprint through checkpointing, Mario allows for preposing more forward computation into the pipeline bubbles, making more room for further overlapping with greater flexibility, and thus exploiting the bubbles. Then we design a lightweight pipeline simulator to model execution behavior w/o|w/ Mario. Finally, we introduce an automatic pipeline scheduler specifically for Mario, capable of searching for near optimal combination of checkpointing and pipeline configurations within minutes. Experimental results on GPT3 and LLaMA2 models show that Mario can speed up existing state-of-the-art pipeline schemes (w/o|w/ checkpointing) including 1F1B, Chimera, and Interleave by 1.16×|1.57× on average. This work paves a new direction for effective low-cost pipeline training. Mingzhen Li 0001, Guangming Tan, Weile Jia |
PPoPP | 2 |
| 2025 | An interpretable DeePMD-kit performance model for emerging supercomputersabstractAbstract Deep potential (DP) scheme has increased the simulation temporal and spatial scales while maintaining the ab initio accuracy of the molecular dynamics. DeePMD-kit is an outstanding application that implements DP scheme efficiently. However, current performance model cannot accurately measure the resource utilization of DeePMD-kit operators and predict the execution time. We introduce DP-perf, an interpretable performance model for DeePMD-kit. DP-perf can accurately measure the resource utilization of the individual DeePMD-kit operators, communication pattern, and the overall application by exploiting physical system properties and machine configurations. It can be easily applied to mainstream supercomputers including Tianhe-3F, the new Sunway, Fugaku, and Summit. With DP-perf, users can select the optimal machine and decide the corresponding configuration for various purposes (e.g., lower cost, less time) without real runs. Evaluation of four top supercomputers shows that DP-perf can fit overall execution time with a low mean absolute percentage error of 5.7 %/8.1%/14.3%/13.1% on Tianhe-3F/new Sunway/Fugaku/Summit. On the prediction scenario, DP-perf can predict the total execution time with a mean absolute percentage error of less than 20%. Xiangyu Meng 0005, Xun Wang 0010, Mingzhen Li 0001, Guangming Tan, Weile Jia |
CCF Trans. High Perform. Comput. | 3 |
| 2025 | 29-Billion Atoms Molecular Dynamics Simulation With Ab Initio Accuracy on 35 Million Cores of New Sunway SupercomputerabstractPhysical phenomena such as bond breaking and phase transitions require molecular dynamics (MD) withab initioaccuracy, involving up to billions of atoms and over nanosecond timescales. Previous state-of-the-art work has demonstrated that neural network molecular dynamics (NNMD) like deep potential molecular dynamics (DeePMD), can successfully extend the temporal and spatial scales of MD withab initioaccuracy on both ARM and GPU platforms. However, the DeePMD-kit package is currently unable to fully exploit the computational potential of the new Sunway supercomputer due to its unique many-core architecture, memory hierarchy, and low precision capability. In this paper, we re-design the DeePMD-kit to harness the massive computing power of the new Sunway, enabling the MD with over ten billion atoms. We first design a large-scale parallelization scheme to exploit the massive parallelism of the new Sunway. Then we devise specialized optimizations for the time-consuming operators. Finally, we design a novel mixed precision method for DeePMD-kit customized operators to leverage the low precision computing power of the new Sunway. The optimized DeePMD-kit achieves 67.6 / 56.5$\boldsymbol{\times}$speedup for water / copper systems on the new Sunway. Meanwhile, it can perform 29 billion atoms simulation for the water system on 35 million cores (i.e., 90,000 computing nodes, around 84% of the whole supercomputer) with a peak performance of 57.1 PFLOPs, which is 7.9$\boldsymbol{\times}$bigger and 1.2$\boldsymbol{\times}$faster than state-of-the-art results. This paves the way for investigating more realistic scenarios, such as studying the mechanical properties of metals, semiconductor devices, batteries, and other materials and physical systems. Xun Wang 0010, Xiangyu Meng 0005, Zhuoqiang Guo, Mingzhen Li 0001, Mingfan Li, Ninghui Sun, Guangming Tan, Weile Jia |
IEEE Trans. Computers | 4 |
| 2024 | Accelerating Large-Scale Sparse LU Factorization for RF Circuit Simulation
Guofeng Feng, Zhuoqiang Guo, Mingzhen Li 0001, Zhou Jin 0001, Weile Jia, Guangming Tan, Ninghui Sun |
Euro-Par (3) | 4 |
| 2024 | Retrospection on the Performance Analysis Tools for Large-Scale HPC ProgramsabstractAs the performance gap between hardware and software widens, performance analysis tools are essential for understanding the behavior of large-scale High-Performance Computing (HPC) programs. These tools provide insights into the performance bottlenecks and help in optimizing the performance of the programs. In this paper, we present a comprehensive study of performance analysis tools for large-scale HPC systems including both sampling-based and instrumentation-based tools that are commonly adopted in the HPC community. We investigate the abundance and overheads of data collection as well as the analysis capabilities of HPCToolkit, TAU, and Scalasca with representative programs at scale. Our study shows that different performance analysis tools have distinct strengths and weaknesses, and the choice of a performance analysis tool depends on the specific requirements of the user. We also discuss the challenges and future directions in the field of performance analysis tools for large-scale HPC systems. Zhibo Xuan, Xin You 0001, Hailong Yang 0002, Mingzhen Li 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
HiPC | 4 |
| 2024 | Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per DayabstractPhysical phenomena such as chemical reactions, bond breaking, and phase transition require molecular dynamics (MD) simulation with ab initio accuracy ranging from milliseconds to microseconds. However, previous state-of-the-art neural network based MD packages such as DeePMD-kit can only reach 4.7 nanoseconds per day on the Fugaku supercomputer. In this paper, we present a novel node-based parallelization scheme to reduce communication by 81%, then optimize the computationally intensive kernels with sve-gemm and mixed precision. Finally, we implement intra-node load balance to further improve the scalability. Numerical results on the Fugaku supercomputer show that our work has significantly improved the time-to-solution of the DeePMD-kit by a factor of 31.7 x, reaching 149 nanoseconds per day on 12,000 computing nodes. This work has opened the door for millisecond simulation with ab initio accuracy within one week for the first time. Zhuoqiang Guo, Mingzhen Li 0001, Enji Li, Guojun Yuan, Zhan Wang 0003, Guangming Tan, Weile Jia |
SC | 4 |
| 2024 | Building a domain-specific compiler for emerging processors with a reusable approach
Mingzhen Li 0001, Yi Liu 0013, Bangduo Chen, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
Sci. China Inf. Sci. | 1 |
| 2024 | Towards optimized tensor code generation for deep learning on sunway many-core processor
Mingzhen Li 0001, Changxi Liu, Jianjin Liao, Xuegui Zheng, Hailong Yang 0002, Rujun Sun, Lin Gan 0001, Guangwen Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
Frontiers Comput. Sci. | 1 |
| 2024 | ElasticBatch: A Learning-Augmented Elastic Scheduling System for Batch Inference on MIGabstractAs deep learning (DL) technologies become ubiquitous, GPU clusters are deployed for inference tasks with consistent service level objectives (SLOs). Efficiently utilizing multiple GPUs is crucial for throughput and cost-effectiveness. This article addresses the challenges posed by dynamic input and NVIDIA MIG in scheduling DL workloads. We present ElasticBatch, a scheduling system that simplifies configuration through bucketization and employs a machine learning-based pipeline to optimize settings. Our experiments demonstrate that ElasticBatch achieves a 50% reduction in GPU instances compared to MIG disablement, increases GPU utilization by 1.4% to 6.5% over an ideal scheduler and significantly reduces profiling time. This research contributes to the discourse on efficient utilization of GPU clusters. ElasticBatch's effectiveness in mitigating challenges posed by dynamic inputs and NVIDIA MIG underscores its potential to optimize GPU cluster performance, providing tangible benefits in terms of reduced instances, increased utilization, and significant time savings in real-world deployment scenarios. Jiaxing Qi, Wencong Xiao, Mingzhen Li 0001, Chaojie Yang, Yong Li 0045, Wei Lin 0016, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | Exploiting Subgraph Similarities for Efficient Auto-tuning of Tensor ProgramsabstractThe requirement for deploying deep learning (DL) models efficiently has boosted the research of DL compilers. Especially, the difficulty of generating optimized tensor programs has driven DL compilers to commonly adopt the auto-tuning approaches. Consequently, there are increasing demands to improve the effectiveness of auto-tuning in terms of both search efficiency and search quality. However, existing auto-tuning approaches commonly treat subgraphs individually and overlook the similarities among them, and thus fail to generate better tensor programs under limited time budget. To address the above drawbacks, we propose FamilySeer, an auto-tuning framework that can generate better tensor programs by exploiting the subgraph similarities. Specifically, FamilySeer organizes similar subgraphs into subgraph families, where the cost models are built at family basis with improved accuracy for estimating high potential program candidates. To further leverage the similarity, FamilySeer uses the accurate cost model per family to reduce the number of program candidates for costly hardware measurements without degrading search quality. The experiment results on various DL models demonstrate that FamilySeer can achieve better search efficiency/quality on both CPU and GPU platforms compared to the state-of-the-art auto-tuning framework. Mingzhen Li 0001, Hailong Yang 0002, Shanjun Zhang, Fengwei Yu, Ruihao Gong, Yi Liu 0013, Zhongzhi Luan, Depei Qian 0001 |
ICPP | 1 |
| 2023 | Exploiting Input Tensor Dynamics in Activation Checkpointing for Efficient Training on GPUabstractLarger deep learning models usually lead to higher model quality, however with an ever-increasing GPU memory footprint. Although several tensor checkpointing techniques have been proposed to enable training under a restricted GPU memory budget, they fail to exploit the input tensor dynamics due to diverse datasets and subsequent data augmentation, and thus leave the training optimization on table. In this paper, we propose Mimose, an input-aware tensor checkpointing planner respecting the memory budget while enabling efficient model training on GPU. Mimose builds a lightweight but accurate prediction model of GPU memory usage online, without pre-analyzing the model. It generates a tensor checkpointing plan based on per-layer memory prediction and applies it to the training process on the fly. Our experiments show that Mimose achieves superior training throughput compared to state-of-the-art checkpointing frameworks under the same GPU memory budgets. Jianjin Liao, Mingzhen Li 0001, Hailong Yang 0002, Qingxiao Sun, Biao Sun 0002, Jiwei Hao, Tianyu Feng, Fengwei Yu, Shengdong Chen, Zhongzhi Luan, Depei Qian 0001 |
IPDPS | 2 |
| 2023 | EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsabstractDistributed synchronized GPU training is commonly used for deep learning. The resource constraint of using a fixed number of GPUs makes large-scale training jobs suffer from long queuing time for resource allocation, and lowers the cluster utilization. Adapting to resource elasticity can alleviate this but often introduces inconsistent model accuracy, due to lacking of capability to decouple model training procedure from resource allocation. We propose EasyScale, an elastic training system that achieves consistent model accuracy under resource elasticity for both homogeneous and heterogeneous GPUs. EasyScale preserves the data-parallel training behaviors strictly, traces the consistency-relevant factors carefully, utilizes the deep learning characteristics for EasyScaleThread abstraction and fast context-switching. To utilize heterogeneous cluster, EasyScale dynamically assigns workers based on the intra-/inter-job schedulers, minimizing load imbalance and maximizing aggregated job throughput. Deployed in an online serving cluster, EasyScale powers the training jobs to utilize idle GPUs opportunistically, improving overall cluster utilization by 62.1%. Mingzhen Li 0001, Wencong Xiao, Hailong Yang 0002, Biao Sun 0002, Shiru Ren, Zhongzhi Luan, Xianyan Jia, Yi Liu 0013, Yong Li 0045, Wei Lin 0016, Depei Qian 0001 |
SC | 1 |
| 2023 | Adapting combined tiling to stencil optimizations on sunway processor
Biao Sun 0002, Mingzhen Li 0001, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
CCF Trans. High Perform. Comput. | 2 |
| 2022 | Toward accelerated stencil computation by adapting tensor core unit on GPUabstractThe Tensor Core Unit (TCU) has been increasingly adopted on modern high performance processors, specialized in boosting the performance of general matrix multiplication (GEMM). Due to its highly optimized hardware design, TCU can significantly accelerate GEMM-based operations widely used in scientific as well as deep learning applications. However, there is few work exploiting TCU to accelerate non-GEMM operations such as stencil computation that is also important in the field of high performance computing. To the best of our knowledge, there is no previous work that adapts stencil computation to TCU efficiently by considering its unique characteristics. In this paper, we propose a new method called TCstencil to adapt TCU for accelerating stencil computation. Specifically, we re-design the stencil computation as a series of reduction and summation operations in order to leverage the computing power of TCU. In addition, we propose corresponding optimizations for better exploiting TCU and memory hierarchy on GPU. We evaluate our method with different stencils and input mesh sizes on NVIDIA A100 and V100 GPUs. The experiment results demonstrate our method can achieve superior performance compared to the state-of-the-art stencil optimization frameworks. Yi Liu 0013, Hailong Yang 0002, Jianjin Liao, Mingzhen Li 0001, Zhongzhi Luan, Depei Qian 0001 |
ICS | 5 |
| 2022 | CoGNN: Efficient Scheduling for Concurrent GNN Training on GPUsabstractGraph neural networks (GNNs) suffer from low GPU utilization due to frequent memory accesses. Existing concurrent training mechanisms cannot be directly adapted to GNNs because they fail to consider the impact of input irregularity. This requires pre-profiling the memory footprint of concurrent tasks based on input dimensions to ensure successful co-location on GPU. Moreover, massive training tasks generated from scenarios such as hyper-parameter tuning require flexible scheduling strategies. To address these problems, we propose CoGNN that enables efficient management of GNN training tasks on GPUs. Specifically, the CoGNN organizes the tasks in a queue and estimates the memory consumption of each task based on cost functions at operator basis. In addition, the CoGNN implements scheduling policies to generate task groups, which are iteratively submitted for execution. The experiment results show that the CoGNN can achieve shorter completion and queuing time for training tasks from diverse GNN models. Qingxiao Sun, Yi Liu 0013, Hailong Yang 0002, Ruizhe Zhang 0012, Ming Dun, Mingzhen Li 0001, Wencong Xiao, Yong Li 0020, Zhongzhi Luan, Depei Qian 0001 |
SC | 6 |
| 2022 | QoS-aware dynamic resource allocation with improved utilization and energy efficiency on GPU
Qingxiao Sun, Liu Yi, Hailong Yang 0002, Mingzhen Li 0001, Zhongzhi Luan, Depei Qian 0001 |
Parallel Comput. | 4 |
| 2021 | PriPro: Towards Effective Privacy Protection on Edge-Cloud System running DNN InferenceabstractThe huge computation demand for deep learning models and limited computation resources on the edge devices calls for the cooperation between the edge device and cloud service. On a typical edge-cloud system accommodating DNN inference, a deep model is split into two partial models running on the edge device and the cloud service, respectively. The two partial models collaborate closely to satisfy the DNN inference requested by the user. However, user's privacy is vulnerable when transferring the intermediate results generated by the partial model at edge device to cloud service. Existing research works rely on metrics that are either impractical or insufficient to measure the effectiveness of privacy protection methods in the above scenario, especially from a single input aspect. In this paper, we first thoroughly analyze the state-of-the-art methods and drawbacks of existing methods from the aspects of both evaluation metrics and proposed techniques. Then, we propose a new metric system, including privacy accuracy (PA) and privacy index (PI), that can accurately measure the effectiveness of privacy protection methods. Furthermore, we propose PriPro, a privacy protection method that can dynamically inject noise to the intermediate results at various layers regarding the input features through the self-attention mechanism. The experiment results demonstrate our method outperforms existing methods for protecting user privacy on deep models such as AlexNet, VGG, and ResNet. Ruiyuan Gao 0001, Hailong Yang 0002, Shaohan Huang, Ming Dun, Mingzhen Li 0001, Zerong Luan, Zhongzhi Luan, Depei Qian 0001 |
CCGRID | 5 |
| 2021 | Automatic Code Generation and Optimization of Large-scale Stencil Computation on Many-core ProcessorsabstractStencil computation is an indispensable building block of many scientific applications and is widely used by the numerical solvers of partial differential equations (PDEs). Due to the complex computation patterns of different stencils and the various hardware targets (e.g., many-core processors), many domain-specific languages (DSLs) have been proposed to optimize stencil computation. However, existing stencil DSLs mostly focus on the performance optimizations on homogeneous many-core processors such as CPUs and GPUs, and fail to embrace emerging heterogeneous many-core processors such as Sunway. In addition, few of them can support expressing stencil with multiple time dependencies and optimizations from both spatial and temporal dimensions. Moreover, most stencil DSLs are unable to generate codes that can run efficiently in large scale, which limits their practical applicability. In this paper, we propose MSC, a new stencil DSL designed to express stencil computation in both spatial and temporal dimensions. It can generate high-performance stencil codes for large-scale execution on emerging many-core processors. Specially, we design several optimization primitives for improving parallelism and data locality, and a communication library for efficient halo exchange in large scale execution. The experiment results show that our MSC achieves better performance compared to the state-of-the-art stencil DSLs. Mingzhen Li 0001, Yi Liu 0013, Hailong Yang 0002, Yongmin Hu, Qingxiao Sun, Bangduo Chen, Xin You 0001, Zhongzhi Luan, Depei Qian 0001 |
ICPP | 1 |
| 2021 | The Deep Learning Compiler: A Comprehensive SurveyabstractThe difficulty of deploying various deep learning (DL) models on diverse DL hardware has boosted the research and development of DL compilers in the community. Several DL compilers have been proposed from both industry and academia such as Tensorflow XLA and TVM. Similarly, the DL compilers take the DL models described in different DL frameworks as input, and then generate optimized codes for diverse DL hardware as output. However, none of the existing survey has analyzed the unique design architecture of the DL compilers comprehensively. In this paper, we perform a comprehensive survey of existing DL compilers by dissecting the commonly adopted design in details, with emphasis on the DL oriented multi-level IRs, and frontend/backend optimizations. Specifically, we provide a comprehensive comparison among existing DL compilers from various aspects. In addition, we present detailed analysis on the design of multi-level IRs and illustrate the commonly adopted optimization techniques. Finally, several insights are highlighted as the potential research directions of DL compiler. This is the first survey paper focusing on the design architecture of DL compilers, which we hope can pave the road for future research towards DL compiler. Mingzhen Li 0001, Yi Liu 0013, Qingxiao Sun, Xin You 0001, Hailong Yang 0002, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Real-Time Polyp Detection for Colonoscopy Video on CPUabstractColorectal cancer(CRC), which is derived from polyps, is a common malignant tumor with a high incidence and mortality. Currently the colonoscopy is the most effective approach for early detection of the colorectal cancer. With the development of object detection, some polyp detection algorithms have achieved satisfied results. However, most of them need the support of hardware accelerators such as GPU, which is unsuitable for actual clinical environment due to its high power consumption and loud noise. To solve the problem, this paper proposes a solution to implement real-time polyp detection for colonoscopy video on general-purpose computing platforms, i.e. CPU. Our solution incorporates several methods to improve efficiency of polyp detection. Firstly, by analyzing the similarity and difference between successive frames in colonoscopy videos, we use the abrupt shot detection method to reduce the unnecessary detection for frames that contain non-polyps; secondly, for the frames that need to be detected, we iterated a procedure of sparse training, pruning and knowledge distillation to search for compressed models. Our approach is evaluated with both public datasets and actual colonoscopy videos, and results show that our approach can process 35.28 frames per second on colonoscopy videos with satisfied precision on CPU. Xuetong Li, Rui Liu 0020, Mingzhen Li 0001, Yi Liu 0013, Lianghui Jiang, Changhong Zhou |
ICTAI | 3 |
| 2020 | Accelerating Sparse Cholesky Factorization on Sunway Manycore ArchitectureabstractTo improve the performance of sparse Cholesky factorization, existing research divides the adjacent columns of the sparse matrix with the same nonzero patterns into supernodes for parallelization. However, due to the various structures of sparse matrices, the computation of the generated supernodes varies significantly, and thus hard to optimize when computed by dense matrix kernels. Therefore, how to efficiently map sparse Choleksy factorization to the emerging architectures, such as Sunway many-core processor, remains an active research direction. In this article, we propose swCholesky, which is a highly optimized implementation of sparse Cholesky factorization on Sunway processor. Specifically, we design three kernel task queues and a dense matrix library to dynamically adapt to the kernel characteristics and architecture features. In addition, we propose an auto-tuning mechanism to search for the optimal settings of the important parameters in swCholesky. Our experiments show that swCholesky achieves better performance than state-of-the-art implementations. Mingzhen Li 0001, Yi Liu 0013, Hailong Yang 0002, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |