Li Li 0012

dblp:53/2189-12 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
8since 2021 · last 2025
0009-0005-6099-614XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 8 since 2021Computer networks · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Generating Microservice Graphs with Production Characteristics for Efficient Resource Scaling
abstract
A production microservice application can have multiple services with varying call graphs, and a microservice may be shared across different call graphs.Improving resource efficiency in such complex applications requires proper benchmarks, but production traces are often too large to be used in experiments.To this end, we propose a Service Dependency Graph Generator (DGG) that comprises a Data Handler and a Graph Generator, to generate service dependency graphs of benchmarks that incorporate production-level characteristics from traces.The data handler constructs fine-grained call graphs with dynamic interface and repeated calling features from the trace, and then clusters these call graphs based on the topological and invocation types.The graph generator uses a random graph model to simulate real microservice invocations, generating multiple call graphs and merging them into small-scale service dependency graphs with production-level characteristics.Case studies show that * Fanrong Du and Jiuchen Shi contributed equally to this work.
Fanrong Du, Jiuchen Shi, Quan Chen 0002, Pu Pang, Li Li 0012, Minyi Guo
ICS5
2025 Repurpose Accel-Sim for Next Generation NVIDIA Jetson GPU Architectural Design
abstract
The growing adoption of NVIDIA Jetson devices in edge-AI applications highlights the need for accurate architecture simulation tools on their integrated GPUs. Existing cycle-accurate GPU simulators primarily target traditional discrete GPUs and exhibit significant inaccuracies when applied to Jetson integrated GPUs. While Accel-Sim serves as the most widely used academic simulator for NVIDIA GPU research, its lack of support for the latest Jetson integrated GPUs severely hinders architectural exploration for next generation edge-AI devices.We propose Accel-Sim-J, which bridges the gap by repurposing Accel-Sim simulation framework to NVIDIA Jetson GPUs. We refine three major Accel-Sim framework components by applying tuner modifications, GPGPU-Sim performance model enhancements, and correlator adjustments. These improvements enable precise Jetson GPU simulation support, reducing simulation cycle errors from 29.0% to 22.7% on the Rodinia benchmark and from 26.1% to 16.1% on a transformer block. Furthermore, our enhanced architectural support for Ampere GPUs achieves a considerable reduction in simulation error (from 140.1% to 50.2%) for GEMM kernels.Based on Accel-Sim-J, we conduct a case study investigating the architecture design difference between an edge GPU and a traditional one. Specifically, we compare the optimal Compute-to-Cache (C2C) ratio by changing the L2 cache size of Jetson AGX Orin and RTX 3090. We conclude that Jetson GPUs demonstrate a higher optimal C2C ratio than discrete GPUs for the same workloads. We suggest that designers reduce the on-chip area proportion of the L2 cache in the next generation Jetson GPU design for better performance and efficiency.
Chao Li 0009, Xiaofeng Hou, Yaqian Zhao, Jingwen Leng, Li Li 0012, Minyi Guo
ISLPED7
2025 Reducing Load-Balancing Cost for Multithreading Applications on Asymmetric NUMA Machine
Yuhang Fang, Pu Pang, Quan Chen 0002, Li Li 0012, Minyi Guo
NPC (2)4
2023 AdaptGear: Accelerating GNN Training via Adaptive Subgraph-Level Kernels on GPUs
abstract
Graph neural networks (GNNs) are powerful tools for exploring and learning from graph structures and features. As such, achieving high-performance execution for GNNs becomes crucially important. Prior works have proposed to explore the sparsity (i.e., low density) in the input graph to accelerate GNNs, which uses the full-graph-level or block-level sparsity format. We show that they fail to balance the sparsity benefit and kernel execution efficiency. In this paper, we propose a novel system, referred to as AdaptGear, that addresses the challenge of optimizing GNNs performance by leveraging kernels tailored to the density characteristics at the subgraph level. Meanwhile, we also propose a method that dynamically chooses the optimal set of kernels for a given input graph. Our evaluation shows that AdaptGear can achieve a significant performance improvement, up to 6.49× (1.87× on average), over the state-of-the-art works on two mainstream NVIDIA GPUs across various datasets.
Yangjie Zhou 0001, Yaoxu Song, Jingwen Leng, Zihan Liu 0002, Weihao Cui, Zhendong Zhang 0004, Cong Guo 0003, Quan Chen 0002, Li Li 0012, Minyi Guo
CF9
2023 DistSim: A performance model of large-scale hybrid distributed DNN training
abstract
With the ever-increasing computational demand of DNN training workloads, distributed training has been widely adopted. A combination of data, model and pipeline parallelism strategy, called hybrid parallelism distributed training, is imported to tackle the problem of deploying large-scale models. However, how to evaluate the hybrid strategy and the utilization of each device remains a challenge since existing works either profile on a real large-scale cluster with high time and money costs or only analyze a specific type of parallelism without considering the hybrid parallelism. In this work, we proposed DistSim, an event-based performance model to accurately analyze each device's computation and communication activities with low profiling costs. DistDim breaks down the model into events according to the given distributed strategy, which can be profiled on two nodes. Then DistSim leverages the hierarchy of different parallel strategies to generate the computation and communication event-flow from layer level to model level and finally the activity timeline of each device participating in training. Experiment shows that DistSim can reach <4% errors when predicting distributing training batch time and <5% errors when predicting a single device's activity time in various hybrid strategy settings. We also provide a use-case of DistSim, automatically evaluate and search the best distributed training strategy, and find a hybrid strategy with at most 7.37× throughput improvement.
Guandong Lu, Runzhe Chen, Yakai Wang, Yangjie Zhou 0001, Rui Zhang 0040, Zheng Hu 0002, Yanming Miao, Zhifang Cai, Li Li 0012, Jingwen Leng, Minyi Guo
CF9
2023 MMExit: Enabling Fast and Efficient Multi-modal DNN Inference with Adaptive Network Exits
Xiaofeng Hou, Jiacheng Liu 0001, Xuehan Tang, Chao Li 0009, Kwang-Ting Cheng, Li Li 0012, Minyi Guo
Euro-Par6
2021 Falcon: Addressing Stragglers in Heterogeneous Parameter Server Via Multiple Parallelism
abstract
The parameter server architecture has shown promising performance advantages when handling deep learning (DL) applications. One crucial issue in this regard is the presence of stragglers, which significantly retards DL training progress. Previous solutions for solving stragglers may not fully exploit the computation resource of the cluster as evidenced by our experiments, especially in the heterogeneous environment. This motivates us to design a heterogeneity-aware parameter server paradigm that addresses stragglers and accelerates DL training from the perspective of computation parallelism. We introduce a novel methodology named straggler projection to give a comprehensive inspection of stragglers and reveal practical guidelines to solve this problem in two aspects: (1) controlling each worker's training speed via elastic training parallelism control and (2) transferring blocked tasks from stragglers to pioneers to fully utilize the computation resource. Following these guidelines, we propose the abstraction of parallelism as an infrastructure and design the Elastic-Parallelism Synchronous Parallel (EPSP) algorithm to handle distributed training and parameter synchronization, supporting both enforcedand slack-synchronization schemes. The whole idea has been implemented into a prototype called Falcon which effectively accelerates the DL training speed with the presence of stragglers. Evaluation under various benchmarks with baseline comparison demonstrates the superiority of our system. Specifically, Falcon reduces the training convergence time, by up to 61.83, 55.19, 38.92, and 23.68 percent shorter than FlexRR, Sync-opt, ConSGD, and DynSGD, respectively.
Qihua Zhou, Song Guo 0001, Haodong Lu 0001, Li Li 0012, Minyi Guo, Yanfei Sun, Kun Wang 0005
IEEE Trans. Computers4
2021 Petrel: Heterogeneity-Aware Distributed Deep Learning Via Hybrid Synchronization
abstract
The parameter server (PS) paradigm has achieved great success in deploying large-scale distributed Deep Learning (DL) systems. However, these systems implicitly assume that the cluster is homogeneous and this assumption does not hold in many realworld cases. Although the previous efforts are paid to address heterogeneity, they mainly prioritize the contribution of fast workers and reduce the involvement of slow workers, resulting in the limitations of workload imbalance and computation inefficiency. We reveal that grouping workers into communities, an abstraction proposed by us, and handling parameter synchronization at the community level can conquer these limitations and accelerate the training convergence progress. The inspiration of community comes from our exploration of prior knowledge about the similarity between workers, which is often neglected by previous work. These observations motivate us to propose a new synchronization mechanism named Community-aware Synchronous Parallel (CASP), which uses the Asynchronous Advantage Actor-Critic (A3C)-based algorithm to intelligently determine community configuration and fully improve the synchronization performance. The whole idea has been implemented in a prototype system called Petrel that achieves a good balance between convergence efficiency and communication overhead. The evaluation under various benchmarks with multiple metrics and baseline comparison demonstrates the effectiveness of Petrel. Specifically, Petrel accelerates the training convergence speed by up to 1.87 x faster and reduces communication traffic by up to 26.85 percent, on average, over the non-community synchronization mechanisms.
Qihua Zhou, Song Guo 0001, Zhihao Qu, Peng Li 0017, Li Li 0012, Minyi Guo, Kun Wang 0005
IEEE Trans. Parallel Distributed Syst.5
2020 Petrel: Community-aware Synchronous Parallel for Heterogeneous Parameter Server
abstract
As to address the impact of heterogeneity in distributed Deep Learning (DL) systems, most previous approaches focus on prioritizing the contribution of fast workers and reducing the involvement of slow workers, incurring the limitations of workload imbalance and computation inefficiency. We reveal that grouping workers into communities, an abstraction proposed by us, and handling parameter synchronization in community level can conquer these limitations and accelerate the training convergence progress. The inspiration of community comes from our exploration of prior knowledge about the similarity between workers, which is often neglected by previous work. These observations motivate us to propose a new synchronization mechanism named Community-aware Synchronous Parallel (CSP), which uses the Asynchronous Advantage Actor-Critic (A3C), a Reinforcement Learning (RL) based algorithm, to intelligently determine community configuration and fully improve the synchronization performance. The whole idea has been implemented in a system called Petrel that achieves a good balance between convergence efficiency and communication overhead. The evaluation under different benchmarks demonstrates our approach can effectively accelerate the training convergence speed and reduce synchro-nization traffic.
Qihua Zhou, Song Guo 0001, Peng Li 0017, Yanfei Sun, Li Li 0012, Minyi Guo, Kun Wang 0005
ICDCS5
2019 Ebird: Elastic Batch for Improving Responsiveness and Throughput of Deep Learning Services
abstract
GPUs have been widely adopted to serve online deep learning-based services that have stringent QoS requirements. However, emerging deep learning serving systems often result in long latency and low throughput of the inference requests that damage user experience and increase the number of GPUs required to host an online service. Our investigation shows that the poor batching operation and the lacking of data transfer-computation overlap are the root causes of the long latency and low throughput. To this end, we propose Ebird, a deep learning serving system that is comprised of a GPU-resident memory pool, a multi-granularity inference engine, and an elastic batch scheduler. The memory pool eliminates the unnecessary waiting of the batching operation and enables data transfer-computation overlap. The inference engine enables concurrent execution of different batches, improving the GPUs resource utilization. The batch scheduler organizes inference requests elastically. Our experimental results on an Nvidia Titan RTX GPU show that Ebird reduces the response latency of inferences by up to 70.9% and improves the throughput by up to 49.3% while guaranteeing the QoS target compared with TensorFlow Serving.
Weihao Cui, Mengze Wei, Quan Chen 0002, Xiaoxin Tang, Jingwen Leng, Li Li 0012, Mingyi Guo
ICCD6
2019 Falcon: Towards Computation-Parallel Deep Learning in Heterogeneous Parameter Server
abstract
Parameter server paradigm has shown great performance superiority for handling deep learning (DL) applications. One crucial issue in this regard is the presence of stragglers, which significantly retards DL training progress. Previous solutions for solving straggler may not fully exploit the computation capacity of a cluster as evidenced by our experiments. This motivates us to make an attempt at building a new parameter server architecture that mitigates and addresses stragglers in heterogeneous DL from the perspective of computation parallelism. We introduce a novel methodology named straggler projection to give a comprehensive inspection of stragglers and reveal practical guidelines for resolving this problem: (1) reducing straggler emergence frequency via elastic parallelism control and (2) transferring blocked tasks to pioneer workers for fully exploiting cluster computation capacity. Following the guidelines, we propose the abstraction of parallelism as an infrastructure and elaborate the Elastic-Parallelism Synchronous Parallel (EPSP) that supports both enforced-and slack-synchronization schemes. The whole idea has been implemented in a prototype called Falcon which efficiently accelerates the DL training progress with the presence of stragglers. Evaluation under various benchmarks with baseline comparison evidences the superiority of our system. Specifically, Falcon yields shorter convergence time, by up to 61.83%, 55.19%, 38.92% and 23.68% reduction over FlexRR, Sync-opt, ConSGD and DynSGD, respectively.
Qihua Zhou, Kun Wang 0005, Song Guo 0001, Haodong Lu 0001, Li Li 0012, Minyi Guo, Yanfei Sun
ICDCS5
2019 Themis: Predicting and Reining in Application-Level Slowdown on Spatial Multitasking GPUs
abstract
Predicting performance degradation of a GPU application when it is co-located with other applications on a spatial multitasking GPU without prior application knowledge is essential in public Clouds. Prior work mainly targets CPU co-location, and is inaccurate and/or inefficient for predicting performance of applications at co-location on spatial multitasking GPUs. Our investigation shows that hardware event statistics caused by co-located applications, which can be collected with negligible overhead, strongly correlate with their slowdowns. Based on this observation, we present Themis, an online slowdown predictor that can precisely and efficiently predict application slowdown without prior application knowledge. We first train a precise slowdown model offline using hardware event statistics collected from representative co-locations. When new applications co-run, Themis collects event statistics and predicts their slowdowns simultaneously. Our evaluation shows that Themis has negligible runtime overhead and can precisely predict application-level slowdown with prediction error smaller than 9.5%. Based on Themis, we also implement an SM allocation engine to rein in application slowdown at co-location. Case studies show that the engine successfully enforces fair sharing and QoS.
Wenyi Zhao, Quan Chen 0002, Jingwen Leng, Chao Li 0009, Wenli Zheng, Li Li 0012, Minyi Guo
IPDPS8
2012 Recommendation of More Interests Based on Collaborative Filtering
abstract
Collaborative Filtering is one of the most important techniques in recommender systems. Current researches on Collaborative Filtering focus on how to improve the accuracy. However, it is of the same importance to recommend more potential interests to users because many of them have more expectations for recommendation list besides the accuracy. Current recommender systems did not address this problem. This paper focuses on how to help users find more interests in the recommendation list. We propose an sampling-based algorithm Probabilistic Top-N Selection to recommend potential interests for users, and propose two metrics, average predicted rating and category coverage, to assess the quality of the recommendation list. Then we conduct a series of experiments on Movie Lens dataset, experimental results demonstrate that our algorithm can significantly improve user experience through providing them with more potential interests.
Feilong Tang 0001, Li Li 0012, Leonard Barolli, Ilsun You, Huakang Li
AINA3
2011 Resource Allocation for OFDMA-Based Cognitive Radio Systems with Primary User Activity Consideration
abstract
In OFDMA-based Cognitive Radio (CR) systems, how to deal with the time-varying nature of avaliliable spectrum resources has become a hotspot and challenging problem in resource allocation. Traditional resource allocation algorithm can not guarantee proportional rates among non-real-time CR user (CRU) because the number of available subchannels is smaller than the number of CRUs in some OFDM symbol durations. In this paper, taking the maximizing of the sum-rate of all CRUs as the optimization objective, we propose a resource allocation algorithm in OFDMA-based CR systems using dual methods which can maintain statistical proportional rates among CRUs while keeping the interference introduced to Primary users (PU) under the specified thresholds. In contrast to traditional resource allocation algorithms, the proposed algorithm can achieve higher transmission rate and guarantee non-real-time services of CRUs.
Li Li 0012
ICC1
2009 Transaction Management for Reliable Grid Applications
abstract
Transaction management in Grids is responsible for ensuring the reliable execution of inherently distributed Grid applications. Grid transaction management is different from existing distributed transaction models because Grid resources are highly autonomous, dynamic and heterogeneous. This paper proposes a Grid transaction service (GridTS) and coordination algorithms that manage short-lived and long-lived Grid transactions respectively, providing reliability support for Grid applications. Unlike existing long-lived transaction models that require application programmers to develop compensating transactions, the GridTS can automatically generate compensating transactions during the execution of long-lived Grid transactions. The feasibility of GridTS and the effectiveness of proposed coordination algorithms are demonstrated through simulation studies.
Feilong Tang 0001, Minyi Guo, Minglu Li 0001, Li Li 0012
AINA4
2009 I-Cache Tag Reduction for Low Power Chip Multiprocessor
abstract
Energy consumption is a major consideration in microprocessor optimization. This paper presents a tag-reduction based approach for energy saving in L1 i-cache (instruction cache) of chip multiprocessors (CMP). To our best knowledge, this is the first work that extends the tag reduction technique to the CMP. We formulate our approach to an equivalent problem which is to find an assignment of the whole instruction pages in the physical memory to a set of cores such that the tag reduction conflicts for each core can be mostly avoided or reduced. We then propose three algorithms using different heuristics for this assignment problem. The experimental results show that our proposed algorithms can save the total power up to 45.33% in average compared to the one that the tag reduction is not used. They outperform significantly the tag reduction based algorithm on single core processor as well.
Long Zheng 0001, Mianxiong Dong, Song Guo 0001, Minyi Guo, Li Li 0012
ISPA5
2009 Editorial: Special issue on mobile P2P networking and computing
Li Li 0012, Jiannong Cao 0001, Kurt Tutschku
Peer-to-Peer Netw. Appl.1