VLDB 2026 Research / reviewers in the wild / expert
Shigang Li 0002
dblp:24/6178-2
· DBLP profile ↗
51ranked-venue papers
13as first author
30since 2021 · last 2026
0000-0003-0022-7865ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 11 first-author · 26 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | C3FT: Computation-Centric Checkpointing for Distributed Large Model Training
Yonghua Huang, Baodong Wu, Shigang Li 0002, Jiahao Ding, Rongtian Fu, Qingping Li, Boxun Li, Zhenhua Zhu 0002 |
APPT | 3 |
| 2026 | Dynamo-MoE: Accelerating Sparse Large Model Inference with Dynamic ParallelizationabstractMixtral-of-Experts (MoE) has become one of the major model structures in LLMs because of its computational efficiency when scaling the model size. However, MoE model inference suffers from critical load imbalance issue caused by the sparsely and dynamically activated experts. In addition, current inference frameworks are oblivious to the real-time workload fluctuation, a common phenomenon in LLM serving. Therefore, the static model deployment of existing frameworks leads to severe performance limitations. To this end, we propose Dynamo-MoE, an out-of-box MoE inference framework to bridge the performance gap by dynamic parallelization strategies. Specifically, Dynamo-MoE integrates a novel load balancing approach based on token sorting and on-demand expert loading to solve the workload imbalance issue in the scenario of high workload (such as Prefill). Dynamo-MoE is also aware of workload varying to adaptively switch between tensor parallelism (for low latency in small batch scenarios) and expert parallelism (for high throughput in large batch scenarios). Furthermore, the model parameter redistribution overhead of dynamic parallelization is smartly overlapped through sophisticated pipeline orchestration. Compared to the SOTA framework vLLM (w/ and w/o EPLB), Dynamo-MoE achieves up to 6.75 × reduction for TTFT, 1.59 × reduction for TPOT, and 1.5 × improvement for throughput. Shigang Li 0002, Rongtian Fu, Tong Wu 0024, Jingkun Dong |
HPDC | 2 |
| 2026 | Omnia: Efficient RAG Serving through Speculative SchedulingabstractRetrieval-Augmented Generation (RAG) has emerged for enhancing Large Language Models (LLMs) by improving factual accuracy and mitigating hallucinations. A typical RAG pipeline executes in three cascaded stages: retrieval, reranking, and generation. The existing serving systems suffer from two critical system-level bottlenecks when applying to RAG serving: the first is the cumulative latency caused by rigid sequential dependencies between reranking and generation, and the second is the system saturation triggered by bursty, high fan-in reranking workloads. Rongtian Fu, Shigang Li 0002, Youxuan Xu, Tong Wu 0024, Jinliang Shi |
HPDC | 2 |
| 2025 | ParGNN: A Scalable Graph Neural Network Training Framework on multi-GPUsabstractFull-batch Graph Neural Network (GNN) training is indispensable for interdisciplinary applications. Although fullbatch training has advantages in convergence accuracy and speed, it still faces challenges such as severe load imbalance and high communication traffic overhead. In order to address these challenges, we propose ParGNN, an efficient full-batch training system for GNNs, which adopts a profiler-guided adaptive load balancing method along with graph over-partition to alleviate load imbalance. Based on the over-partition results, we present a subgraph pipeline algorithm to overlap communication and computation while maintaining the accuracy of GNN training. Extensive experiments demonstrate that ParGNN can not only obtain the highest accuracy but also reach the preset accuracy in the shortest time. In the end-to-end experiments performed on the four datasets, ParGNN outperforms the two state-of-theart full-batch GNN systems, PipeGCN and DGL, achieving the highest speedup of $2.7 \times$ and $21.8 \times$ times respectively. Junyu Gu, Shunde Li, Rongqiang Cao, Jue Wang 0013, Shigang Li 0002, Chunbao Zhou, Yangang Wang 0002, Xuebin Chi |
DAC | 8 |
| 2025 | CoLa: Towards Communication-efficient Distributed Sparse Matrix-Matrix Multiplication on GPUsabstractSparse Matrix-Matrix Multiplication (SpMM) is a critical operator in many applications, such as graph neural networks (GNNs).However, when SpMM is scaled to multiple GPUs, existing works face significant challenges: (1) massive redundant communication and (2) unawareness of heterogeneous links.To address the issue of massive redundant communication, we introduce a communication redundancy-free distributed SpMM algorithm that efficiently reutilizes fetched remote data to reduce communication volume.To tackle the unawareness of heterogeneous links, we propose two link-aware optimization techniques: communication fusion, which leverages local GPUs as embedding caches to reduce communication over slow links; and requestcoalesced communication, which coalesces necessary requested remote data into bulk transfers to maximize bandwidth utilization and minimize communication volume over proxy-based links.Based on these techniques, we develop CoLa, a highly communication-efficient distributed SpMM framework.Extensive evaluations on real-world datasets under different multi-GPU settings demonstrate that CoLa achieves geomean speedups of 8.56×, 9.12×, and 57.97× over CAGNET, MGGCN, and MGG, respectively. Lixing Zhang, Yingxia Shao, Shigang Li 0002 |
ICS | 3 |
| 2025 | FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresabstractSparse Matrix-matrix Multiplication (SpMM) and Sampled Dense-dense Matrix Multiplication (SDDMM) are important sparse operators in scientific computing and deep learning. Tensor Core Units (TCUs) enhance modern accelerators with superior computing power, which is promising to boost the performance of matrix operators to a higher level. However, due to the irregularity of unstructured sparse data, it is difficult to deliver practical speedups on TCUs. To this end, we propose FlashSparse, a novel approach to bridge the gap between sparse workloads and the TCU architecture. Specifically, FlashSparse minimizes the sparse granularity for SpMM and SDDMM on TCUs through a novel swap-and-transpose matrix multiplication strategy. Benefiting from the minimum sparse granularity, the computation redundancy is remarkably reduced while the computing power of TCUs is fully utilized. Besides, FlashSparse is equipped with a memory-efficient thread mapping strategy for coalesced data access and a sparse matrix storage format to save memory footprint. Extensive experimental results on H100 and RTX 4090 GPUs show that FlashSparse sets a new state-of-the-art for sparse matrix multiplications (geometric mean 5.5x speedup over DTC-SpMM and 3.22x speedup over RoDe). Jinliang Shi, Shigang Li 0002, Youxuan Xu, Rongtian Fu, Xueying Wang 0003, Tong Wu 0024 |
PPoPP | 2 |
| 2025 | Hypertron: Efficiently Scaling Large Models by Exploring High-Dimensional Parallelization SpaceabstractLarge models are evolving towards massive scale, diverse model architectures (dense and sparse) and long-context processing, which makes it very challenging to efficiently scale large models on parallel machines. The current widely-used parallelization strategies are often sub-optimal due to their limited parallelization strategy space. To this end, we propose Hypertron, a scalable parallel large-model training framework which incorporates an unprecedented high-dimensional (up to 7D) parallelization space, a holistic scheme for efficient dimension fusion, and a comprehensive performance model to guide the high-dimensional exploration. By exploiting the high-dimensional space to discover the optimal strategy which is not supported by existing frameworks, Hypertron significantly reduces memory and communication cost while improving parallel scalability. Extensive evaluations demonstrate that Hypertron achieves up to 56.7% Model FLOPs Utilization (MFU) on 2,048 new-generation Ascend NPU accelerators (scaling with supernodes) for different large models (such as sparse 141B and dense 310B), with 1.33x speedup over the best configuration of the state-of-the-art frameworks. Shigang Li 0002, Jingkun Dong, Jihao Chen, Zhongzhe Hu |
SC | 1 |
| 2025 | SparkAttention: high-performance multi-head attention for large models on Volta GPU architecture
Youxuan Xu, Tong Wu 0024, Shigang Li 0002, Xueying Wang 0003 |
CCF Trans. High Perform. Comput. | 3 |
| 2025 | OptiFX: Automatic Optimization for Convolutional Neural Networks with Aggressive Operator Fusion on GPUsabstractConvolutional Neural Networks (CNNs) are fundamental to advancing computer vision technologies. As CNNs become more complex and larger, optimizing model inference remains a critical challenge in both industry and academia. On modern GPU platforms, CNN operators are typically memory-bound, leading to significant performance degradation due to memory wall effects. While recent advancements have utilized operator fusion–merging multiple operators into one–to enhance inference performance, the fusion of multiple region-based operators like convolution is seldom addressed. This article introduces AFusion , a novel operator fusion technique aimed at improving inference performance, and OptiFX, an automatic optimization framework based on this approach. OptiFX employs a cost-based backtracking search to identify optimal sub-graphs for fusion and utilizes template-based code generation to create efficient kernels for these fused sub-graphs. We evaluate OptiFX across seven prominent CNN architectures–GoogLeNet, ResNet, DenseNet, MobileNet, SqueezeNet, NasNet, and UNet–on Nvidia A6000 Ada, RTX 4090, and Jetson AGX Orin platforms. Our results demonstrate that OptiFX significantly outperforms existing methods, achieving average speedups of \(2.91\times\) , \(3.30\times\) , and \(2.09\times\) in accelerating inference performance on these platforms, respectively. Xueying Wang 0003, Shigang Li 0002, Fan Luo 0003, Zhaoyang Hao, Tong Wu 0024, Ruiyuan Xu, Huimin Cui, Xiaobing Feng 0002, Guangli Li, Jingling Xue |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | HE-ASR-IT: Hybrid Excitation and Adaptive Style Recombination for Unpaired Image-to-Image TranslationabstractUnpaired image-to-image translation aims to translate an image from the source domain to the target domain without paired training data. Some recent works have applied the self-attention to this task and achieved impressive results. Nevertheless, it may generate unsatisfactory results for scenes with complex image content or strong geometric variations between image domains. In this paper, We propose a novel method based on hybrid excitation and adaptive style recombination for unpaired image-to-image translation, which has these main advantages. 1) A hybrid perceptual excitation module is used to capture different ranges of contextual information in complex scenes dynamically. 2) An adaptive cross-attention code recombination module is designed to recombine the content code and style code adaptively according to different positions. 3) We also propose multi-source NCE loss to constrain the generated image content by two aspects: image-level and patch-level. Experiments show that our method achieves more reasonable results than the state-of-the-art methods on several benchmark datasets. Juanjuan Luo, Mingxin Du, Shigang Li 0002, Wenbin Yao, Zhibin Huang |
IJCNN | 4 |
| 2024 | A High-Performance Design, Implementation, Deployment, and Evaluation of The Slim Fly Network
Nils Blach, Maciej Besta, Daniele De Sensi, Jens Domke, Hussein Harake, Shigang Li 0002, Patrick Iff, Marek Konieczny, Kartik Lakhotia, Ales Kubicek, Marcel Ferrari, Fabrizio Petrini, Torsten Hoefler |
NSDI | 6 |
| 2024 | POSTER: ParGNN: Efficient Training for Large-Scale Graph Neural Network on GPU ClustersabstractFull-batch graph neural network (GNN) training is essential for interdisciplinary applications. Large-scale graph data is usually divided into subgraphs and distributed across multiple compute units to train GNN. The state-of-the-art load balancing method based on direct graph partition is too rough to effectively achieve true load balancing on GPU clusters. We propose ParGNN, which employs a profiler-guided load balance workflow in conjunction with graph repartition to alleviate load imbalance and minimize communication traffic. Experiments have verified that ParGNN has the capability to scale to larger clusters. Shunde Li, Junyu Gu, Jue Wang 0013, Tiechui Yao, Yumeng Shi, Shigang Li 0002, Weiting Xi, Shushen Li, Chunbao Zhou, Yangang Wang 0002, Xuebin Chi |
PPoPP | 7 |
| 2024 | AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth CostabstractRecent advances in deep learning are driven by the growing scale of computation, data, and models. However, efficiently training large-scale models on distributed systems requires an intricate combination of data, operator, and pipeline parallelism, which exerts heavy burden on machine learning practitioners. To this end, we propose AutoDDL, a distributed training framework that automatically explores and exploits new parallelization schemes with near-optimal bandwidth cost. AutoDDL facilitates the description and implementation of different schemes by utilizing OneFlow'sSplit,Broadcast, andPartial Sum(SBP) abstraction. AutoDDL is equipped with an analytical performance model combined with a customized Coordinate Descent algorithm, which significantly reduces the scheme searching overhead. We conduct evaluations on Multi-Node-Single-GPU and Multi-Node-Multi-GPU machines using different models, including VGG and Transformer. Compared to the expert-optimized implementations, AutoDDL reduces the end-to-end training time by up to 31.1% and 10% for Transformer and up to 17.7% and 71.5% for VGG on the two parallel systems, respectively. Jinfan Chen, Shigang Li 0002, Jinhui Yuan, Torsten Hoefler |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | RDRM: Real-Time Dynamic Replica Management With Joint Optimization for Edge ComputingabstractThe combination of edge computing and replication technology provides service guarantee for edge applications. However, optimizing replica creation and placement to enhance system performance is challenging due to the limited resources available at the edge. In this context, effective replica management becomes crucial for efficient and reliable edge computing. This study proposes a real-time dynamic replica management model to address the challenges of replica creation and placement in the edge computing environment. Firstly, we design a prediction-based dynamic proactive replica creation algorithm. This algorithm integrates data popularity and node load, utilizing fuzzy membership functions to model data and node states, effectively handling the state uncertainty in certain conditions. It also defines overheating and undercooling similarities to assess the trend of state changes, thereby determining the optimal timing for replica creation. To prevent latency in replica creation, the algorithm employs an Long Short-Term Memory (LSTM) model with a deviation feedback mechanism, which helps prevent lag in replica creation and minimizes unnecessary replica generation. Secondly, we formula replica placement as a multi-objective optimization problem considering the node load and access degree. We use a joint optimization replica placement algorithm that combines Evolutionary Gradient Search (EGS) and Sorting Genetic Algorithm-II to solve the multi-objective replica placement problem. Finally, we conduct extensive experiments on the replica management model. The results demonstrate significant improvements in average response time, effective network utilization rate, storage space utilization rate, and system load balancing, which validate the effectiveness of the proposed method. Xikang Zhu, Wenbin Yao, Yingying Hou, Shigang Li 0002, Juanjuan Luo, Zhibin Huang, Shengdong Fu |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | Asynch-SGBDT: Train Stochastic Gradient Boosting Decision Trees in an Asynchronous Parallel MannerabstractGradient Boosting Decision Tree (GBDT) is a costly machine learning model. Current parallel GBDT algorithms generally follow a synchronous parallel design: Fork-join parallel manner, like MapReduce. Fork-join parallel manner needs considerable time. Thus, we propose whether synchronization is necessary for GBDT training and is asynchronous training manner efficient. In this paper, we solve the above problem by offering an asynchronous algorithm. We try to build a stochastic optimization problem by sampling, which shares the same output with original GBDT training problem and use asynchronous parallel SGD manner to train Gradient step GBDT. We name our algorithm as asynch-SGBDT. Our theoretical and experimental results indicate that compared with the serial GBDT training process, when the datasets’ high sample diversity is high and using Gradient step training GBDT, asynch-SGBDT does not slow down convergence speed on the epoch, and the sample diversity of current high-dimensional sparse datasets is usually high. We conduct experiments on a 32-node cluster using four different datasets. The results show that with LightGBM using a single worker as the baseline, LightGBM (the state-of-the-art synchronous parallel algorithm implement) on 32 workers achieves 5x-7x speedup, while our asynch-SGBDT on 32 workers increases the speedup to 11x-15x. Daning Cheng, Shigang Li 0002, Yunquan Zhang |
IPDPS | 2 |
| 2023 | A Scalable Hybrid Total FETI Method for Massively Parallel FEM SimulationsabstractThe Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method plays an important role in solving large-scale and complex engineering problems. This method needs to handle numerous matrix-vector multiplications. Directly calling the vendor-optimized library for general matrix-vector multiplication (gemv) on GPU leads to low performance, since it does not consider optimizations for different matrix sizes in HTFETI, i.e. different row and column sizes. In addition, state-of-the-art graph partitioning methods cannot guarantee load balancing for HTFETI, since the matrix size is determined by the length of the subdomain boundary. To solve the problems above, we first port gemv to the multi-stream pipeline scheme and develop a new batched kernel function on GPU, which brings 15%~30% throughput improvement and 37% average GFLOPs improvement, respectively. We also propose a multi-grained load-balancing scheme based on graph repartitioning and work-stealing, and the load imbalance ratio is down to 1.05~1.09 from 1.5. We have successfully applied the scalable HTFETI method to simulate the whole core assembly of China Experimental Fast Reactor (CEFR) for steady-state analysis, and the efficiencies of weak scalability and strong scalability reach 78% and 72% on 12,288 GPUs, respectively. As far as we know, this is the first time that HTFETI has been used in large-scale and high-fidelity whole core assembly simulation. Kehao Lin, Chunbao Zhou, Ningming Nie, Jue Wang 0013, Shigang Li 0002, Yangde Feng, Yangang Wang 0002, Kehan Yao, Tiechui Yao, Jian Wan 0001 |
PPoPP | 6 |
| 2023 | Co-design Hardware and Algorithm for Vector SearchabstractVector search has emerged as the foundation for large-scale information retrieval and machine learning systems, with search engines like Google and Bing processing tens of thousands of queries per second on petabyte-scale document datasets by evaluating vector similarities between encoded query texts and web documents. As performance demands for vector search systems surge, accelerated hardware offers a promising solution in the post-Moore's Law era. We introduce FANNS, an end-to-end and scalable vector search framework on FPGAs. Given a user-provided recall requirement on a dataset and a hardware resource budget, FANNS automatically co-designs hardware and algorithm, subsequently generating the corresponding accelerator. The framework also supports scale-out by incorporating a hardware TCP/IP stack in the accelerator. FANNS attains up to 23.0× and 37.2× speedup compared to FPGA and CPU baselines, respectively, and demonstrates superior scalability to GPUs, achieving 5.5× and 7.6× speedup in median and 95th percentile (P95) latency within an eight-accelerator configuration. The remarkable performance of FANNS lays a robust groundwork for future FPGA integration in data centers and AI supercomputers. Wenqi Jiang 0001, Shigang Li 0002, Johannes de Fine Licht, Zhenhao He, Runbin Shi, Cédric Renggli, Shuai Zhang 0007, Theodoros Rekatsinas, Torsten Hoefler, Gustavo Alonso |
SC | 2 |
| 2023 | ANT-MOC: Scalable Neutral Particle Transport Using 3D Method of Characteristics on Multi-GPU SystemsabstractThe Method Of Characteristic (MOC) to solve the Neutron Transport Equation (NTE) is the core of full-core simulation for reactors. High resolution is enabled by discretizing the NTE through massive tracks to traverse the 3D reactor geometry. However, the 3D full-core simulation is prohibitively expensive because of the high memory consumption and the severe load imbalance. To deal with these challenges, we develop ANT-MOC1. Specifically, we build a performance model for memory footprint, computation and communication, based on which a track management strategy is proposed to overcome the resolution bottlenecks caused by limited GPU memory. Furthermore, we implement a novel multi-level load mapping strategy to ensure load balancing among nodes, GPUs, and CUs. ANT-MOC enables a 3D full-core reactor simulation with 100 billion tracks on 16,000 GPUs, with 70.69% and 89.38% parallel efficiency for strong scalability and weak scalability, respectively. Shunde Li, Zongguo Wang, Lingkun Bu, Jue Wang 0013, Zhikuang Xin, Shigang Li 0002, Yangang Wang 0002, Yangde Feng, Peng Shi 0006, Xuebin Chi |
SC | 6 |
| 2023 | Large-Scale Simulation of Structural Dynamics Computing on GPU ClustersabstractStructural dynamics simulation plays an important role in research on reactor design and complex engineering. The Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method combined with Newmark method is an efficient way to solve large-scale structural dynamics problems. However, the sparse direct solver and the load imbalance caused by inconsistent density models are two critical issues limiting the performance and the scalability of structural dynamics computing. For the former, we propose an efficient variable-size batched method to accelerate SpMV on GPUs. For the latter, we establish an online performance prediction model, based on which we then design a novel inter-cluster subdomain fine-tuning algorithm to balance the workload of HTFETI parallel computing. We are the first to achieve the high-fidelity structural dynamics simulation of China Experimental Fast Reactor core assembly with up to 53.4 billion grids. The weak and strong scalability efficiencies reach 91.77% and 86.13% on 12,800 GPUs, respectively. Yumeng Shi, Ningming Nie, Jue Wang 0013, Kehao Lin, Chunbao Zhou, Shigang Li 0002, Kehan Yao, Shunde Li, Yangde Feng, Yangang Wang 0002 |
SC | 6 |
| 2023 | AGCM-3DLF: Accelerating Atmospheric General Circulation Model via 3-D Parallelization and Leap-FormatabstractThe atmospheric general circulation model (AGCM) has been an important research tool in the study of climate change for decades. As the demand for high-resolution simulation is becoming urgent, the scalability and simulation efficiency is faced with great challenges, especially for the latitude-longitude mesh-based models. In this paper, we propose a highly scalable 3-D atmospheric general circulation model based on leap-format, namely AGCM-3DLF. First, it utilizes a 3-D decomposition method allowing for parallelism release in all three physical dimensions. Then the leap-format difference computation scheme is adopted to maintain computational stability in grid updating and avoid additional filtering at the high latitudes. A novel shifting window communication algorithm is designed for parallelization of the unified model. Furthermore, a series of optimizations are conducted to improve the effectiveness of large-scale simulations. Experiment results in different platforms demonstrate good efficiency and scalability of the model. AGCM-3DLF scales up to the entire CAS-Xiandao1 supercomputer (196,608 CPU cores), attaining the speed of 11.1 simulation-year-per-day (SYPD) at a high resolution of 25KM. In addition, simulations conducted on the Sunway TaihuLight supercomputer exhibit a 1.06 million cores scalability with 36.1% parallel efficiency. He Zhang 0005, Yunquan Zhang, Baodong Wu, Kun Li 0016, Shigang Li 0002, Pengqi Lu, Junmin Xiao |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | A data-centric optimization framework for machine learningabstractRapid progress in deep learning is leading to a diverse set of quickly changing models, with a dramatically growing demand for compute. However, as frameworks specialize performance optimization to patterns in popular networks, they implicitly constrain novel and diverse models that drive progress in research. We empower deep learning researchers by defining a flexible and user-customizable pipeline for optimizing training of arbitrary deep neural networks, based on data movement minimization. The pipeline begins with standard networks in PyTorch or ONNX and transforms computation through progressive lowering. We define four levels of general-purpose transformations, from local intra-operator optimizations to global data movement reduction. These operate on a data-centric graph intermediate representation that expresses computation and data movement at all levels of abstraction, including expanding basic operators such as convolutions to their underlying computations. Central to the design is the interactive and introspectable nature of the pipeline. Every part is extensible through a Python API, and can be tuned interactively using a GUI. We demonstrate competitive performance or speedups on ten different networks, with interactive optimizations discovering new opportunities in EfficientNet. Oliver Rausch, Tal Ben-Nun, Nikoli Dryden, Andrei Ivanov, Shigang Li 0002, Torsten Hoefler |
ICS | 5 |
| 2022 | Near-optimal sparse allreduce for distributed deep learningabstractCommunication overhead is one of the major obstacles to train large deep learning models at scale. Gradient sparsification is a promising technique to reduce the communication volume. However, it is very challenging to obtain real performance improvement because of (1) the difficulty of achieving an scalable and efficient sparse allreduce algorithm and (2) the sparsification overhead. This paper proposes Ok-Topk, a scheme for distributed training with sparse gradients. Ok-Topk integrates a novel sparse allreduce algorithm (less than 6k communication volume which is asymptotically optimal) with the decentralized parallel Stochastic Gradient Descent (SGD) optimizer, and its convergence is proved. To reduce the sparsification overhead, Ok-Topk efficiently selects the top-k gradient values according to an estimated threshold. Evaluations are conducted on the Piz Daint supercomputer with neural network models from different deep learning domains. Empirical results show that Ok-Topk achieves similar model accuracy to dense allreduce. Compared with the optimized dense and the state-of-the-art sparse allreduces, Ok-Topk is more scalable and significantly improves training throughput (e.g., 3.29x-12.95x improvement for BERT on 256 GPUs). Shigang Li 0002, Torsten Hoefler |
PPoPP | 1 |
| 2022 | HammingMesh: A Network Topology for Large-Scale Deep LearningabstractNumerous microarchitectural optimizations unlocked tremendous processing power for deep neural networks that in turn fueled the AI revolution. With the exhaustion of such optimizations, the growth of modern AI is now gated by the performance of training systems, especially their data movement. Instead of focusing on single accelerators, we investigate data-movement characteristics of large-scale training at full system scale. Based on our workload analysis, we design HammingMesh, a novel network topology that provides high bandwidth at low cost with high job scheduling flexibility. Specifically, HammingMesh can support full bandwidth and isolation to deep learning training jobs with two dimensions of parallelism. Furthermore, it also supports high global bandwidth for generic traffic. Thus, HammingMesh will power future large-scale deep learning systems with extreme bandwidth requirements. Torsten Hoefler, Tommaso Bonato, Daniele De Sensi, Salvatore Di Girolamo, Shigang Li 0002, Marco Heddes, Jon Belk, Deepak Goel, Miguel Castro 0001, Steve Scott |
SC | 5 |
| 2022 | Efficient Quantized Sparse Matrix Operations on Tensor CoresabstractThe exponentially growing model size drives the continued success of deep learning, but it brings prohibitive computation and memory cost. From the algorithm perspective, model sparsification and quantization have been studied to alleviate the problem. From the architecture perspective, hardware vendors provide Tensor cores for acceleration. However, it is very challenging to gain practical speedups from sparse, low-precision matrix operations on Tensor cores, because of the strict requirements for data layout and lack of support for efficiently manipulating the low-precision integers. We propose Magicube, a high-performance sparse-matrix library for low-precision integers on Tensor cores. Magicube supports SpMM and SDDMM, two major sparse operations in deep learning with mixed precision. Experimental results on an NVIDIA A100show that Magicube achieves on average 1.44x (up to 2.37x) speedup over the vendor-optimized library for sparse kernels, and 1.43x speedup over the state-of-the-art with a comparable accuracy for end-to-end sparse Transformer inference. Shigang Li 0002, Kazuki Osawa, Torsten Hoefler |
SC | 1 |
| 2022 | VenusAI: An artificial intelligence platform for scientific discovery on supercomputers
Tiechui Yao, Jue Wang 0013, Meng Wan, Zhikuang Xin, Yangang Wang 0002, Rongqiang Cao, Shigang Li 0002, Xuebin Chi |
J. Syst. Archit. | 7 |
| 2021 | Asynchronous Decentralized SGD with Quantized and Local UpdatesabstractDecentralized optimization is emerging as a viable alternative for scalable distributed machine learning, but also introduces new challenges in terms of synchronization costs. To this end, several communication-reduction techniques, such as non-blocking communication, quantization, and local steps, have been explored in the decentralized setting. Due to the complexity of analyzing optimization in such a relaxed setting, this line of work often assumes \emph{global} communication rounds, which require additional synchronization. In this paper, we consider decentralized optimization in the simpler, but harder to analyze, \emph{asynchronous gossip} model, in which communication occurs in discrete, randomly chosen pairings among nodes. Perhaps surprisingly, we show that a variant of SGD called \emph{SwarmSGD} still converges in this setting, even if \emph{non-blocking communication}, \emph{quantization}, and \emph{local steps} are all applied \emph{in conjunction}, and even if the node data distributions and underlying graph topology are both \emph{heterogenous}. Our analysis is based on a new connection with multi-dimensional load-balancing processes. We implement this algorithm and deploy it in a super-computing environment, showing that it can outperform previous decentralized methods in terms of end-to-end training time, and that it can even rival carefully-tuned large-batch SGD for certain tasks. Giorgi Nadiradze, Amirmojtaba Sabour, Peter Davies-Peck, Shigang Li 0002, Dan Alistarh |
NeurIPS | 4 |
| 2021 | Chimera: efficiently training large-scale neural networks with bidirectional pipelinesabstractTraining large deep learning models at scale is very challenging. This paper proposes Chimera, a novel pipeline parallelism scheme which combines bidirectional pipelines for efficiently training large-scale models. Chimera is a synchronous approach and therefore no loss of accuracy, which is more convergence-friendly than asynchronous approaches. Compared with the latest synchronous pipeline approach, Chimera reduces the number of bubbles by up to 50%; benefiting from the sophisticated scheduling of bidirectional pipelines, Chimera has a more balanced activation memory consumption. Evaluations are conducted on Transformer based language models. For a GPT-2 model with 1.3 billion parameters running on 2,048 GPU nodes of the Piz Daint supercomputer, Chimera improves the training throughput by 1.16x-2.34x over the state-of-the-art synchronous and asynchronous pipeline approaches. Shigang Li 0002, Torsten Hoefler |
SC | 1 |
| 2021 | Flare: flexible in-network allreduceabstractThe allreduce operation is one of the most commonly used communication routines in distributed applications. To improve its bandwidth and to reduce network traffic, this operation can be accelerated by offloading it to network switches, that aggregate the data received from the hosts, and send them back the aggregated result. However, existing solutions provide limited customization opportunities and might provide suboptimal performance when dealing with custom operators and data types, with sparse data, or when reproducibility of the aggregation is a concern. To deal with these problems, in this work we design a flexible programmable switch by using as a building block PsPIN, a RISC-V architecture implementing the sPIN programming model. We then design, model, and analyze different algorithms for executing the aggregation on this architecture, showing performance improvements compared to state-of-the-art approaches. Daniele De Sensi, Salvatore Di Girolamo, Saleh Ashkboos, Shigang Li 0002, Torsten Hoefler |
SC | 4 |
| 2021 | Why Dataset Properties Bound the Scalability of Parallel Machine Learning Training AlgorithmsabstractAs the training dataset size and the model size of machine learning increase rapidly, more computing resources are consumed to speedup the training process. However, the scalability and performance reproducibility of parallel machine learning training, which mainly uses stochastic optimization algorithms, are limited. In this paper, we demonstrate that the sample difference in the dataset plays a prominent role in the scalability of parallel machine learning algorithms. We propose to use statistical properties of dataset to measure sample differences. These properties include the variance of sample features, sample sparsity, sample diversity, and similarity in sampling sequences. We choose four types of parallel training algorithms as our research objects: (1) the asynchronous parallel SGD algorithm (Hogwild! algorithm), (2) the parallel model average SGD algorithm (minibatch SGD algorithm), (3) the decentralization optimization algorithm, and (4) the dual coordinate optimization (DADM algorithm). Our results show that the statistical properties of training datasets determine the scalability upper bound of these parallel training algorithms. Daning Cheng, Shigang Li 0002, Hanping Zhang, Fen Xia, Yunquan Zhang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | Breaking (Global) Barriers in Parallel Stochastic Optimization With Wait-Avoiding Group AveragingabstractDeep learning at scale is dominated by communication time. Distributing samples across nodes usually yields the best performance, but poses scaling challenges due to global information dissemination and load imbalance across uneven sample lengths. State-of-the-art decentralized optimizers mitigate the problem, but require more iterations to achieve the same accuracy as their globally-communicating counterparts. We present Wait-Avoiding Group Model Averaging (WAGMA) SGD, a wait-avoiding stochastic optimizer that reduces global communication via subgroup weight exchange. The key insight is a combination of algorithmic changes to the averaging scheme and the use of a group allreduce operation. We prove the convergence of WAGMA-SGD, and empirically show that it retains convergence rates similar to Allreduce-SGD. For evaluation, we train ResNet-50 on ImageNet; Transformer for machine translation; and deep reinforcement learning for navigation at scale. Compared with state-of-the-art decentralized SGD variants, WAGMA-SGD significantly improves training throughput (e.g., 2.1× on 1,024 GPUs for reinforcement learning), and achieves the fastest time-to-solution (e.g., the highest score using the shortest training time for Transformer). Shigang Li 0002, Tal Ben-Nun, Giorgi Nadiradze, Salvatore Di Girolamo, Nikoli Dryden, Dan Alistarh, Torsten Hoefler |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | A Highly Efficient Dynamical Core of Atmospheric General Circulation Model based on Leap-FormatabstractThe finite-difference dynamical core based on the equal-interval latitude-longitude mesh has been widely used for numerical simulations of the Atmospheric General Circulation Model (AGCM). Previous work utilizes different filtering schemes to alleviate the instability problem incurred by the unequal physical spacing at different latitudes, but they all incur high communication and computation overhead and become a scaling bottleneck. This paper proposes a new leap-format finite-difference computing scheme. It generalizes the usual finite-difference format with adaptive wider intervals and is able to maintain the computational stability in the grid updating. Therefore, the costly filtering scheme is eliminated. The new scheme is parallelized with a shifting communication method and implemented with fine communication optimizations based on a 3D decomposition. With the proposed leap-format computation scheme, the communication overhead of the AGCM is significantly reduced and good load balance is exhibited. The simulation results verify the correctness of the new leap-format scheme. The new scheme achieves the speed of 16.6 simulation-year-per-day (SYPD) and up to 3.3x speedup over the latest implementation. He Zhang 0005, Baodong Wu, Shigang Li 0002, Pengqi Lu, Yunquan Zhang, Yongjun Xu 0001 |
IPDPS | 5 |
| 2020 | Taming unbalanced training workloads in deep learning with partial collective operationsabstractLoad imbalance pervasively exists in distributed deep learning training systems, either caused by the inherent imbalance in learned tasks or by the system itself. Traditional synchronous Stochastic Gradient Descent (SGD) achieves good accuracy for a wide variety of tasks, but relies on global synchronization to accumulate the gradients at every training step. In this paper, we propose eager-SGD, which relaxes the global synchronization for decentralized accumulation. To implement eager-SGD, we propose to use two partial collectives: solo and majority. With solo allreduce, the faster processes contribute their gradients eagerly without waiting for the slower processes, whereas with majority allreduce, at least half of the participants must contribute gradients before continuing, all without using a central parameter server. We theoretically prove the convergence of the algorithms and describe the partial collectives in detail. Experiments are conducted on a variety of neural networks and datasets. The results on load-imbalanced environments show that eager-SGD achieves 2.64 X speedup (ResNet-50 on ImageNet) over the asynchronous centralized SGD, and achieves 1.29 X speedup (ResNet-50 on ImageNet) and 1.27X speedup (LSTM on UCF101) over the state-of-the-art synchronous decentralized SGDs, without losing accuracy. Shigang Li 0002, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh, Torsten Hoefler |
PPoPP | 1 |
| 2020 | WP-SGD: Weighted parallel SGD for distributed unbalanced-workload training system
Daning Cheng, Shigang Li 0002, Yunquan Zhang |
J. Parallel Distributed Comput. | 2 |
| 2020 | FastNBL: fast neighbor lists establishment for molecular dynamics simulation based on bitwise operations
Kun Li 0016, Shigang Li 0002, Yunquan Zhang |
J. Supercomput. | 2 |
| 2019 | Using Gradient Based Multikernel Gaussian Process and Meta-Acquisition Function to Accelerate SMBOabstractAutomatic machine learning (automl) is a crucial technology in machine learning. Sequential model-based optimisation algorithms (SMBO) (e.g., SMAC, TPE) are state-of-the-art hyperparameter optimisation methods in automl. However, SMBO does not consider known information, like the best hyperparameters high possibility range and gradients. In this paper, we accelerate the traditional SMBO method and name our method as accSMBO. In accSMBO, we build a gradient-based multikernel Gaussian process with a good generalisation ability and we design meta-acquisition function which encourages that SMBO puts more attention on the best hyperparameters high possibility range. In L2 norm regularised logistic loss function experiments, our method exhibited state-of-the-art performance. Daning Cheng, Hanping Zhang, Fen Xia, Shigang Li 0002, Yunquan Zhang |
ICTAI | 4 |
| 2019 | OpenKMC: a KMC design for hundred-billion-atom simulation using millions of cores on Sunway TaihulightabstractWith more attention attached to nuclear energy, the formation mechanism of the solute clusters precipitation within complex alloys becomes intriguing research in the embrittlement of nuclear reactor pressure vessel (RPV) steels. Such phenomenon can be simulated with atomic kinetic Monte Carlo (AKMC) software, which evaluates the interactions of solute atoms with point defects in metal alloys. In this paper, we propose OpenKMC to accelerate large-scale KMC simulations on Sunway many-core architecture. To overcome the constraints caused by complex many-core architecture, we employ six levels of optimization in OpenKMC: (1) a new efficient potential computation model; (2) a group reaction strategy for fast event selection; (3) a software cache strategy; (4) combined communication optimizations; (5) a Transcription-Translation-Transmission algorithm for many-core optimization; (6) vectorization acceleration. Experiments illustrate that our OpenKMC has high accuracy and good scalability of applying hundred-billion-atom simulation over 5.2 million cores with a performance of over 80.1% parallel efficiency. Kun Li 0016, Honghui Shang, Yunquan Zhang, Shigang Li 0002, Baodong Wu, Dexun Chen, Zhiqiang Wei 0004 |
SC | 4 |
| 2019 | Efficient parallel optimizations of a high-performance SIFT on GPUs
Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Shice Liu, Shigang Li 0002 |
J. Parallel Distributed Comput. | 5 |
| 2019 | Correction to: FastNBL: fast neighbor lists establishment for molecular dynamics simulation based on bitwise operations
Kun Li 0016, Shigang Li 0002, Yunquan Zhang |
J. Supercomput. | 2 |
| 2018 | AGCM3D: A Highly Scalable Finite-Difference Dynamical Core of Atmospheric General Circulation Model Based on 3D DecompositionabstractIt is commonly recognized that the dynamical core of the atmospheric model based on latitude-longitude mesh has poor parallel scalability, since it has to perform the costly polar or high-latitude filtering to dump out the unwanted modes. To parallelize the algorithm, only two dimensions can be partitioned even for a 3-dimensional mesh because of the costly filtering, which hinders the scalability of the algorithm. In this paper, we develop a highly scalable finite-difference dynamical core based on the latitude-longitude mesh using a 3D decomposition method, named as AGCM3D. Different from the traditional methods, our method releases the parallelism in all three dimensions, namely latitude, longitude, and level. To replace the costly Fast Fourier Transform (FFT) filtering, we propose a novel adaptive Gaussian filtering scheme, whose filtering strength increases as the latitude increases. Compared with the parallel FFT filtering, the parallel adaptive Gaussian filtering is far more efficient. In addition, we use the techniques of communication avoiding and message aggregation to further reduce the communication overhead. Experiments are conducted on Tianhe-2 supercomputer, and the resolution of the model is set as 0.5°x0.5°(50 km). Results show that our implementation scales up to 32,768 CPU cores in strong scaling and achieves the maximal simulation speed of 15.6 simulation-year-per-day (SYPD). Baodong Wu, Shigang Li 0002, Yunquan Zhang, He Zhang 0005, Junmin Xiao |
ICPADS | 2 |
| 2018 | Massively Scaling the Metal Microscopic Damage Simulation on Sunway TaihuLight SupercomputerabstractThe limitation of simulation scales leads to a gap between simulation results and physical phenomena. This paper reports our efforts on increasing the scalability of metal material microscopic damage simulation on the Sunway TaihuLight supercomputer. We use a multiscale modeling approach that couples Molecular Dynamics (MD) with Kinetic Monte Carlo (KMC). According to the characteristics of metal materials, we design a dedicated data structure to record the neighbor atoms for MD, which significantly reduces the memory consumption. Data compaction and double buffer are used to reduce the data transfer overhead between the main memory and the local store. We propose an on-demand communication strategy for KMC to remarkably reduce the communication overhead. We simulate 4 * 1012 atoms on 6,656,000 master+slave cores using MD with 85% parallel efficiency. Using the coupled MD-KMC approach, we simulate 3.2 * 1010 atoms in 19.2 days temporal scale on 6,240,000 master+slave cores with runtime of 8.6 hours. Shigang Li 0002, Baodong Wu, Yunquan Zhang, Xianmeng Wang, Jianjiang Li, Changjun Hu, Jue Wang 0013, Yangde Feng, Ningming Nie |
ICPP | 1 |
| 2018 | Communication-Avoiding for Dynamical Core of Atmospheric General Circulation ModelabstractDynamical core is one of the most time-consuming parts in the global atmospheric general circulation model, which is widely used for the numerical simulation of the dynamic evolution process of global atmosphere. Due to its complicated calculation procedures and the non-uniformity of latitude-longitude mesh, the parallelization suffers from high communication overhead. In this paper, we deduce the operator form of the calculating flow in the dynamical core. Furthermore, it is abstracted out that the stencil and collection alternate action is the basic operation in the dynamic core. Based on the operator form of the calculation flow, we propose the corresponding optimization strategy for each operator. In the end, we develop a communication-avoiding algorithm to reduce communication overhead in the dynamic core. Our experiments show that the communication-avoiding algorithm reduces the total runtime by 54% at most for a 50 km resolution model running 10 years. Especially for communication reduction, the new algorithm achieves 1.4x speedup on average for the collective communication and 3.9x speedup on average for the communication involved in the stencil computation. Junmin Xiao, Shigang Li 0002, Baodong Wu, He Zhang 0005, Kun Li 0016, Erlin Yao, Yunquan Zhang, Guangming Tan |
ICPP | 2 |
| 2018 | Cache-Oblivious MPI All-to-All Communications Based on Morton OrderabstractMany-core systems with a rapidly increasing number of cores pose a significant challenge to parallel applications to use their complex memory hierarchies efficiently. Many such applications rely on collective communications in performance-critical phases, which become a bottleneck if they are not optimized. We address this issue by proposing cache-oblivious algorithms for MPI_Alltoall, MPI_Allgather, and the MPI neighborhood collectives to exploit the data locality. To implement the cache-oblivious algorithms, we allocate the send and receive buffers on a shared heap and use Morton order to guide the memory copies. Our analysis shows that our algorithm for MPI_Alltoall is asymptotically optimal. We show an extension to our algorithms to minimize the communication distance on NUMA systems while maintaining optimality within each socket. We further demonstrate how the cache-oblivious algorithms can be applied to multi-node machines. Experiments are conducted on different many-core architectures. For MPI_Alltoall, our implementation achieves on average 1.40X speedup over the naive implementation based on shared heap for small and medium block sizes (less than 16 KB) on a Xeon Phi KNC, achieves on average 3.03X speedup over MVAPICH2 on a Xeon E7-8890, and achieves on average 2.23X speedup over MVAPICH2 on a 256-node Xeon E5-2680 cluster for block sizes less than 1 KB. Shigang Li 0002, Yunquan Zhang, Torsten Hoefler |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | POSTER: Cache-Oblivious MPI All-to-All Communications on Many-Core ArchitecturesabstractIn the many-core era, the performance of MPI collectives is more dependent on the intra-node communication component. However, the communication algorithms generally inherit from the inter-node version and ignore the cache complexity. We propose cache-oblivious algorithms for MPI all-to-all operations, in which data blocks are copied into the receive buffers in Morton order to exploit data locality. Experimental results on different many-core architectures show that our cache-oblivious implementations significantly outperform the naive implementations based on shared heap and the highly optimized MPI libraries. Shigang Li 0002, Yunquan Zhang, Torsten Hoefler |
PPoPP | 1 |
| 2016 | Parallel Processing Systems for Big Data: A SurveyabstractThe volume, variety, and velocity properties of big data and the valuable information it contains have motivated the investigation of many new parallel data processing systems in addition to the approaches using traditional database management systems (DBMSs). MapReduce pioneered this paradigm change and rapidly became the primary big data processing system for its simplicity, scalability, and fine-grain fault tolerance. However, compared with DBMSs, MapReduce also arouses controversy in processing efficiency, low-level abstraction, and rigid dataflow. Inspired by MapReduce, nowadays the big data systems are blooming. Some of them follow MapReduce's idea, but with more flexible models for general-purpose usage. Some absorb the advantages of DBMSs with higher abstraction. There are also specific systems for certain applications, such as machine learning and stream data processing. To explore new research opportunities and assist users in selecting suitable processing systems for specific applications, this survey paper will give a high-level overview of the existing parallel data processing systems categorized by the data input as batch processing, stream processing, graph processing, and machine learning processing and introduce representative projects in each category. As the pioneer, the original MapReduce system, as well as its active variants and extensions on dataflow, data access, parameter tuning, communication, and energy optimizations will be discussed at first. System benchmarks and open issues for big data processing will also be studied in this survey. Yunquan Zhang, Shigang Li 0002, Xinhui Tian, Haipeng Jia, Athanasios V. Vasilakos |
Proc. IEEE | 3 |
| 2016 | A Cross-Platform SpMV Framework on Many-Core ArchitecturesabstractSparse Matrix-Vector multiplication (SpMV) is a key operation in engineering and scientific computing. Although the previous work has shown impressive progress in optimizing SpMV on many-core architectures, load imbalance and high memory bandwidth remain the critical performance bottlenecks. We present our novel solutions to these problems, for both GPUs and Intel MIC many-core architectures. First, we devise a new SpMV format, called Blocked Compressed Common Coordinate (BCCOO). BCCOO extends the blocked Common Coordinate (COO) by using bit flags to store the row indices to alleviate the bandwidth problem. We further improve this format by partitioning the matrix into vertical slices for better data locality. Then, to address the load imbalance problem, we propose a highly efficient matrix-based segmented sum/scan algorithm for SpMV, which eliminates global synchronization. At last, we introduce an autotuning framework to choose optimization parameters. Experimental results show that our proposed framework has a significant advantage over the existing SpMV libraries. In single precision, our proposed scheme outperforms clSpMV COCKTAIL format by 255% on average on AMD FirePro W8000, and outperforms CUSPARSE V7.0 by 73.7% on average and outperforms CSR5 by 53.6% on average on GeForce Titan X; in double precision, our proposed scheme outperforms CUSPARSE V7.0 by 34.0% on average and outperforms CSR5 by 16.2% on average on Tesla K20, and has equivalent performance compared with CSR5 on Intel MIC. Yunquan Zhang, Shigang Li 0002, Shengen Yan, Huiyang Zhou |
ACM Trans. Archit. Code Optim. | 2 |
| 2015 | Analyzing MPI-3.0 Process-Level Shared Memory: A Case Study with Stencil ComputationsabstractThe recently released MPI-3.0 standard introduced a process-level shared-memory interface which enables processes within the same node to have direct load/store access to each others' memory. Such an interface allows applications to declare data structures that are shared by multiple MPI processes on the node. In this paper, we study the capabilities and performance implications of using MPI-3.0 shared memory, in the context of a five-point stencil computation. Our analysis reveals that the use of MPI-3.0 shared memory has several unforeseen performance implications including disrupting certain compiler optimizations and incorrectly using suboptimal page sizes inside the OS. Based on this analysis, we propose several methodologies for working around these issues and improving communication performance by 40-85% compared to the current MPI-1.0 based approach. Junchao Zhang 0002, Kazutomo Yoshii, Shigang Li 0002, Yunquan Zhang, Pavan Balaji |
CCGRID | 4 |
| 2015 | Automatic tuning of sparse matrix-vector multiplication on multicore clusters
Shigang Li 0002, Changjun Hu, Junchao Zhang 0002, Yunquan Zhang |
Sci. China Inf. Sci. | 1 |
| 2013 | NUMA-aware shared-memory collective communication for MPI
Shigang Li 0002, Torsten Hoefler, Marc Snir |
HPDC | 1 |
| 2013 | Asynchronous Work Stealing on Distributed Memory SystemsabstractWork stealing is a popular policy for dynamic load balancing of irregular applications. However, communication overhead incurred by work stealing may make it less efficient, especially on distributed memory systems. In this work we propose an asynchronous work stealing (AsynchWS) strategy which exploits opportunities to overlap communication with local residual tasks. Profiling information is collected locally to optimize task granularity and guide the asynchronous work stealing. AsynchWS is implemented in Unified Parallel C (UPC), which effectively supports non-blocking one-sided communication and facilitates the implementation. Experiments are conducted on a 32 nodes Xeon X5650 cluster using a set of irregular applications. Results show that up to 16% better performance than the state-of-the-art strategies on distributed memory. Shigang Li 0002, Chongchong Zhao |
PDP | 1 |
| 2011 | Extending Synchronization Constructs in OpenMP to Exploit Pipeline Parallelism on Heterogeneous Multi-core
Shigang Li 0002, Shucai Yao, Haohu He, Lili Sun, Yi Chen 0003 |
ICA3PP (2) | 1 |
| 2010 | Support for OpenMP Tasks on Cell Architecture
Changjun Hu, Haohu He, Shigang Li 0002 |
ICA3PP (2) | 5 |