EDBT 2026 Demo / reviewers in the wild / expert
Jianlong Zhong
dblp:90/9635
· DBLP profile ↗
9ranked-venue papers
5as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-authorArtificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
GPUs and heterogeneous computing · 55% Parallel and multicore computing · 24% High-performance computing · 11% | |
| Computer networks
1 paper |
Network performance modeling · 100% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU runtime systems |
0.4 | 2 | 2014 | Medusa: Simplified Graph Processing on GPUs · IEEE Trans. Parallel Distributed Syst. 2014 Kernelet: High-Throughput GPU Kernel Executions with Dynamic Slicing and Scheduling · IEEE Trans. Parallel Distributed Syst. 2014 |
High-performance computing
collective communication |
0.4 | 2 | 2015 | Network Performance Aware MPI Collective Communication Operations in the Cloud · IEEE Trans. Parallel Distributed Syst. 2015 An overview of CMPI: network performance aware MPI in the cloud · PPoPP 2012 |
GPUs and heterogeneous computing
GPU graph processing |
0.4 | 2 | 2015 | Optimization of asynchronous graph processing on GPU with hybrid coloring model · PPoPP 2015 An overview of Medusa: simplified graph processing on GPUs · PPoPP 2012 |
GPUs and heterogeneous computing › GPU programming
GPU programming frameworks |
0.4 | 2 | 2014 | Medusa: Simplified Graph Processing on GPUs · IEEE Trans. Parallel Distributed Syst. 2014 Parallel Graph Processing on Graphics Processors Made Easy · Proc. VLDB Endow. 2013 |
Parallel and multicore computing
parallel programming models |
0.3 | 2 | 2013 | Parallel Graph Processing on Graphics Processors Made Easy · Proc. VLDB Endow. 2013 An overview of Medusa: simplified graph processing on GPUs · PPoPP 2012 |
GPUs and heterogeneous computing › GPU graph processing
asynchronous graph processing |
0.2 | 1 | 2015 | Optimization of asynchronous graph processing on GPU with hybrid coloring model · PPoPP 2015 |
Parallel and multicore computing
graph coloring |
0.2 | 1 | 2015 | Optimization of asynchronous graph processing on GPU with hybrid coloring model · PPoPP 2015 |
GPUs and heterogeneous computing › GPU kernel
GPU kernel execution |
0.2 | 1 | 2014 | Kernelet: High-Throughput GPU Kernel Executions with Dynamic Slicing and Scheduling · IEEE Trans. Parallel Distributed Syst. 2014 |
GPUs and heterogeneous computing
multi-GPU computing |
0.2 | 1 | 2014 | Medusa: Simplified Graph Processing on GPUs · IEEE Trans. Parallel Distributed Syst. 2014 |
GPUs and heterogeneous computing
GPU programming |
0.2 | 1 | 2013 | Parallel Graph Processing on Graphics Processors Made Easy · Proc. VLDB Endow. 2013 |
Parallel and multicore computing
parallel graph algorithms |
0.2 | 1 | 2013 | Parallel Graph Processing on Graphics Processors Made Easy · Proc. VLDB Endow. 2013 |
Distributed systems
graph processing systems |
0.1 | 1 | 2012 | An overview of Medusa: simplified graph processing on GPUs · PPoPP 2012 |
Parallel and multicore computing
MPI |
0.1 | 1 | 2012 | An overview of CMPI: network performance aware MPI in the cloud · PPoPP 2012 |
Graph algorithms and graph theory › graph traversal
breadth-first search |
0.1 | 1 | 2014 | Medusa: Simplified Graph Processing on GPUs · IEEE Trans. Parallel Distributed Syst. 2014 |
Graph algorithms and graph theory
graph algorithms |
0.1 | 1 | 2014 | Medusa: Simplified Graph Processing on GPUs · IEEE Trans. Parallel Distributed Syst. 2014 |
Methods — techniques the papers use, named apart from their topics
graph-centric optimization · 0.7network performance modeling · 0.3simulation · 0.2hybrid coloring · 0.2benchmarking · 0.2scheduling · 0.2markov chain · 0.2dynamic slicing · 0.2runtime scheduling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Optimization of asynchronous graph processing on GPU with hybrid coloring modelabstractModern GPUs have been widely used to accelerate the graph processing for complicated computational problems regarding graph theory. Many parallel graph algorithms adopt the asynchronous computing model to accelerate the iterative convergence. Unfortunately, the consistent asynchronous computing requires locking or the atomic operations, leading to significant penalties/overheads when implemented on GPUs. To this end, coloring algorithm is adopted to separate the vertices with potential updating conflicts, guaranteeing the consistency/correctness of the parallel processing. We propose a light-weight asynchronous processing framework called Frog with a hybrid coloring model. We find that majority of vertices (about 80%) are colored with only a few colors, such that they can be read and updated in a very high degree of parallelism without violating the sequential consistency. Accordingly, our solution will separate the processing of the vertices based on the distribution of colors. Xuanhua Shi, Junling Liang, Sheng Di, Bingsheng He, Hai Jin 0001, Lu Lu 0006, Jianlong Zhong |
PPoPP | 9 |
| 2015 | Network Performance Aware MPI Collective Communication Operations in the CloudabstractThis paper examines the performance of collective communication operations in message passing interfaces (MPI) in the cloud computing environment. The awareness of network topology has been a key factor in performance optimizations for existing MPI implementations. However, virtualization in the cloud environment not only hides the network topology information from the users, but also causes traffic interference and dynamics to network performance. Existing topology-aware optimizations are no longer feasible in the cloud environment. Therefore, we develop novel network performance aware algorithms for a series of collective communication operations including broadcast, reduce, gather and scatter. We further implement two common applications, N-body and conjugate gradient (CG). We have conducted our experiments with two complementary methods (on Amazon EC2 and simulations). Our experimental results show that the network performance awareness results in 25.4 and 28.3 percent performance improvement over MPICH2 on Amazon EC2 and on simulations, respectively. Evaluations on N-body and CG show 41.6 and 14.3 percent respectively on application performance improvement. Yifan Gong 0003, Bingsheng He, Jianlong Zhong |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | Kernelet: High-Throughput GPU Kernel Executions with Dynamic Slicing and SchedulingabstractGraphics processors, or GPUs, have recently been widely used as accelerators in shared environments such as clusters and clouds. In such shared environments, many kernels are submitted to GPUs from different users, and throughput is an important metric for performance and total ownership cost. Despite recently improved runtime support for concurrent GPU kernel executions, the GPU can be severely underutilized, resulting in suboptimal throughput. In this paper, we propose Kernelet, a runtime system to improve the throughput of concurrent kernel executions on the GPU. Kernelet embraces transparent memory management and PCI-e data transfer techniques, and dynamic slicing and scheduling techniques for kernel executions. With slicing, Kernelet divides a GPU kernel into multiple sub-kernels (namely slices ). Each slice has tunable occupancy to allow co-scheduling with other slices for high GPU utilization. We develop a novel Markov chain-based performance model to guide the scheduling decision. Our experimental results demonstrate up to 31 percent and 23 percent performance improvement on NVIDIA Tesla C2050 and GTX680 GPUs, respectively. Jianlong Zhong, Bingsheng He |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Medusa: Simplified Graph Processing on GPUsabstractGraphs are common data structures for many applications, and efficient graph processing is a must for application performance. Recently, the graphics processing unit (GPU) has been adopted to accelerate various graph processing algorithms such as BFS and shortest paths. However, it is difficult to write correct and efficient GPU programs and even more difficult for graph processing due to the irregularities of graph structures. To simplify graph processing on GPUs, we propose a programming framework called Medusa which enables developers to leverage the capabilities of GPUs by writing sequential C/C++ code. Medusa offers a small set of user-defined APIs and embraces a runtime system to automatically execute those APIs in parallel on the GPU. We develop a series of graph-centric optimizations based on the architecture features of GPUs for efficiency. Additionally, Medusa is extended to execute on multiple GPUs within a machine. Our experiments show that 1) Medusa greatly simplifies implementation of GPGPU programs for graph processing, with many fewer lines of source code written by developers and 2) the optimization techniques significantly improve the performance of the runtime system, making its performance comparable with or better than manually tuned GPU graph operations. Jianlong Zhong, Bingsheng He |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Towards GPU-Accelerated Large-Scale Graph Processing in the CloudabstractRecently, we have witnessed that cloud providers start to offer heterogeneous computing environments. There have been wide interests in both clusters and cloud of adopting graphics processors (GPUs) as accelerators for various applications. On the other hand, large-scale graph processing is important for many data-intensive applications in the cloud. In this paper, we propose to leverage GPUs to accelerate large-scale graph processing in the cloud. Specifically, we develop an in-memory graph processing engine G2 with three non-trivial GPU-specific optimizations. Firstly, we adopt fine-grained APIs to take advantage of the massive thread parallelism of the GPU. Secondly, G2 embraces a graph partition based approach for load balancing on heterogeneous CPU/GPU architectures. Thirdly, a runtime system is developed to perform transparent memory management on the GPU, and to perform scheduling for an improved throughput of concurrent kernel executions from graph tasks. We have conducted experiments on an Amazon EC2 virtual cluster of eight nodes. Our preliminary results demonstrate that 1) GPU is a viable accelerator for cloud-based graph processing, and 2) the proposed optimizations improve the performance of GPU-based graph processing engine. We further present the lessons learnt and open problems towards large-scale graph processing with GPU accelerations. Jianlong Zhong, Bingsheng He |
CloudCom (1) | 1 |
| 2013 | Simulation of Information Propagation over Complex Networks: Performance Studies on Multi-GPUabstractGeneral Purpose Graphics Processing Units (GPGPU) have been used in high performance computing platforms to accelerate the performance of scientific applications such as simulations. With the increased computing resources required for large-scale network simulation, one GPU device may not have enough memory and computation capacities. It is therefore necessary to enhance the system scalability by introducing multiple GPU devices. It is also attractive to investigate the performance scalability of Multi-GPU simulations. This paper describes the simulation of information propagation on multiple GPU devices, including the optimized network simulation algorithms, the network partitioning and replication strategy, and the data synchronization scheme. The experimental results for scalable random networks show that the number of simulation steps, computation time, synchronization time, and data transfer time all affect the overall simulation performance. In order to compare with random networks, we also conduct simulations of scale-free networks. We can observe that the node replication ratio in scale-free networks is smaller than that in random networks and therefore the cost of data transfer and synchronization is significantly reduced. This indicates that the network structure is also an important factor that influences the simulation performance in a Multi-GPU system. Jiangming Jin, Stephen John Turner, Bu-Sung Lee, Jianlong Zhong, Bingsheng He |
DS-RT | 4 |
| 2013 | Parallel Graph Processing on Graphics Processors Made EasyabstractThis paper demonstrates Medusa, a programming framework for parallel graph processing on graphics processors (GPUs). Medusa enables developers to leverage the massive parallelism and other hardware features of GPUs by writing sequential C/C++ code for a small set of APIs. This simplifies the implementation of parallel graph processing on the GPU. The runtime system of Medusa automatically executes the user-defined APIs in parallel on the GPU, with a series of graph-centric optimizations based on the architecture features of GPUs. We will demonstrate the steps of developing GPU-based graph processing algorithms with Medusa, and the superior performance of Medusa with both real-world and synthetic datasets. Jianlong Zhong, Bingsheng He |
Proc. VLDB Endow. | 1 |
| 2012 | An overview of CMPI: network performance aware MPI in the cloudabstractCloud computing enables users to perform distributed computing tasks on many virtual machines, without owning a physical cluster. Recently, various distributed computing tasks such as scientific applications are being moved from supercomputers and private clusters to public clouds. Message passing interface (MPI) is a key and common component in distributed computing tasks. The virtualized computing environment of the public cloud hides the network topology information from the users, and existing topology-aware optimizations for MPI are no longer feasible in the cloud environment. We propose a network performance aware MPI library named CMPI. CMPI embraces a new model for capturing the network performance among different virtual machines in the cloud. Based on the network performance model, we develop novel network performance aware algorithms for communication operations. This poster gives an overview of CMPI design, and presents some preliminary results on collective operations such as broadcast.We demonstrate the effectiveness of our network performance aware optimizations on Amazon EC2. Yifan Gong 0003, Bingsheng He, Jianlong Zhong |
PPoPP | 3 |
| 2012 | An overview of Medusa: simplified graph processing on GPUsabstractGraphs are the de facto data structures for many applications, and efficient graph processing is a must for the application performance. GPUs have an order of magnitude higher computational power and memory bandwidth compared to CPUs and have been adopted to accelerate several common graph algorithms. However, it is difficult to write correct and efficient GPU programs and even more difficult for graph processing due to the irregularities of graph structures. To address those difficulties, we propose a programming framework named Medusa to simplify graph processing on GPUs. Medusa offers a small set of APIs, based on which developers can define their application logics by writing sequential code without awareness of GPU architectures. The Medusa runtime system automatically executes the developer defined APIs in parallel on the GPU, with a series of graph-centric optimizations. This poster gives an overview of Medusa, and presents some preliminary results. Jianlong Zhong, Bingsheng He |
PPoPP | 1 |