Chenguang Zheng

dblp:48/8490 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Neuromorphic Circuit With Supramodal Attention Effects Based on Cognitive Resource Limitation
abstract
When organisms face complex environments, the cognitive resource limitation is an important mechanism to ensure the quality of perceived information and prevent information overload. Organisms allocate cognitive resources rationally by regulating attention, thus promoting more important cognitive orientations. However, this phenomenon has been scarcely investigated within the realm of memristive biomimetic circuits. The prefrontal cortex (PFC), as the highest central hub for attention control, achieves goal-oriented attentional selection by assigning the basal ganglion to suppress irrelevant information. Based on this biological mechanism, a neuromorphic circuit has been designed in this paper to implement the attentional regulation function of the PFC between bimodal sensory inputs. When organisms face multi-sensory information input, the enhancement and inhibition effects in the supramodal attention effects are considered. In addition, the circuit realizes biological phenomena such as temporal consistency, semantic consistency, emotional attention, and attention fatigue. The attention regulation mechanism is further extended to more senses with variable sensitivities, providing variable strategies for performing tasks in different scenarios. Performance analysis results demonstrate that the circuit exhibits excellent robustness. This work provides guidance for the further development of information processing in brain-inspired intelligence.
Mei Guo, Chenguang Zheng, Gang Dou, Herbert H. C. Iu
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Systems for Scalable Graph Analytics and Machine Learning: Trends and Methods
abstract
Graph-theoretic algorithms and graph machine learning models are essential tools for addressing many real-life problems, such as social network analysis and bioinformatics. To support large-scale graph analytics, graph-parallel systems have been actively developed for over one decade, such as Google's Pregel and Spark's GraphX, which (i) promote a think-like-a-vertex computing model and target (ii) iterative algorithms and (iii) those problems that output a value for each vertex. However, this model is too restricted for supporting the rich set of heterogeneous operations for graph analytics and machine learning that many real applications demand. In recent years, two new trends emerge in graph-parallel systems research: (1) a novel think-like-a-task computing model that can efficiently support the various computationally expensive problems of subgraph search; and (2) scalable systems for learning graph neural networks. These systems effectively complement the diversity needs of graph-parallel tools that can flexibly work together in a comprehensive graph processing pipeline for real applications, with the capability of capturing structural features. This tutorial will provide an effective categorization of the recent systems in these two directions based on their computing models and adopted techniques, and will review the key design ideas of these systems.
Da Yan 0001, Lyuheng Yuan, Akhlaque Ahmad, Chenguang Zheng, James Cheng
KDD4
2024 GE2: A General and Efficient Knowledge Graph Embedding Learning System
abstract
Graph embedding learning computes an embedding vector for each node in a graph and finds many applications in areas such as social networks, e-commerce, and medicine. We observe that existing graph embedding systems (e.g., PBG, DGL-KE, and Marius) have long CPU time and high CPU-GPU communication overhead, especially when using multiple GPUs. Moreover, it is cumbersome to implement negative sampling algorithms on them, which have many variants and are crucial for model quality. We propose a new system called GE 2 , which achieves both generality and efficiency for graph embedding learning. In particular, we propose a general execution model that encompasses various negative sampling algorithms. Based on the execution model, we design a user-friendly API that allows users to easily express negative sampling algorithms. To support efficient training, we offload operations from CPU to GPU to enjoy high parallelism and reduce CPU time. We also design COVER, which, to our knowledge, is the first algorithm to manage data swap between CPU and multiple GPUs for small communication costs. Extensive experimental results show that, comparing with the state-of-the-art graph embedding systems, GE 2 trains consistently faster across different models and datasets, where the speedup is usually over 2x and can be up to 7.5x.
Chenguang Zheng, Guanxian Jiang, Xiao Yan 0002, Peiqi Yin, Qihui Zhou, James Cheng
Proc. ACM Manag. Data1
2023 DSP: Efficient GNN Training with Multiple GPUs
abstract
Jointly utilizing multiple GPUs to train graph neural networks (GNNs) is crucial for handling large graphs and achieving high efficiency. However, we find that existing systems suffer from high communication costs and low GPU utilization due to improper data layout and training procedures. Thus, we propose a system dubbed Distributed Sampling and Pipelining (DSP) for multi-GPU GNN training. DSP adopts a tailored data layout to utilize the fast NVLink connections among the GPUs, which stores the graph topology and popular node features in GPU memory. For efficient graph sampling with multiple GPUs, we introduce a collective sampling primitive (CSP), which pushes the sampling tasks to data to reduce communication. We also design a producer-consumer-based pipeline, which allows tasks from different mini-batches to run congruently to improve GPU utilization. We compare DSP with state-of-the-art GNN training frameworks, and the results show that DSP consistently outperforms the baselines under different datasets, GNN models and GPU counts. The speedup of DSP can be up to 26x and is over 2x in most cases.
Zhenkun Cai, Qihui Zhou, Xiao Yan 0002, Da Zheng 0004, Xiang Song 0003, Chenguang Zheng, James Cheng, George Karypis
PPoPP6
2022 G-Tran: A High Performance Distributed Graph Database with a Decentralized Architecture
abstract
Graph transaction processing poses unique challenges such as random data access due to the irregularity of graph structures, low throughput and high abort rate due to the relatively large read/write sets in graph transactions. To address these challenges, we present G-Tran, a remote direct memory access (RDMA)-enabled distributed in-memory graph database with serializable and snapshot isolation support. First, we propose a graph-native data store to achieve good data locality and fast data access for transactional updates and queries. Second, G-Tran adopts a fully decentralized architecture that leverages RDMA to process distributed transactions with the massively parallel processing (MPP) model, which can achieve high performance by utilizing all computing resources. In addition, we propose a new multi-version optimistic concurrency control (MV-OCC) protocol with two optimizations to address the issue of large read/write sets in graph transactions. Extensive experiments show that G-Tran achieves competitive performance compared with other popular graph databases on benchmark workloads.
Changji Li, Chenguang Zheng, Chenghuan Huang, Juncheng Fang, James Cheng, Jie Zhang 0046
Proc. VLDB Endow.3
2022 ByteGraph: A High-Performance Distributed Graph Database in ByteDance
abstract
Most products at ByteDance, e.g., TikTok, Douyin, and Toutiao, naturally generate massive amounts of graph data. To efficiently store, query and update massive graph data is challenging for the broad range of products at ByteDance with various performance requirements. We categorize graph workloads at ByteDance into three types: online analytical, transaction, and serving processing, where each workload has its own characteristics. Existing graph databases have different performance bottlenecks in handling these workloads and none can efficiently handle the scale of graphs at ByteDance. We developed ByteGraph to process these graph workloads with high throughput, low latency and high scalability. There are several key designs in ByteGraph that make it efficient for processing our workloads, including edge-trees to store adjacency lists for high parallelism and low memory usage, adaptive optimizations on thread pools and indexes, and geographic replications to achieve fault tolerance and availability. ByteGraph has been in production use for several years and its performance has shown to be robust for processing a wide range of graph workloads at ByteDance.
Changji Li, Yingqian Hu, Xiangchen Li, Dongqing Han, Huiming Zhu, Xuwei Fu, Tingwei Wu, Hongfei Tan, Hengtian Ding, Mengjin Liu, Kangcheng Wang, Ting Ye, Chenguang Zheng, James Cheng
Proc. VLDB Endow.23
2022 ByteGNN: Efficient Graph Neural Network Training at Large Scale
abstract
Graph neural networks (GNNs) have shown excellent performance in a wide range of applications such as recommendation, risk control, and drug discovery. With the increase in the volume of graph data, distributed GNN systems become essential to support efficient GNN training. However, existing distributed GNN training systems suffer from various performance issues including high network communication cost, low CPU utilization, and poor end-to-end performance. In this paper, we propose ByteGNN, which addresses the limitations in existing distributed GNN systems with three key designs: (1) an abstraction of mini-batch graph sampling to support high parallelism, (2) a two-level scheduling strategy to improve resource utilization and to reduce the end-to-end GNN training time, and (3) a graph partitioning algorithm tailored for GNN workloads. Our experiments show that ByteGNN outperforms the state-of-the-art distributed GNN systems with up to 3.5--23.8 times faster end-to-end execution, 2--6 times higher CPU utilization, and around half of the network communication cost.
Chenguang Zheng, Yuxuan Cheng, Zhezheng Song, Yifan Wu 0002, Changji Li, James Cheng
Proc. VLDB Endow.1