Xiangwen Liu

dblp:52/6977 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
4since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2023 uGrapher: High-Performance Graph Operator Computation via Unified Abstraction for Graph Neural Networks
abstract
As graph neural networks (GNNs) have achieved great success in many graph learning problems, it is of paramount importance to support their efficient execution. Different graphs and different operators present different patterns during execution. However, there is still a gap in the existing GNN acceleration research to explore adaptive parallelism. We show that existing GNN frameworks rely on handwritten static kernels, which fail to achieve the best performance across different graph operators and input graph structures. In this work, we propose uGrapher, a unified interface that achieves general high performance for different graph operators and datasets. The existing GNN frameworks can easily integrate our design for its simple and unified API. We take a principled approach that decouples a graph operator’s computation and schedule to achieve that. We first build a GNN-specific operator abstraction that incorporates the semantics of graph tensors and graph loops. We explore various schedule strategies based on the abstraction that can balance the well-established trade-off relationship between parallelism, locality, and efficiency. Our evaluation shows that uGrapher can bring up to 29.1× (3.5× on average) performance improvement over the state-of-the-art baselines on two studied NVIDIA GPUs.
Yangjie Zhou 0001, Jingwen Leng, Yaoxu Song, Shuwen Lu, Chao Li 0009, Minyi Guo, Wenting Shen, Yong Li 0045, Wei Lin 0016, Xiangwen Liu
ASPLOS (2)11
2023 High-Throughput GPU Random Walk with Fine-Tuned Concurrent Query Processing
abstract
Random walk serves as a powerful tool in dealing with large-scale graphs, reducing data size while preserving structural information. Unfortunately, existing system frameworks all focus on the execution of a single walker task in serial. We propose CoWalker, a high-throughput GPU random walk framework tailored for concurrent random walk tasks. It introduces a multi-level concurrent execution model to allow concurrent random walk tasks to efficiently share GPU resources with low overhead. Our system prototype confirms that the proposed system could outperform (up to 54%) the state-of-the-art in a wide spectral of scenarios.
Chao Li 0009, Pengyu Wang 0003, Xiaofeng Hou, Jing Wang 0055, Shixuan Sun, Minyi Guo, Dongbai Chen, Xiangwen Liu
PPoPP10
2023 DRAGON: Dynamic Recurrent Accelerator for Graph Online Convolution
abstract
Despite the extraordinary applicative potentiality that dynamic graph inference may entail, its practical-physical implementation has been a topic seldom explored in literature. Although graph inference through neural networks has received plenty of algorithmic innovation, its transfer to the physical world has not found similar development. This is understandable since the most preeminent Euclidean acceleration techniques from CNN have little implication in the non-Euclidean nature of relational graphs. Instead of coping with the challenges arising from forcing naturally sparse structures into more inflexible stochastic arrangements, in DRAGON, we embrace this characteristic in order to promote acceleration. Inspired by high-performance computing approaches like Parallel Multi-moth Flame Optimization for Link Prediction (PMFO-LP), we propose and implement a novel efficient architecture, capable of producing similar speed-up and performance than baseline but at a fraction of its hardware requirements and power consumption. We leverage the hidden parallelistic capacity of our previously developed static graph convolutional processor ACE-GCN and expanded it with RNN structures, allowing the deployment of a multi-processing network referenced around a common pool of proximity-based centroids. Experimental results demonstrate outstanding acceleration. In comparison with the fastest CPU-based software implementation available in the literature, DRAGON has achieved roughly 191× speed-up. Under the largest configuration and dataset, DRAGON was also able to overtake a more power-hungry PMFO-LP by almost 1.59× in speed, and at around 89.59% in power efficiency. More importantly than raw acceleration, we demonstrate the unique functional qualities of our approach as a flexible and fault-tolerant solution that makes it an interesting alternative for an anthology of applicative scenarios.
José Romero Hung, Chao Li 0009, Taolei Wang, Jinyang Guo 0001, Pengyu Wang 0003, Chuanming Shao, Jing Wang 0055, Guoyong Shi, Xiangwen Liu
ACM Trans. Design Autom. Electr. Syst.9
2022 HyFarM: Task Orchestration on Hybrid Far Memory for High Performance Per Bit
abstract
Tapping into secondary memory resources, i.e., far memory (FM), has shown huge potential to improve the cost-efficiency of data centers. Recent advances in both storage-based vertical FM and network-based horizontal FM have raised new questions about leveraging hybrid FM tiers to achieve the best performance per bit of memory. It is still unclear how to efficiently place tasks when far memory access is enabled.In this work, we propose HyFarM, a novel task management strategy for hybrid FM clusters. We analyze FM sensitivity and cooperatively co-locate tasks to enable high utilization and scalability. Further, by tapping into dynamic memory adaption within and across servers, our strategy allows one to consistently deliver high performance on memory-intensive tasks. We evaluate our design with a heavily instrumented testbench. Compared with the state-of-the-art designs, HyFarM respectively improves memory utilization and the overall performance per bit (PPB) by up to 17.6% and 20.5%, with minor overhead.
Jing Wang 0055, Chao Li 0009, Junyi Mei, Taolei Wang, Pengyu Wang 0003, Lu Zhang 0049, Minyi Guo, Dongbai Chen, Xiangwen Liu
ICCD11
2018 Improved Expressivity Through Dendritic Neural Networks
abstract
A typical biological neuron, such as a pyramidal neuron of the neocortex, receives thousands of afferent synaptic inputs on its dendrite tree and sends the efferent axonal output downstream. In typical artificial neural networks, dendrite trees are modeled as linear structures that funnel weighted synaptic inputs to the cell bodies. However, numerous experimental and theoretical studies have shown that dendritic arbors are far more than simple linear accumulators. That is, synaptic inputs can actively modulate their neighboring synaptic activities; therefore, the dendritic structures are highly nonlinear. In this study, we model such local nonlinearity of dendritic trees with our dendritic neural network (DENN) structure and apply this structure to typical machine learning tasks. Equipped with localized nonlinearities, DENNs can attain greater model expressivity than regular neural networks while maintaining efficient network inference. Such strength is evidenced by the increased fitting power when we train DENNs with supervised machine learning tasks. We also empirically show that the locality structure can improve the generalization performance of DENNs, as exemplified by DENNs outranking naive deep neural network architectures when tested on 121 classification tasks from the UCI machine learning repository.
Xundong Wu, Xiangwen Liu
NeurIPS2
2017 Personalized extended (α, k)-anonymity model for privacy-preserving data publishing
abstract
Summary General (α,k)‐anonymity model is a widely used method in privacy‐preserving data publishing, but it cannot provide personalized anonymity. At present, two main schemes for personalized anonymity are the individual‐oriented anonymity and the sensitive value‐oriented anonymity. Unfortunately, the existing personalized anonymity models, designed for any of the aforementioned schemes for privacy‐preserving data publishing, are not effective enough to meet the personalized privacy preservation requirement. In this paper, we propose a novel personalized extended scheme to provide the personalized services in general (α,k)‐anonymity model. The sensitive value‐oriented anonymity is combined with the individual‐oriented anonymity in the new personalized extended (α,k)‐anonymity model by the following two steps: (1) The sensitive attribute values are divided into several groups according to their sensitivities, and each group is assigned with its own frequency constraint threshold. (2) A guarding node is set for each individual to replace his/her sensitive value if necessary. We implement the personalized extended (α,k)‐anonymity model with a clustering algorithm. The performance evaluation finally shows that our model can provide stronger privacy preservation efficiently as well as achieving the personalized service. Copyright © 2016 John Wiley & Sons, Ltd.
Xiangwen Liu, Qing-Qing Xie, Liangmin Wang 0001
Concurr. Comput. Pract. Exp.1
2005 GLB-DMECR: Geographic Location-Based Decentralized Minimum Energy Consumption Routing in Wireless Sensor Networks
abstract
Energy efficient routing is an important measure to promote the energy efficiency of wireless sensor networks (WSN). In this paper, we proposed a minimum energy consumption (MEC) routing algorithm for WSN, GLB-DMECR. This algorithm exploits the utility of sensor nodes’ geographical location information to calculate the ideal MEC route and use the ideal MEC path to guide the routing procedure, thus to find an practical MEC path. Moreover, GLB-DMECR adopts a localized and decentralized routing decision mechanism and has good scalability and stability, which are more suitable for some special characters of WSN such as free of centralized control, nodes’ vulnerability to failure, large network dimension, and etc. The implementation of GLB-DMECR is relatively simple, and it has been verified by simulation that GLB-DMECR outperforms or equivalent to many other present algorithms in MEC performance under wide network circumstance. Therefore, it will have bright future in WSN.
Huifeng Hou, Xiangwen Liu, Hongyi Yu, Hanying Hu
PDCAT2
2005 Coverage and Energy Efficient Information Gathering Protocol in Wireless Sensor Networks
abstract
Coverage and energy efficiency are two important concerns in wireless sensor networks. In this paper, a coverage and energy efficient (CEE) task assigning strategy is presented. The basic idea of CEE is to assign tasks according to node’s coverage degree and remaining energy. Based on this idea, grid-based coverage and energy efficient (GCEE) information gathering protocol is designed. GCEE is simulated with NS-2. How the various parameters affect the fraction of the survived nodes and the area energy balancing of the network is shown in simulation results.
Xiangwen Liu, Huifeng Hou, Jinya Yang, Hongyi Yu, Hanying Hu
PDCAT1