EDBT 2026 Demo / reviewers in the wild / expert
Wangdong Yang
dblp:84/8385
· DBLP profile ↗
46ranked-venue papers
6as first author
33since 2021 · last 2026
0000-0003-2681-7898ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 3 first-author · 22 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoProteus: LLM-Driven Multi-Version Operator Generation for Energy-Aware Scheduling in Heterogeneous Cloud-Edge Environments
Haotian Wang 0006, Junshuang Ma, Zicong Wang, Wangdong Yang, Kenli Li 0001 |
Euro-Par (2) | 5 |
| 2026 | SPECTRA: Revitalizing image-based iterative method selection for sparse linear systems
Dali Chang, Mingguan Yang, Jing Zhao 0016, Wangdong Yang |
Expert Syst. Appl. | 5 |
| 2026 | Generating Sparsity Patterns for Inverse Preconditioning on SIMD ArchitecturesabstractThe Conjugate Gradient method is a prevalent iterative approach for solving sparse linear systems$Ax = b$, whereAis a symmetric positive definite matrix. In this scenario, the Factorized Sparse Approximate Inverse (FSAI) preconditioner is commonly used. The FSAI approximates$A^{-1}$by the matrix productGTG, whereGis a lower triangular matrix. The numerical properties of FSAI mainly depend on the sparsity pattern ofG. This work presents an approach to generate FSAI sparsity patterns tailored to SIMD architectures, which play a pivotal role in current high-performance computing systems. First, we implement FSAI-S, an FSAI preconditioner based on the SELL-C-σ format for SIMD architectures. To further improve iterative performance, we propose a padding-aware preconditioner, FSAI-P, which leverages the zero padding inherent in the SELL-C-σ to extend the sparsity pattern with minimal computational and storage overhead. We evaluate our approach on three SIMD architectures: a Skylake processor implementing the AVX-512 ISA, a RISC-V processor supporting the RISC-V “V” Vector ISA, and an Nvidia A100 GPU. Experimental results show that FSAI-P provides performance gains across all evaluated platforms. Specifically, average speedups are 12.51% (vs. FSAI) and 11.66% (vs. FSAI-cache) on Skylake, 57.79% (vs. FSAI) on RISC-V, and 18.33% (vs. FSAI) on the A100 GPU. Hantao Xiong, Haotian Wang 0006, Chubo Liu, Wangdong Yang, Kenli Li 0001, Marc Casas |
IEEE Trans. Computers | 4 |
| 2026 | Application of LLM-powered Multimodal Driver Emotion Recognition in IoV SystemabstractIn the Internet of Vehicles (IoV) systems, recognizing driver emotions is crucial to alleviate dangerous driving behaviors caused by emotional instability. Current research predominantly utilizes multimodal data generated by various types of sensors in IoV systems as input to analyze driver emotion changes using multimodal models. However, existing methods are not enough to fully exploit the advantages of large language models (LLM) in information extraction and multimodal feature fusion, which limits the inference capability of emotion recognition models. Therefore, this article proposes an LLM-auxiliary supervision module, which assists in the training phase through LLM to enhance the performance of multimodal emotion recognition models. Specifically, we designed a label text feature extraction (LTFE) module that employs LLM for text data augmentation and extraction, converting label text into semantically informative feature representations. Additionally, we proposed the label-auxiliary supervision (LAS) strategy, which effectively integrates the LLM label text features learned from the LTFE module with the multimodal emotion recognition model during the training phase to enhance the model’s inference ability. Notably, the LTFE and LAS modules are used only during the training phase, ensuring that the backbone model requires minimal computational resources during inference, making it compatible with the computational constraints of intelligent vehicular devices. Extensive experiments conducted on the PPB-Emo, RAVDESS, and IEMOCAP datasets demonstrate that the proposed method outperforms existing approaches in driver emotion recognition tasks. Yiming Wu 0004, Ronghui Cao, Zhuo Tang, Wangdong Yang, Huilong Pi |
ACM Trans. Internet Things | 5 |
| 2026 | Polarity-Aware and Adaptive Sparse Aggregation for Implicit Heterophilic Graph ClassificationabstractGraph-structured data appears in domains such as molecular analysis, social networks, and program optimization, where graphs often exhibit implicit heterogeneity, as nodes may look homogeneous in type yet differ significantly in semantics or functionality. Graph Neural Networks (GNNs), while powerful on homophilic graphs, tend to degrade in such settings due to polarity confusion, over-smoothing, and inefficiency caused by dense propagation. We propose a polarity-aware framework for graph classification that addresses these challenges through adaptive directional sparse aggregation. The framework introduces a polarity-aware propagation mechanism that adaptively reinforces or inverts neighbor signals, mitigating contamination under heterophily. A polarity-guided sparse aggregation operator further alleviates over-smoothing, improves scalability by constraining redundant connections, and condenses information flow into more effective representations, while maintaining unbiased estimation with controlled variance. We provide theoretical analyses that characterize the computational complexity, stability properties, and expressive behavior of signed directional aggregation, offering theoretical insights into its computational, stability, and expressive properties. Extensive experiments on molecular and social graph benchmarks with implicit heterophily demonstrate consistent improvements in graph classification accuracy and efficiency. Our method achieves a 2.36% improvement when compared with the strongest baseline on each dataset. In addition, it improves accuracy by 4.53% on average on program optimization strategy recognition tasks, reaching 80.12% overall. Haotian Wang 0006, Yan Ding 0004, Wangdong Yang, Zhuo Tang, Chubo Liu, Kenli Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2026 | cuFastTuckerPlusTC: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor CoresabstractSparse tensors are prevalent in real-world applications, often characterized by their large-scale, high-order, and high dimensional nature. Directly handling raw tensors is impractical due to the significant memory and computational overhead involved. The current mainstream approach involves compressing or decomposing the original tensor. One popular tensor decomposition algorithm is the Tucker decomposition. However, existing state-of-the-art algorithms for large-scale Tucker decomposition typically relax the original optimization problem into multiple convex optimization problems to ensure polynomial convergence. Unfortunately, these algorithms tend to converge slowly. In contrast, tensor decomposition exhibits a simple optimization landscape, making local search algorithms capable of converging to a global (approximate) optimum much faster. In this paper, we propose the FastTuckerPlus algorithm, which decomposes the original optimization problem into two non-convex optimization problems and solves them alternately using the Stochastic Gradient Descent method. Furthermore, we introduce cuFastTuckerPlusTC, a fine-grained parallel algorithm designed for GPU platforms, leveraging the performance of tensor cores. This algorithm minimizes memory access overhead and computational costs, surpassing the state-of-the-art algorithms. Our experimental results demonstrate that the proposed method achieves a 2× to 8× improvement in convergence speed and a 3× to 5× improvement in per-iteration execution speed compared with state-of-the-art algorithms. Mingxing Duan, Huizhang Luo, Wangdong Yang, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2026 | cuFastTucker-2L: A Two-Level Optimization Parallel Algorithm for Solving FastTucker Decomposition on GPU Platform
Haotian Wang 0006, Wangdong Yang, Keqin Li 0001, Kenli Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | An Input-Aware Sparse Tensor Compiler Empowered by Vectorized AccelerationabstractSparsity is widely prevalent in real-world applications, yet existing compiler optimizations and code generation techniques for sparse computations remain underdeveloped. Sparse matrix-matrix multiplication (SpMM) is a representative operator in sparse computations, whose performance is often limited by the design of sparse formats and the extent of hardware architecture optimization. Most existing solutions achieve highperformance SpMM through two approaches: (1) meticulously designed kernels and specialized sparse formats, which require extensive manual effort, or (2) tensor compilers that support code generation, though these typically offer limited support for sparse patterns, making it challenging to adapt to complex sparsity patterns in practical applications. This paper presents SpMMTC, an input-aware sparse tensor compiler. Given a sparse matrix as input, SpMMTC analyzes its non-zero distribution and generates a vectorized kernel optimized for SpMM on the specific matrix. We evaluated SpMMTC on various workloads. It achieves speedups of 1.21 x to 2.97 x over state-of-the-art methods such as TACO, TVM, and ASpT on different multi-core processors. It also provides a speedup of up to $\mathbf{1. 5 2 x}$ for sparse MobileNetV1 inference on the edge device. Xianhao He, Haotian Wang 0006, Jiapeng Zhang 0001, Wangdong Yang, Anthony T. Chronopoulos, Kenli Li 0001 |
DAC | 4 |
| 2025 | SSpMV: A Sparsity-aware SpMV Framework Empowered by Multimodal Machine LearningabstractSparse Matrix-Vector Multiplication (SpMV) is an essential sparse operation in scientific computing and artificial intelligence. Efficiently adapting SpMV algorithms to diverse matrices and architectures requires a framework capable of accurately recognizing sparse patterns and selecting the optimal implementation. In this work, we introduce Sparsity-aware SpMV (SSpMV), a framework that integrates expert-designed features with multimodal representations to adaptively predict the best-performing algorithm and parameters. For this purpose, we design a multimodal neural network called MM-Adapter, to capture diverse modalities to represent the computational features of SpMV. Experimental results demonstrate that MMAdapter achieves the highest accuracy of $81.05 \%$, outperforming existing SpMV prediction models. Furthermore, SSpMV consistently delivers substantial performance improvements over state-of-the-art sparse libraries across various multi-core platforms. Shengle Lin, Chubo Liu, Yan Ding 0004, Joey Tianyi Zhou, Kenli Li 0001, Wangdong Yang |
DAC | 6 |
| 2025 | Mitigating Channel Redundancy for Multivariate Time Series ForecastingabstractTransformer-based methods have been widely used in multivariate time series forecasting (MTSF), typically following either channel-independent (CI) or channel-dependent (CD) mod-eling approaches. However, CI methods overlook inter-channel information, while CD methods struggle with the complexity and redundancy of channel relationships, leading to suboptimal performance and high computational costs. In this paper, we propose a novel Channel Aggregation Network, dubbed CANet, to efficiently model both intra and inter-channel dependencies. Specifically, CANet embeds input sequences into patch tokens and uses probabilistic masking in the Channel Aggregator to effectively filter out redundant and noisy information. The refined inter-channel features are then injected into the temporal dimension, allowing CANet to focus on temporal attention while efficiently leveraging cross-channel information. After that, the Temporal Sampler selects key temporal tokens along the time axis, enhancing long-range dependency modeling and accelerating global representation learning. Extensive experiments on multiple real-world datasets demonstrate that CANet achieves state-of-the-art performance with superior accuracy and significantly reduced computational complexity. Guoqing Xiao 0001, Boxiang Qin, Yuedan Chen, Wangdong Yang |
IJCNN | 4 |
| 2025 | DIDS: A distributed inference framework with dynamic scheduling capability
Yuwei Yan, Yikun Hu 0001, Qinyun Tsai, Wangdong Yang, Kenli Li 0001 |
Future Gener. Comput. Syst. | 4 |
| 2025 | SASTC: Spatial-Aware Sparse Tensor Completion for Large-Scale Traffic Data Recovery
Renqiu Ouyang, Haotian Wang 0006, Yikun Hu 0001, Wangdong Yang, Kenli Li 0001 |
IEEE Internet Things J. | 5 |
| 2025 | MM-AutoSolver: A multimodal machine learning method for the auto-selection of iterative solvers and preconditioners
Hantao Xiong, Wangdong Yang, Weiqing He, Shengle Lin, Keqin Li 0001, Kenli Li 0001 |
J. Parallel Distributed Comput. | 2 |
| 2025 | A Context-Awareness and Hardware-Friendly Sparse Matrix Multiplication Kernel for CNN Inference AccelerationabstractSparsification technology is crucial for deploying convolutional neural networks in resource-constrained environments. However, the efficiency of sparse models is hampered by irregular memory access patterns in sparse matrix multiplication kernels. Hardware-level support for 2:4 granularity in sparse tensor cores presents an opportunity for designing efficient sparse matrix multiplication kernels. Existing approaches often involve adjusting sparse structures or secondary sparsification, introducing additional computational errors. To tackle this challenge, we introduce a flexible 2:4 structured adaptive sparse matrix multiplication (FS-AMM) method, a hardware-friendly sparse matrix multiplication kernel that leverages model context to accelerate convolutional neural networks. First, we propose a model context-aware matrix pre-processing method that employs heuristic algorithms to estimate a loss of accuracy due to weight sparsity at each layer. Second, we design a hardware-friendly sparse storage format that combines 2:4 sparse and dense storage formats, enabling more versatile sparsity ratio selection. Third, we implement efficient matrix multiplication kernels to optimize GPU utilization. Finally, experimental results on A100 GPUs show that our method effectively utilizes the sparse tensor kernel and obtains an average 3.09 times speedup ratio compared to other sparse methods while maintaining a high accuracy. Haotian Wang 0006, Yan Ding 0004, Weichen Liu 0001, Chubo Liu, Wangdong Yang, Kenli Li 0001 |
IEEE Trans. Computers | 6 |
| 2025 | DCGG: A Dynamically Adaptive and Hardware-Software Coordinated Runtime System for GNN Acceleration on GPUsabstractGraph neural networks (GNNs) are a prominent trend in graph-based deep learning, known for their capacity to produce high-quality node embeddings. However, the existing GNN framework design is only implemented from the algorithm level, and the hardware architecture of the GPU is not fully utilized.To this end, we propose DCGG, a dynamic runtime adaptive framework, which can accelerate various GNN workloads on GPU platforms. DCGG has carried out deeper optimization work mainly in terms of load balancing and software and hardware matching. Accordingly, three optimization strategies are proposed. First, we propose dynamic 2D workload management methods and perform customized optimization based on it, effectively reducing additional memory operations. Second, a new slicing strategy is adopted, combined with hardware features, to effectively improve the efficiency of data reuse. Third, DCGG uses the Quantitative Dimension Parallel Strategy to optimize dimensions and parallel methods, greatly improving load balance and data locality. Extensive experiments demonstrate that DCGG outperforms the state-of-the-art GNN computing frameworks, such as Deep Graph Library (up to 3.10× faster) and GNNAdvisor (up to 2.80× faster), on mainstream GNN architectures across various datasets. Guoqing Xiao 0001, Yuedan Chen, Hongyang Chen 0001, Wangdong Yang |
IEEE Trans. Computers | 5 |
| 2025 | High Performance OpenCL-Based GEMM Kernel Auto-Tuned by Bayesian Optimization
Shengle Lin, Guoqing Xiao 0001, Haotian Wang 0006, Wangdong Yang, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | PUSHGNN: A Low-communication Runtime System for GNN Acceleration on Multi-GPUsabstractThe need for multi-GPU platforms in graph neural networks (GNNs) has been driven by the growing size of input graphs. However, although the existing multi-gpu GNN framework has been optimized from the perspective of optimizing computing and communication operations, communication competition still exists. To this end, we introduce PUSHGNN, a runtime system designed to reduce communication overhead across GPUs, boosting GNN performance. Therefore, we designed a push-based pipeline communication model and made custom tuning to significantly reduce pipeline contention. Comparative assessments demonstrate that PUSHGNN consistently outperforms leading full-graph GNN systems on average 1.97× and 7.57× faster than MGG and MGG-UVM, respectively. Guoqing Xiao 0001, Yuedan Chen, Wangdong Yang |
IEEE Big Data | 4 |
| 2024 | TAPMM: A Traffic-Aware Page Mapping Method for Multi-level NUMA SystemsabstractWith the development of chiplet technology, the architecture of Non-Uniform Memory Access (NUMA) has become increasingly intricate. The placement of memory page significantly influences application performance in NUMA systems. We found that memory access bottlenecks occur between high-level NUMA domains consisting of multiple chiplets. In this paper, we introduce a Traffic-Aware Page Mapping Method (TAPMM) designed for multi-level NUMA systems. TAPMM conceptualizes the multi-level NUMA system as a memory access tree, utilizing hardware performance events to be aware of system traffic and identify the optimal page mapping method for bandwidth efficiency. Our experiments demonstrate that TAPMM achieves a speedup of up to 2.12× on a real commodity machine compared to existing optimization tools. Fengkun Dong, Guoqing Xiao 0001, Haotian Wang 0006, Yikun Hu 0001, Kenli Li 0001, Wangdong Yang |
DAC | 6 |
| 2024 | zeroTT: A Two-Step State Transition Avoidance Scheme for MLC STT-RAMabstractCompared with conventional SRAM, Spin-Transfer Torque Random Access Memory(STT-RAM) is expected to play a crucial role in future memory technologies with the increasing demands for higher storage density and lower power consumption for modern embedded systems. Moreover, Multi-Level Cell (MLC) STT-RAM outperforms Single-Level Cell (SLC) STT-RAM since it has higher bit density. However, MLC STT-RAM suffers from write performance due to the two-step state transitions (TTs) in memory cells' soft domain. State-of-the-art approaches mitigate this issue by reducing TTs with efficient data coding. Unfortunately, none of the existing works can fully eliminate the TTs. In this work, zeroTT, an optimal (3, 4)-based expansion coding method that eliminates TTs for MLC STT-RAM. The design of ZeroTT considers space overhead and coding complexity, and our experimental results demonstrate that zeroTT can completely avoid TTs, leading to a more efficient MLC STT-RAM memory in terms of access latency, energy consumption, and device lifetime. Huizhang Luo, Jeff Zhang 0001, Mingxing Duan, Wangdong Yang, Zhuo Tang, Kenli Li 0001 |
DAC | 5 |
| 2024 | COALA: A Compiler-Assisted Adaptive Library Routines Allocation Framework for Heterogeneous SystemsabstractExperienced developers often leverage well-tuned libraries and allocate their routines for computing tasks to enhance performance when building modern scientific and engineering applications. However, such well-tuned libraries are meticulously customized for specific target architectures or environments. Additionally, the performance of their routines is significantly impacted by the actual input data of computing tasks, which often remains uncertain until runtime. Accordingly, statically allocating these library routines may hinder the adaptability of applications and compromise performance, particularly in the context of heterogeneous systems. To address this issue, we propose the Compiler-Assisted Adaptive Library Routines Allocation (COALA) framework for heterogeneous systems. COALA is a fully automated mechanism that employs compiler assistance for dynamic allocation of the most suitable routine to each computing task on heterogeneous systems. It allows the deployment of varying allocation policies tailored to specific optimization targets. During the application compilation process, COALA reconstructs computing tasks and inserts a probe for each of these tasks. Probes serve the purpose of conveying vital information about the requirements of each task, including its computing objective, data size, and computing flops, to a user-level allocation component at runtime. Subsequently, the allocation component utilizes the probe information along with the allocation policy to assign the most optimal library routine for executing the computing tasks. In our prototype, we further introduce and deploy a performance-oriented allocation policy founded on a machine learning-based performance evaluation method for library routines. Experimental verification and evaluation on two heterogeneous systems reveal that COALA can significantly improve application performance, with gains of up to 4.3x for numerical simulation software and 4.2x for machine learning applications, and enhance system utilization by up to 27.8%. Qinyun Tsai, Guanghua Tan, Wangdong Yang, Xianhao He, Yuwei Yan, Keqin Li 0001, Kenli Li 0001 |
IEEE Trans. Computers | 3 |
| 2024 | Parallel algorithm design and optimization of geodynamic numerical simulation application on the Tianhe new-generation high-performance computer
Wangdong Yang, Ruixuan Qi, Qinyun Tsai, Shengle Lin, Fengkun Dong, Kenli Li 0001, Keqin Li 0001 |
J. Supercomput. | 2 |
| 2024 | BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core AccelerationabstractSparse tensor contraction (SpTC) is an important operator in tensor networks, which tends to generate a large amount of sparse high-dimensional data, placing higher demands on the computational performance and storage bandwidth of the processor. Using GPUs with powerful arithmetic characteristics is a reliable choice for accelerating SpTC, however, the high dimensionality and sparsity of tensor makes GPU-accelerated SpTC operators suffer from the difficulties of low computational intensity and high memory consumption. The recent introduction of Tensor Core Units (TCUs) on GPUs brings even more powerful arithmetic, which exacerbates the memory wall problem. To cope with the challenges, this paper proposes a new BCB format that linearizes the indices of multidimensional blocks to reduce block index accesses and uses a bitmap to store the distribution of non-zero elements in a block to reduce the storage overhead. A parallel blocking algorithm of BCB-SpTC is designed to divide the binary linear indices into free and contracted indexes to improve the pairing overhead of computational tasks. Then based on the characteristic computation method of TCUs, the proprietary filling method of TCUs is designed to overcome the inefficiency of parallel computation of sparse data on TCUs. Finally, experimental results on the A100 dataset show that BCB-SpTC improves the acceleration ratio by$1.1\times$to$21.3\times$over the existing SpTC GPU method. Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Keqin Li 0001, Kenli Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | Parallel Algorithm Design and Optimization for Numerical Simulation Application of Ion Implantation in SiliconabstractWe design a computational framework for ion implantation into silicon that can simulate the reaction process of cascade collisions in silicon and the subsequent annealing process. Since simulating materials requires large-scale particle data and high-precision simulation requirements, which often take up a lot of computing time, the conventional computing process needs to be accelerated in parallel to solve current problems. However, the current field of ion implantation focuses on the development of algorithms, often neglecting the design of parallel algorithms, and has not tested them on large-scale supercomputing clusters. Based on Tianhe new generation high-performance computers, this paper proposes a parallel computing framework for numerical simulation of ion implanted silicon based on multi-core processor architecture. By optimizing the data structure, the local continuity of the data is enhanced to improve the overall communication efficiency. Divide three-dimensional data for processes and adapt to MPI parallel computing strategies. Finally, a multi-level parallel computing strategy based on OpenMP is added to improve computing efficiency. After being deployed on the Tianhe new generation high-performance computer platform, experimental results show that on a single node, the multi-level parallel computing framework can increase the computing speed by 20.81X. After conducting tests at different scales, we can see that the parallel efficiency can be maintained at a high level. Ruixuan Qi, Wangdong Yang, Xiuwen Yan, Honglu Li, Kenli Li 0001 |
ICPADS | 2 |
| 2023 | IAP-SpTV: An input-aware adaptive pipeline SpTV via GCN on CPU-GPU
Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Kenli Li 0001, Keqin Li 0001 |
J. Parallel Distributed Comput. | 2 |
| 2023 | A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-AccelerationabstractAnalysis of multi-dimensional data, especially tensor decomposition, which extracts latent information, is becoming considerably popular. Although multi-dimensional sparse data is typically processed on multi-core processors, developing highly optimized GPU-basedSparseTensorMatrixChainMultiplication (SpTMCM) is challenging. The purpose of this paper is to investigate a novel approach named SpTMCM and to explore the discovery of SpTMCM coupled with the emerging computing core, Tensor Core Unit (TCU). In contrast to prior work, the proposed novel approach enables a uniform storage format and optimization approach for SpTMCM. We design a hybrid tensor format based on multi-dimensional tiling that divides the tensor depending on the tile threshold to address the inefficient memory accesses caused by the irregular nonzero distribution of the sparse tensor. Further, we develop a TCU-based tensor parallel algorithm with our novel approach to increase the memory bandwidth. Compared to state-of-the-art works, our method achieves$1.16\sim 24.12\times$speedup for SpMTTKRP and$5.07\sim 7.15\times$speedup for SpTTMChain across NVIDIA A100 GPU on a range of real-world sparse tensors. Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Distributed Set Label-Constrained Reachability Queries over Billion-Scale GraphsabstractSet label-constrained reachability (SLCR) query in edge-labeled graphs is a building block of many graph-based applications. Formally, given two sets$S$and$T$of source and target vertices and a label set (, it returns all reachable vertex pairs (s, t) under the constraint of (, where$s$∊$S$and$t$∊T. There have been abundant index-based approaches to be applied to process the SLCR query. However, distributed approaches are desirable to process large-scale graphs because of the advantages of good scalability and real-time response. Now, there is no efficient distributed approach to the SLCR query. Most index-based approaches face limitations in terms of index construction and query performance when being extended to the distributed environment for processing large-scale graphs. To alleviate these problems, we first build a boundary graph-based index (BoundG) to reduce the time overhead of index construction. Consider the query performance of the BoundG-based approach has no noticeable improvement. We further construct a novel two layers 2-hop index (TL2hop), and a TL2hop-based query algorithm (TLQA) is designed by integrating an early termination strat-egy that reduces the communication overhead and boosts the query performance. Experimental results over eight data graphs demonstrate that the index time of BoundG is comparable to that of the state-of-the-art, and TL2hop significantly outperforms the state-of-the-art technique in terms of query response time (up to 4 orders of magnitude speedup). Wangdong Yang, Xu Zhou 0001, Guoqing Xiao 0001, Yunjun Gao, Kenli Li 0001 |
ICDE | 2 |
| 2022 | LAP: Latency-aware automated pruning with dynamic-based filter selection
Zailong Chen, Chubo Liu, Wangdong Yang, Kenli Li 0001, Keqin Li 0001 |
Neural Networks | 3 |
| 2022 | An Efficient Parallel Reinforcement Learning Approach to Cross-Layer Defense Mechanism in Industrial Control SystemsabstractThe ongoing digitalization enables stable control processes and smooth operations of Industrial Control Systems (ICSs). A direct consequence of the highly interconnected architecture of ICSs is the introduced cyber vulnerability and increasing cyber security threats to ICSs. Numerous researches pay attention to the security problem of ICSs. However, most current researches face two challenges. First, the interaction problem between cyber layer and physical layer of ICSs may result incorrect attack response strategies. Second, ICSs are real-time systems, but existing defense decision algorithms based on game theory or reinforcement learning techniques have high computational complexity, which prevents it from making decisions quickly. In this paper, we design a new multi-attribute based reward quantitative method and propose a multi-attribute based Q-learning algorithm to resolve the interaction problem. In addition, to overcome the limitation of slow convergence, we develop an effective parallel Q-learning (PQL) algorithm to quickly find the optimal strategy. The experimental results show the effectiveness of the PQL algorithm. Compared with the Q-learning algorithm (QL) and the deep Q-network (DQN) algorithm, our proposed solution can reduce the average completion time by 12.5%-37%. Kai Zhong 0004, Zhibang Yang, Guoqing Xiao 0001, Xingpei Li, Wangdong Yang, Kenli Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | STM-multifrontal QR: streaming task mapping multifrontal QR factorization empowered by GCNabstractMultifrontal QR algorithm, which consists of symbolic analysis and numerical factorization, is a high-performance algorithm for orthogonal factorizing sparse matrix. In this work, a graph convolutional network (GCN) for adaptively selecting the optimal reordering algorithm is proposed in symbolic analysis. Using our GCN adaptive classifier, the average numerical factorization time is reduced by 20.78% compared with the default approach, and the additional memory overhead is approximately 4% higher than that of prior work. Moreover, for numerical factorization, an optimized tasks stream parallel processing strategy is proposed and a more efficient computing task mapping framework for NUMA architecture is adopted in this paper, which called STM-Multifrontal QR factorization. Numerical experiments on the TaiShan Server show average 1.22x performance gains over the original SuiteSparseQR. Nearly 80% of datasets have achieved better performance compared with the MKL sparse QR on Intel Xeon 6248. Shengle Lin, Wangdong Yang, Haotian Wang 0006, Qinyun Tsai, Kenli Li 0001 |
SC | 2 |
| 2021 | Guest Editorial Special Issue on Smart IoT System: Opportunities by Linking Cloud, Edge, and AI
Wangdong Yang, Laurence T. Yang, Anthony T. Chronopoulos |
IEEE Internet Things J. | 1 |
| 2021 | Distributed matrix factorization based on fast optimization for implicit feedback recommendation
Lian Chen, Wangdong Yang, Kenli Li 0001, Keqin Li 0001 |
J. Intell. Inf. Syst. | 2 |
| 2021 | Performance analysis and optimization for SpMV based on aligned storage formats on an ARM processor
Yufeng Zhang 0001, Wangdong Yang, Kenli Li 0001, Dahai Tang, Keqin Li 0001 |
J. Parallel Distributed Comput. | 2 |
| 2021 | Velocity-Aware Parallel Encryption Algorithm with Low Energy Consumption for StreamsabstractIn the environment of cloud computing, the data produced by massive users form a data stream and need to be protected by encryption for maintaining confidentiality. Traditional serial encryption algorithms are poor in performance and consume more energy without considering the property of streams. Therefore, we propose a velocity-aware parallel encryption algorithm with low energy consumption (LECPAES) for streams in cloud computing. The algorithm parallelizes Advanced Encryption Standard (AES) based on heterogeneous many-core architecture, adopts a sliding window to stabilize burst flows, senses the velocity of streams using the thresholds of the window computed by frequency ratios, and dynamically scales the frequency of Graphics Processing Units (GPUs) to lower down energy consumption. The experiments for streams at different velocities and the comparisons with other related algorithms show that the algorithm can reduce energy consumption, but only slightly increases retransmission rate and slightly decreases throughput. Therefore, LECPAES is an excellent algorithm for fast and energy-saving stream encryption. Xiongwei Fei, Kenli Li 0001, Wangdong Yang, Keqin Li 0001 |
IEEE Trans. Big Data | 3 |
| 2020 | Optimizing partitioned CSR-based SpGEMM on the Sunway TaihuLight
Yuedan Chen, Guoqing Xiao 0001, Wangdong Yang |
Neural Comput. Appl. | 3 |
| 2020 | Analysis of energy efficiency of a parallel AES algorithm for CPU-GPU heterogeneous platforms
Xiongwei Fei, Kenli Li 0001, Wangdong Yang, Keqin Li 0001 |
Parallel Comput. | 3 |
| 2019 | A Pipeline Computing Method of SpTV for Three-Order Tensors on CPU and GPUabstractTensors have drawn a growing attention in many applications, such as physics, engineering science, social networks, recommended systems. Tensor decomposition is the key to explore the inherent intrinsic data relationship of tensor. There are many sparse tensor and vector multiplications (SpTV) in tensor decomposition. We analyze a variety of storage formats of sparse tensors and develop a piecewise compression strategy to improve the storage efficiency of large sparse tensors. This compression strategy can avoid storing a large number of empty slices and empty fibers in sparse tensors, and thus the storage space is significantly reduced. A parallel algorithm for the SpTV based on the high-order compressed format based on slices is designed to greatly improve its computing performance on graphics processing unit. Each tensor is cut into multiple slices to form a series of sparse matrix and vector multiplications, which form the pipelined parallelism. The transmission time of the slices can be hidden through pipelined parallel to further optimize the performance of the SpTV. Wangdong Yang, Kenli Li 0001, Keqin Li 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2019 | Performance-Aware Model for Sparse Matrix-Matrix Multiplication on the Sunway TaihuLight SupercomputerabstractGeneral sparse matrix-sparse matrix multiplication (SpGEMM) is one of the fundamental linear operations in a wide variety of scientific applications. To implement efficient SpGEMM for many large-scale applications, this paper proposes scalable and optimized SpGEMM kernels based on COO, CSR, ELL, and CSC formats on the Sunway TaihuLight supercomputer. First, a multi-level parallelism design for SpGEMM is proposed to exploit the parallelism of over 10 millions cores and better control memory based on the special Sunway architecture. Optimization strategies, such as load balance, coalesced DMA transmission, data reuse, vectorized computation, and parallel pipeline processing, are applied to further optimize performance of SpGEMM kernels. Second, we thoroughly analyze the performance of the proposed kernels. Third, a performance-aware model for SpGEMM is proposed to select the most appropriate compressed storage formats for the sparse matrices that can achieve the optimal performance of SpGEMM on the Sunway. The experimental results show the SpGEMM kernels have good scalability and meet the challenge of the high-speed computing of large-scale data sets on the Sunway. In addition, the performance-aware model for SpGEMM achieves an absolute value of relative error rate of 8.31 percent on average when the kernels are executed in one single process and achieves 8.59 percent on average when the kernels are executed in multiple processes. It is proved that the proposed performance-aware model can perform at high accuracy and satisfies the precision of selecting the best formats for SpGEMM on the Sunway TaihuLight supercomputer. Yuedan Chen, Kenli Li 0001, Wangdong Yang, Guoqing Xiao 0001, Xianghui Xie 0001, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | A parallel computing method using blocked format with optimal partitioning for SpMV on GPU
Wangdong Yang, Kenli Li 0001, Keqin Li 0001 |
J. Comput. Syst. Sci. | 1 |
| 2017 | A hybrid computing method of SpMV on CPU-GPU heterogeneous computing systems
Wangdong Yang, Kenli Li 0001, Keqin Li 0001 |
J. Parallel Distributed Comput. | 1 |
| 2017 | A parallel solving method for block-tridiagonal equations on CPU-GPU heterogeneous computing systems
Wangdong Yang, Kenli Li 0001, Keqin Li 0001 |
J. Supercomput. | 1 |
| 2016 | Practical parallel AES algorithms on cloud for massive users and their performance evaluationabstractSummary Many e‐business or social network servers have been constructed on cloud. On such open environments, private data of massive users have to be protected by encrypting, such as using Advanced Encryption Standard (AES), and furthermore, this process must be finished in a short time for users' better experience. This gives huge pressure on cloud servers, especially common servers, such as web servers. We urgently need an inexpensive and highly efficient method to relieve cloud servers' pressure. Fortunately, many cores of a graphics processing unit (GPU) can undertake this hard mission because of stronger computing power and lower price. The GPU environments can be virtualized on demand by cloud through the vCUDA technology. Of course, for those clouds not equipped with a GPU, a central processing unit (CPU) can still work as multithreads in parallel. Thus, in a cloud, AES can be parallelized using many cores of a GPU or multicores of a CPU with high efficiency and low cost. For typical cloud applications, such as web services, there are massive users and each one has short plaintext. If we simply parallelize AES in such an application, we cannot obtain better performance because of the GPU's extra data transferring cost. Thus, we coalesce the massive users' data and cut these data into same‐length slices for improving the performance of parallel AES as much as possible. So we design six parallel AES algorithms using GPU parallelism or CPU parallelism, which differ in parallel scope and whether data are coalesced or cut to slices. Specifically, they are coalescent and sliced GPU (GCS), coalescent and unsliced GPU, uncoalescent GPU, coalescent and sliced CPU, coalescent and unsliced CPU, and uncoalescent CPU. Moreover, we implement them on two representative platforms and evaluate their performance. Through comparing their performance, GCS has the best performance among these algorithms. In a cloud with Nvidia GPUs, GCS is a more powerful algorithm for massive users' data encrypting, relatively. Copyright © 2015 John Wiley & Sons, Ltd. Xiongwei Fei, Kenli Li 0001, Wangdong Yang, Keqin Li 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | A secure and efficient file protecting system based on SHA3 and parallel AES
Xiongwei Fei, Kenli Li 0001, Wangdong Yang, Keqin Li 0001 |
Parallel Comput. | 3 |
| 2016 | A Hybrid Parallel Solving Algorithm on GPU for Quasi-Tridiagonal System of Linear EquationsabstractThere are some quasi-tridiagonal system of linear equations arising from numerical simulations, and some solving algorithms encounter great challenge on solving quasi-tridiagonal system of linear equations with more than millions of dimensions as the scale of problems increases. We present a solving method which mixes direct and iterative methods, and our method needs less storage space in a computing process. A quasi-tridiagonal matrix is split into a tridiagonal matrix and a sparse matrix using our method and then the tridiagonal equation can be solved by the direct methods in the iteration processes. Because the approximate solutions obtained by the direct methods are closer to the exact solutions, the convergence speed of solving the quasi-tridiagonal system of linear equations can be improved. Furthermore, we present an improved cyclic reduction algorithm using a partition strategy to solve tridiagonal equations on GPU, and the intermediate data in computing are stored in shared memory so as to significantly reduce the latency of memory access. According to our experiments on 10 test cases, the average number of iterations is reduced significantly by using our method compared with Jacobi, GS, GMRES, and BiCG respectively, and close to those of BiCGSTAB, BiCRSTAB, and TFQMR. For parallel mode, the parallel computing efficiency of our method is raised by partition strategy, and the performance using our method is better than those of the commonly used iterative and direct methods because of less amount of calculation in an iteration. Kenli Li 0001, Wangdong Yang, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | An iteration-based hybrid parallel algorithm for tridiagonal systems of equations on multi-core architecturesabstractSummary An optimized parallel algorithm is proposed to solve the problem occurred in the process of complicated backward substitution of cyclic reduction during solving tridiagonal linear systems. Adopting a hybrid parallel model, this algorithm combines the cyclic reduction method and the partition method. This hybrid algorithm has simple backward substitution on parallel computers comparing with the cyclic reduction method. In this paper, the operation count and execution time are obtained to evaluate and make comparison for these methods. On the basis of results of these measured parameters, the hybrid algorithm using the hybrid approach with a multi‐threading implementation achieves better efficiency than the other parallel methods, that is, the cyclic reduction and the partition methods. In particular, the approach involved in this paper has the least scalar operation count and the shortest execution time on a multi‐core computer when the size of equations meets some dimension threshold. The hybrid parallel algorithm improves the performance of the cyclic reduction and partition methods by 19.2% and 13.2%, respectively. In addition, by comparing the single‐iteration and multi‐iteration hybrid parallel algorithms, it is found that increasing iteration steps of the cyclic reduction method does not affect the performance of the hybrid parallel algorithm very much. Copyright © 2015 John Wiley & Sons, Ltd. Guangping Tang, Wangdong Yang, Kenli Li 0001, Guoqing Xiao 0001, Keqin Li 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Performance Optimization Using Partitioned SpMV on GPUs and Multicore CPUsabstractThis paper presents a sparse matrix partitioning strategy to improve the performance of SpMV on GPUs and multicore CPUs. This method has wide adaptability for different types of sparse matrices, and is different from existing methods which only adapt to some particular sparse matrices. In addition, our partitioning method can obtain dense blocks by analyzing the probability distribution of non-zero elements in a sparse matrix, and result in very low proportion of zero padded. We make the following significant contributions. (1) We present a partitioning strategy of sparse matrices based on probabilistic modeling of non-zero elements in a row. (2) We prove that our method has the highest mean density compared with other strategies according to certain given ratios of partition obtained from the computing powers of heterogeneous processors. (3) We develop a CPU-GPU hybrid parallel computing model for SpMV on GPUs and multicore CPUs in a heterogeneous computing platform. Our partitioning strategy has balanced load distribution and the performance of SpMV is significantly improved when a sparse matrix is partitioned into dense blocks using our method. The average performance improvement of our solution for SpMV is about 15.75 percent on multicore CPUs, compared to that of the other solutions. By considering the rows of a matrix in a unique order based on the probability mass function of the number of non-zeros in a row, the average performance improvement of our solution for SpMV is about 33.52 percent on GPUs and multicore CPUs of a heterogeneous computing platform, compared to that of the partitioning methods based on the original row order of a matrix. Wangdong Yang, Kenli Li 0001, Zeyao Mo, Keqin Li 0001 |
IEEE Trans. Computers | 1 |
| 2015 | Performance Analysis and Optimization for SpMV on GPU Using Probabilistic ModelingabstractThis paper presents a unique method of performance analysis and optimization for sparse matrix-vector multiplication (SpMV) on GPU. This method has wide adaptability for different types of sparse matrices and is different from existing methods which only adapt to some particular sparse matrices. In addition, our method does not need additional benchmarks to get optimized parameters, which are calculated directly through the probability mass function (PMF). We make the following contributions. (1) We present a PMF to analyze precisely the distribution pattern of non-zero elements in a sparse matrix. The PMF can provide theoretical basis for the compression of a sparse matrix. (2) Compression efficiency of COO, CSR, ELL, and HYB can be analyzed precisely through the PMF, and combined with the hardware parameters of GPU, the performance of SpMV based on COO, CSR, ELL, and HYB can be estimated. Furthermore, the most appropriate format for SpMV can be selected according to estimated value of the performance. Experiments prove that the theoretical estimated values and the tested values have high consistency. (3) For HYB, the optimal segmentation threshold can be found through the PMF to achieve the optimal performance for SpMV. Our performance modeling and analysis are very accurate. The order of magnitude of the estimated speedup and that of the tested speedup for each of the ten tested sparse matrices based on the three formats COO, CSR, and ELL are the same. The percentage of relative difference between an estimated value and a tested value is less than 20 percent for over 80 percent cases. The performance improvement of our algorithm is also effective. The average performance improvement of the optimal solution for HYB is over 15 percent compared with that of the automatic solution provided by CUSPARSE lib. Kenli Li 0001, Wangdong Yang, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |