Haotian Wang 0006

dblp:63/11345-6 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-0086-6301ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MoProteus: LLM-Driven Multi-Version Operator Generation for Energy-Aware Scheduling in Heterogeneous Cloud-Edge Environments
Haotian Wang 0006, Junshuang Ma, Zicong Wang, Wangdong Yang, Kenli Li 0001
Euro-Par (2)2
2026 Generating Sparsity Patterns for Inverse Preconditioning on SIMD Architectures
abstract
The Conjugate Gradient method is a prevalent iterative approach for solving sparse linear systems$Ax = b$, whereAis a symmetric positive definite matrix. In this scenario, the Factorized Sparse Approximate Inverse (FSAI) preconditioner is commonly used. The FSAI approximates$A^{-1}$by the matrix productGTG, whereGis a lower triangular matrix. The numerical properties of FSAI mainly depend on the sparsity pattern ofG. This work presents an approach to generate FSAI sparsity patterns tailored to SIMD architectures, which play a pivotal role in current high-performance computing systems. First, we implement FSAI-S, an FSAI preconditioner based on the SELL-C-σ format for SIMD architectures. To further improve iterative performance, we propose a padding-aware preconditioner, FSAI-P, which leverages the zero padding inherent in the SELL-C-σ to extend the sparsity pattern with minimal computational and storage overhead. We evaluate our approach on three SIMD architectures: a Skylake processor implementing the AVX-512 ISA, a RISC-V processor supporting the RISC-V “V” Vector ISA, and an Nvidia A100 GPU. Experimental results show that FSAI-P provides performance gains across all evaluated platforms. Specifically, average speedups are 12.51% (vs. FSAI) and 11.66% (vs. FSAI-cache) on Skylake, 57.79% (vs. FSAI) on RISC-V, and 18.33% (vs. FSAI) on the A100 GPU.
Hantao Xiong, Haotian Wang 0006, Chubo Liu, Wangdong Yang, Kenli Li 0001, Marc Casas
IEEE Trans. Computers2
2026 Enhancing Large Language Models Reasoning via Multi-Path Optimization on Knowledge Graph
Jiyong Liao, Chubo Liu, Yan Ding 0004, Haotian Wang 0006, Zhuo Tang, Kenli Li 0001, Keqin Li 0001
IEEE Trans. Knowl. Data Eng.4
2026 Polarity-Aware and Adaptive Sparse Aggregation for Implicit Heterophilic Graph Classification
abstract
Graph-structured data appears in domains such as molecular analysis, social networks, and program optimization, where graphs often exhibit implicit heterogeneity, as nodes may look homogeneous in type yet differ significantly in semantics or functionality. Graph Neural Networks (GNNs), while powerful on homophilic graphs, tend to degrade in such settings due to polarity confusion, over-smoothing, and inefficiency caused by dense propagation. We propose a polarity-aware framework for graph classification that addresses these challenges through adaptive directional sparse aggregation. The framework introduces a polarity-aware propagation mechanism that adaptively reinforces or inverts neighbor signals, mitigating contamination under heterophily. A polarity-guided sparse aggregation operator further alleviates over-smoothing, improves scalability by constraining redundant connections, and condenses information flow into more effective representations, while maintaining unbiased estimation with controlled variance. We provide theoretical analyses that characterize the computational complexity, stability properties, and expressive behavior of signed directional aggregation, offering theoretical insights into its computational, stability, and expressive properties. Extensive experiments on molecular and social graph benchmarks with implicit heterophily demonstrate consistent improvements in graph classification accuracy and efficiency. Our method achieves a 2.36% improvement when compared with the strongest baseline on each dataset. In addition, it improves accuracy by 4.53% on average on program optimization strategy recognition tasks, reaching 80.12% overall.
Haotian Wang 0006, Yan Ding 0004, Wangdong Yang, Zhuo Tang, Chubo Liu, Kenli Li 0001
IEEE Trans. Knowl. Data Eng.1
2026 cuFastTucker-2L: A Two-Level Optimization Parallel Algorithm for Solving FastTucker Decomposition on GPU Platform
Haotian Wang 0006, Wangdong Yang, Keqin Li 0001, Kenli Li 0001
IEEE Trans. Parallel Distributed Syst.2
2026 Adaptive Block-Wise Mapping With Intra-Block Resource Allocation for Multi-DNN Workloads on Heterogeneous Accelerator Systems
abstract
Deep neural networks (DNNs) dominate workloads on cloud and edge platforms. Meanwhile, the hardware platform towards the heterogeneous system with various accelerators. By mapping layers to their different preferred accelerators, the computation cost of each layer can be reduced. While mapping these layers on the same accelerator can reduce the inter-accelerator communication cost. These two costs are often competing and difficult to optimize simultaneously. Therefore, the core challenge in achieving efficient execution of DNN workloads on heterogeneous systems is: how to map layers to achieve the best trade-off between computation and communication costs. Existing works group layers into blocks and perform blockwise mapping to reduce inter-layer communication within blocks. However, when grouping layers, they typically rely on modelagnostic rules, which fail to hide critical inter-layer communication within blocks for diverse DNNs. Moreover, after block mapping, the lack of intra-block resource allocation further increases computation cost of block. In this paper, we proposeGHCoM, a novel block-wise mapping framework for exploring the effective cost trade-offs.GHCoMemploys an adaptive grouping strategy to guide layer grouping based on the topology of DNNs and dynamically adjust the grouping according to the trade-off target. Furthermore,GHCoMconsiders the fine-grained allocation of computation (i.e., processing elements) and communication (i.e., on-chip bandwidth) resources within each block to mitigate interlayer resource contention. To jointly optimize layer grouping, block-wise mapping and intra-block resource allocation,GHCoMleverages a two-level genetic algorithm (GA) with tailored encodings and operators that capture the interdependence across the entire design space. Experiments across various workloads and system configurations show thatGHCoMconsistently outperforms state-of-the-art baselines, achieving 1.08× to 4.79× speedup in execution latency and reducing energy consumption by 1.83% to 87.71%.
Zhenyu Nie, Haotian Wang 0006, Anthony T. Chronopoulos, Zhuo Tang, Kenli Li 0001, Chubo Liu
IEEE Trans. Parallel Distributed Syst.2
2026 EBFL: An Efficient Blockchain Framework for Federated Learning Services
abstract
Federated Learning (FL) has emerged as a key framework to deliver AI services, recognized for its capability to construct global models while ensuring individual data. Nevertheless, FL heavily relies on a central server, which introduces significant challenges for participants to collaborate effectively and substantially limits the scalability of FL. Blockchain-based FL (BFL) offers a promising solution by replacing the central server with a decentralized blockchain system, thereby establishing a secure and trustworthy environment for FL. However, current BFL approaches face challenges in balancing high computational overhead, consistency, and security. In view of this, this paper introduces EBFL, an efficient blockchain framework for FL services. EBFL incorporates both asynchronous and synchronous advantages. A DAG-based (Directed Acyclic Graph) asynchronous computation enhances computational efficiency by mitigating delays caused by slow devices and reducing unnecessary waiting due to frequent synchronized consensus. Simultaneously, a periodic synchronized consensus mechanism is introduced during asynchronous training to ensure consistency, thereby improving security and model accuracy. Additionally, taking into account the unique characteristics of FL, we have designed a series of operations tailored for EBFL to further enhance the performance. Experimental results demonstrate that, compared to traditional synchronous BFL (TBFL) approaches, EBFL achieved a maximum speedup of up to 2.38× while retaining 92% of their accuracy. Subsequently, in-depth analytical experiments show that EBFL excels in both convergence speed and security, thereby confirming its potential to balance computational efficiency, consistency, and security.
Ze Yin, Haotian Wang 0006, Chubo Liu, Yan Ding 0004, Keqin Li 0001, Kenli Li 0001
IEEE Trans. Serv. Comput.2
2025 An Input-Aware Sparse Tensor Compiler Empowered by Vectorized Acceleration
abstract
Sparsity is widely prevalent in real-world applications, yet existing compiler optimizations and code generation techniques for sparse computations remain underdeveloped. Sparse matrix-matrix multiplication (SpMM) is a representative operator in sparse computations, whose performance is often limited by the design of sparse formats and the extent of hardware architecture optimization. Most existing solutions achieve highperformance SpMM through two approaches: (1) meticulously designed kernels and specialized sparse formats, which require extensive manual effort, or (2) tensor compilers that support code generation, though these typically offer limited support for sparse patterns, making it challenging to adapt to complex sparsity patterns in practical applications. This paper presents SpMMTC, an input-aware sparse tensor compiler. Given a sparse matrix as input, SpMMTC analyzes its non-zero distribution and generates a vectorized kernel optimized for SpMM on the specific matrix. We evaluated SpMMTC on various workloads. It achieves speedups of 1.21 x to 2.97 x over state-of-the-art methods such as TACO, TVM, and ASpT on different multi-core processors. It also provides a speedup of up to $\mathbf{1. 5 2 x}$ for sparse MobileNetV1 inference on the edge device.
Xianhao He, Haotian Wang 0006, Jiapeng Zhang 0001, Wangdong Yang, Anthony T. Chronopoulos, Kenli Li 0001
DAC2
2025 SASTC: Spatial-Aware Sparse Tensor Completion for Large-Scale Traffic Data Recovery
Renqiu Ouyang, Haotian Wang 0006, Yikun Hu 0001, Wangdong Yang, Kenli Li 0001
IEEE Internet Things J.2
2025 A Context-Awareness and Hardware-Friendly Sparse Matrix Multiplication Kernel for CNN Inference Acceleration
abstract
Sparsification technology is crucial for deploying convolutional neural networks in resource-constrained environments. However, the efficiency of sparse models is hampered by irregular memory access patterns in sparse matrix multiplication kernels. Hardware-level support for 2:4 granularity in sparse tensor cores presents an opportunity for designing efficient sparse matrix multiplication kernels. Existing approaches often involve adjusting sparse structures or secondary sparsification, introducing additional computational errors. To tackle this challenge, we introduce a flexible 2:4 structured adaptive sparse matrix multiplication (FS-AMM) method, a hardware-friendly sparse matrix multiplication kernel that leverages model context to accelerate convolutional neural networks. First, we propose a model context-aware matrix pre-processing method that employs heuristic algorithms to estimate a loss of accuracy due to weight sparsity at each layer. Second, we design a hardware-friendly sparse storage format that combines 2:4 sparse and dense storage formats, enabling more versatile sparsity ratio selection. Third, we implement efficient matrix multiplication kernels to optimize GPU utilization. Finally, experimental results on A100 GPUs show that our method effectively utilizes the sparse tensor kernel and obtains an average 3.09 times speedup ratio compared to other sparse methods while maintaining a high accuracy.
Haotian Wang 0006, Yan Ding 0004, Weichen Liu 0001, Chubo Liu, Wangdong Yang, Kenli Li 0001
IEEE Trans. Computers1
2025 A Privacy-Preserving Scheme With High Utility Over Data Streams in Mobile Crowdsensing
abstract
Both truth discovery and pattern analysis are effective methods for extracting valuable insights from data streams in mobile crowdsensing. However, existing privacy-preserving schemes either suffer from low data utility or provide high utility at the cost of weak privacy protection. To address this challenge, we introduce a robust privacy-preserving scheme that facilitates high-utility truth discovery and pattern analysis over mobile crowdsensing data streams. Concretely, we leverage the Square Wave mechanism, a randomized reporting technique, to perturb the data to prevent privacy breaches. To reduce the utility loss caused by perturbation, we design a budget allocation algorithm. This algorithm ensures that adjacent timestamps with approximate data share a perturbed value derived from their accumulated budgets. Furthermore, to facilitate robust pattern analysis, we propose a data splitting method that divides the perturbed data into two parts: one part records patterns randomly, while the other part recovers the perturbed values. Theoretical analysis confirms that our scheme satisfies ω-event ϵ-differential privacy level. Extensive experiments conducted on four real-world datasets demonstrate that our scheme outperforms existing schemes, delivering more accurate results for both truth discovery and pattern analysis under the same privacy constraints.
Zhimao Gong, Jiapeng Zhang 0001, Haotian Wang 0006, Mingxing Duan, Keqin Li 0001, Kenli Li 0001
IEEE Trans. Inf. Forensics Secur.3
2025 High Performance OpenCL-Based GEMM Kernel Auto-Tuned by Bayesian Optimization
Shengle Lin, Guoqing Xiao 0001, Haotian Wang 0006, Wangdong Yang, Kenli Li 0001, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.3
2025 HiFA: A High-Performance and Flexible Acceleration Framework for Large-Size Number Theoretic Transform
abstract
Zero-Knowledge Proofs (ZKP) and Homomorphic Encryption (HE) are crucial for data privacy in applications like cloud, blockchain, and analytics. However, the real-world adoption often faces performance challenges, particularly in the execution of the Number Theoretic Transform (NTT) required for polynomial multiplication involving sizes beyond \(2^{20}\) and large integer widths (e.g., 256 bits). FPGAs offer a promising platform for acceleration, but efficiently implementing large-size NTTs remains difficult due to the limited on-chip resources. The widely adopted four-step NTT method, used to relieve the need for large on-chip memory, introduces performance bottlenecks. Initially, the traditional dataflow NTT architecture may not fully exploit available compute capability, which hinders achieving peak performance. Furthermore, during the matrix transpose phase, the non-sequential access to external High-Bandwidth Memory (HBM) causes inefficiency. To address these challenges, we introduce HiFA, an FPGA-based automatic accelerator framework designed for high-performance and flexible large-size NTT computations. HiFA utilizes a stacked NTT architecture for high parallelism, maximizing HBM throughput. It supports various decomposed polynomial sizes via a novel reordering module. Additionally, a specialized cyclic shuffle module is integrated to optimize data movement during the matrix transpose step, alleviating random memory access delay. HiFA also provides an automatic Design Space Exploration (DSE) framework that identifies optimal four-step decomposition parameters and generates corresponding hardware configurations. Our experiments show that the FPGA implementation of HiFA achieves an average speedup of 2.97× and up to 7.25× improvement in latency over prior state-of-the-art FPGA solutions. Compared to prior GPU-based methods, HiFA achieves an average energy efficiency gain of 2.24×.
Qilin Hu, Haotian Wang 0006, Chubo Liu, Keqin Li 0001, Kenli Li 0001
ACM Trans. Reconfigurable Technol. Syst.2
2024 TAPMM: A Traffic-Aware Page Mapping Method for Multi-level NUMA Systems
abstract
With the development of chiplet technology, the architecture of Non-Uniform Memory Access (NUMA) has become increasingly intricate. The placement of memory page significantly influences application performance in NUMA systems. We found that memory access bottlenecks occur between high-level NUMA domains consisting of multiple chiplets. In this paper, we introduce a Traffic-Aware Page Mapping Method (TAPMM) designed for multi-level NUMA systems. TAPMM conceptualizes the multi-level NUMA system as a memory access tree, utilizing hardware performance events to be aware of system traffic and identify the optimal page mapping method for bandwidth efficiency. Our experiments demonstrate that TAPMM achieves a speedup of up to 2.12× on a real commodity machine compared to existing optimization tools.
Fengkun Dong, Guoqing Xiao 0001, Haotian Wang 0006, Yikun Hu 0001, Kenli Li 0001, Wangdong Yang
DAC3
2024 BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration
abstract
Sparse tensor contraction (SpTC) is an important operator in tensor networks, which tends to generate a large amount of sparse high-dimensional data, placing higher demands on the computational performance and storage bandwidth of the processor. Using GPUs with powerful arithmetic characteristics is a reliable choice for accelerating SpTC, however, the high dimensionality and sparsity of tensor makes GPU-accelerated SpTC operators suffer from the difficulties of low computational intensity and high memory consumption. The recent introduction of Tensor Core Units (TCUs) on GPUs brings even more powerful arithmetic, which exacerbates the memory wall problem. To cope with the challenges, this paper proposes a new BCB format that linearizes the indices of multidimensional blocks to reduce block index accesses and uses a bitmap to store the distribution of non-zero elements in a block to reduce the storage overhead. A parallel blocking algorithm of BCB-SpTC is designed to divide the binary linear indices into free and contracted indexes to improve the pairing overhead of computational tasks. Then based on the characteristic computation method of TCUs, the proprietary filling method of TCUs is designed to overcome the inefficiency of parallel computation of sparse data on TCUs. Finally, experimental results on the A100 dataset show that BCB-SpTC improves the acceleration ratio by$1.1\times$to$21.3\times$over the existing SpTC GPU method.
Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Keqin Li 0001, Kenli Li 0001
IEEE Trans. Parallel Distributed Syst.2
2023 IAP-SpTV: An input-aware adaptive pipeline SpTV via GCN on CPU-GPU
Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Kenli Li 0001, Keqin Li 0001
J. Parallel Distributed Comput.1
2023 A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration
abstract
Analysis of multi-dimensional data, especially tensor decomposition, which extracts latent information, is becoming considerably popular. Although multi-dimensional sparse data is typically processed on multi-core processors, developing highly optimized GPU-basedSparseTensorMatrixChainMultiplication (SpTMCM) is challenging. The purpose of this paper is to investigate a novel approach named SpTMCM and to explore the discovery of SpTMCM coupled with the emerging computing core, Tensor Core Unit (TCU). In contrast to prior work, the proposed novel approach enables a uniform storage format and optimization approach for SpTMCM. We design a hybrid tensor format based on multi-dimensional tiling that divides the tensor depending on the tile threshold to address the inefficient memory accesses caused by the irregular nonzero distribution of the sparse tensor. Further, we develop a TCU-based tensor parallel algorithm with our novel approach to increase the memory bandwidth. Compared to state-of-the-art works, our method achieves$1.16\sim 24.12\times$speedup for SpMTTKRP and$5.07\sim 7.15\times$speedup for SpTTMChain across NVIDIA A100 GPU on a range of real-world sparse tensors.
Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Kenli Li 0001, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.1
2021 STM-multifrontal QR: streaming task mapping multifrontal QR factorization empowered by GCN
abstract
Multifrontal QR algorithm, which consists of symbolic analysis and numerical factorization, is a high-performance algorithm for orthogonal factorizing sparse matrix. In this work, a graph convolutional network (GCN) for adaptively selecting the optimal reordering algorithm is proposed in symbolic analysis. Using our GCN adaptive classifier, the average numerical factorization time is reduced by 20.78% compared with the default approach, and the additional memory overhead is approximately 4% higher than that of prior work. Moreover, for numerical factorization, an optimized tasks stream parallel processing strategy is proposed and a more efficient computing task mapping framework for NUMA architecture is adopted in this paper, which called STM-Multifrontal QR factorization. Numerical experiments on the TaiShan Server show average 1.22x performance gains over the original SuiteSparseQR. Nearly 80% of datasets have achieved better performance compared with the MKL sparse QR on Intel Xeon 6248.
Shengle Lin, Wangdong Yang, Haotian Wang 0006, Qinyun Tsai, Kenli Li 0001
SC3