VLDB 2026 Research / reviewers in the wild / expert
Xiaochen Hao
dblp:127/4879 · also Xiao-Chen Hao
· DBLP profile ↗
25ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Computer networks · 4 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Co-evolutionary reinforcement learning for industrial optimization: a multi-objective optimization framework for cement rotary kilns
Gang Liu 0033, Shangjian Xie, Xiaochen Hao, Shuaixiang Zhai, Mohan Liu |
Adv. Eng. Informatics | 3 |
| 2026 | Multi-model adversarial and collaborative forecasting network for quality prediction modeling of high dimensional imbalanced sequences in process industry
Gaolu Huang, Xiaochen Hao, Junze Jiao, Xiaodie Ren |
Appl. Intell. | 3 |
| 2026 | Causal strength contrast-based abnormal path detection and root cause diagnosis for nonstationary industrial processes
Libin Wei, Xiaochen Hao, Tianqiang Lu |
Expert Syst. Appl. | 4 |
| 2026 | Predictive Monitoring of Industrial Processes and Quality Indices via Joint Training of Soft Sensor and Time-Series ForecastingabstractPredictive monitoring of industrial processes and quality indices, particularly free calcium oxide (f-CaO) in cement production, is of critical importance. However, it is significantly constrained by the asynchronous sampling between high-frequency process data and delayed laboratory quality measurements. To address this challenge, this article proposes a joint modeling framework that collaboratively optimizes a multiscale attention Kolmogorov–Arnold network (MSAKAN) soft sensor and a MambaMoE time-series forecasting model. The MSAKAN component combines multiscale attention with a Kolmogorov–Arnold network to enhance nonlinear mapping capabilities, while the MambaMoE component employs a mixture-of-experts mechanism for multitask forecasting of future process variables. A cascaded adversarial-inspired joint training strategy further bridges the temporal disparity between process and quality data by enabling mutual adaptation between the two models. Experimental results on a real-world cement plant dataset demonstrate the effectiveness of the proposed framework. The MSAKAN model achieves a pretraining coefficient of determination ($R^{2}$) of 0.8724, and the jointly trained system reduces the mean absolute error of predictive f-CaO monitoring to 0.1573. These results confirm the synergy of the proposed approach and show substantial performance gains over state-of-the-art baselines. Yonghang Li, Xiaochen Hao, Xunian Yang, Xingzhi Zheng, Libin Wei |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | A Unified Synthesis Framework for Dataflow Accelerators Through Multi-level Software and Hardware Intermediate Representations
Xiaochen Hao, Ruifan Xu, Yun Liang 0001 |
APPT | 1 |
| 2025 | StructILU: Dependency-Preserving Incomplete LU with Hierarchical Parallelism for Structured Grid PDEs on GPUsabstractThe Incomplete LU (ILU) computation is a crucial component for solving large-scale sparse linear systems arising from partial differential equations (PDEs), many of which are discretized on structured grids.However, due to inherent loop-carried data dependencies in ILU computation, implementing it on GPUs with massive computing units poses significant challenges.Existing methods either experience Hao Luo 0015, Qianchao Zhu, Xiaochen Hao, Chunxi Lei, Chengdi Ma, Yun Liang 0001, Chao Yang 0002 |
ICS | 3 |
| 2025 | Telos: A Dataflow Accelerator for Sparse Triangular Solver of Partial Differential EquationsabstractPartial Differential Equations (PDEs) serve as the backbone of numerous scientific problems.Their solutions often rely on numerical methods, which transform these equations into large, sparse systems of linear equations.These systems, solved with iterative methods, exhibit structured sparsity patterns when derived from stencil-based numerical schemes.In preconditioned solvers, the sparse triangular solve procedure, SpTRSV, usually dominates the entire execution due to its loop-carried dependencies.Optimizing SpTRSV requires extracting parallelism from dependent computations.However, prior works have struggled to achieve both high parallelism and data locality, leading to suboptimal performance.We propose Telos, a dataflow accelerator for SpTRSV that exploits structured sparsity patterns in PDE solving.The dataflow execution leverages stencil patterns, efficiently utilizing pipeline parallelism to resolve data dependencies with minimal overhead.We tackle the challenge of complex data dependencies by proposing a plane-parallel pipelining technique that maps computations onto processing elements (PEs) while preserving data locality.A cross-plane communication aggregation technique is developed to streamline data transfers into a systolic manner.Our accelerator features effective pipelining of dependent computations and overlapping of computations with memory accesses.Experiments demonstrate that Telos delivers average speedups of 61×, 8×, and 11× over CPUs, GPUs, state-of-the-art accelerator, respectively. Xiaochen Hao, Hao Luo 0015, Chao Yang 0002, Yun Liang 0001 |
ISCA | 1 |
| 2025 | Scaleformer: A hierarchical transformer network utilizing multi-scale spatial-temporal features for electricity consumption prediction in cement raw material grinding system
Gang Liu 0033, Bocheng Cao, Xiaochen Hao, Jiaan Huang |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Dynamic multi-objective optimization method for production index of cement clinker firing process based on collaborative prediction strategy
Gang Liu 0033, Shangjian Xie, Xiaochen Hao, Mengke Yang, Xunian Yang, Xingxing Xu |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Productively Generating a High-Performance Linear Algebra Library on FPGAsabstractLinear algebra computations can be greatly accelerated using spatial accelerators on FPGAs. As a standard building block of linear algebra applications, BLAS covers a wide range of compute patterns that vary vastly in data reuse, bottleneck resources, matrix storage layouts, and data types. However, existing implementations of BLAS routines on FPGAs are stuck in the dilemma of productivity and performance. They either require extensive human effort or fail to leverage the properties of routines for acceleration. We introduce Lasa, a framework composed of a programming model and a compiler, designed to address the dilemma by abstracting (for productivity) and specializing (for performance) the architecture of a spatial accelerator. The programming model realizes systolic arrays using uniform recurrence equations and space-time transforms. Streaming tensors, an intuitive dataflow-style abstraction, is proposed to uniformly describe the movement, storage, and transpose of input and output data across the spatial components. According to streaming tensors, a customized memory hierarchy is automatically built on an FPGA by our compiler. The compiler further specializes the architecture with transparent optimizations on FPGAs. Using this framework, we develop a complete BLAS library, demonstrating performance in parity with expert-written HLS code for BLAS level 3 routines, 76%–94% machine peak for level 1 and 2 routines, and 1.6X–13X speedup by leveraging the matrix properties such as symmetry, triangularity, and bandness. Xiaochen Hao, Mingzhe Zhang 0002, Ce Sun 0001, Zhuofu Tao, Hongbo Rong, Yu Zhang 0086, Lei He 0001, Eric Petit 0002, Yun Liang 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2024 | POPA: Expressing High and Portable Performance across Spatial and Vector Architectures for Tensor ComputationsabstractThis paper aims at high and portable performance for tensor computations across spatial (e.g., FPGAs) and vector architectures (e.g., GPUs). The state-of-the-art usually address performance portability across vector architectures (CPUs and GPUs). However, they either miss FPGAs or do not achieve high performance. Without a common architectural abstraction, they program and optimize spatial and vector devices separately, causing low portability. Xiaochen Hao, Hongbo Rong, Mingzhe Zhang 0002, Ce Sun 0001, Hong H. Jiang, Yun Liang 0001 |
FPGA | 1 |
| 2024 | Hierarchical Power Co-Optimization and Management for LLM Chiplet DesignsabstractThe demand for efficient and high-performance hardware for large language models (LLMs) has driven the development of scalable chiplet design, which requires careful power optimization and management. This paper presents a co-optimization and management methodology for hierarchical power delivery of chiplet designs targeting LLM applications. To model LLM workload mapping and power delivery, we first build a scalable chiplet simulator, which demonstrates different power strategies have notable efficiency impact and require careful and thorough optimizations. We further develop a co-optimization framework ScalePoM for chiplet power management. Based on given LLM model and PPA requirements, ScalePoM can automatically explore the chiplet architecture and workload mapping for optimal hierarchical power delivery. Our co-optimization methodology is evaluated through two scaled LLM chiplets with different interconnect topologies, achieving an average of 45% and up to 62% energy saving for large language model inferences with various sparsity levels. Yanchi Dong, Xiaochen Hao, Yun Liang 0001, Ru Huang 0001, Le Ye |
ICCAD | 3 |
| 2024 | MatFactory: A Framework for High-Performance Matrix Factorization on FPGAsabstractMatrix factorization is a widely used powerful tool in signal processing, machine learning and high performance computing. For accelerating matrix factorization, FPGAs are suitable platforms, as they can build wide and deep pipelines with favorable power efficiency. Factorizing matrices on FPGAs is thus desirable; however, there is no infrastructure on FPGAs for matrix factorization so far, as it involves several challenges: applicability and scalability of the circuit, pipelining of irregular computing patterns, and effective data caching given the limited memory bandwidth. Mingzhe Zhang 0002, Xiaochen Hao, Hongbo Rong |
ICCAD | 2 |
| 2024 | Regression generative adversarial network based on bounded losses for prediction of free calcium oxide in cement clinker
Gaolu Huang, Xiaochen Hao, Hui Dang |
Adv. Eng. Informatics | 2 |
| 2023 | Lasa: Abstraction and Specialization for Productive and Performant Linear Algebra on FPGAsabstractLinear algebra can often be significantly expedited by spatial accelerators on FPGAs. As a broadly-adopted linear algebra library, BLAS requires extensive optimizations for routines that vary vastly in data reuse, bottleneck resources, matrix storage layouts, and data types. Existing solutions are stuck in the dilemma of productivity and performance. We introduce Lasa, a framework composed of a programming model and a compiler, that addresses the dilemma by abstracting (for productivity) and specializing (for performance) the architecture of a spatial accelerator. Lasa abstracts a compute and its I/O as two dataflow graphs. A compiler maps the graphs onto systolic arrays and a customized memory heirarchy. The compiler further specializes the architecture transparently. In this framework, we develop 14 key BLAS routines, and demonstrate performance in parity with expert-written HLS code for BLAS level 3 routines, >=80% machine peak performance for level 2 and 1 routines, and 1.6X-7X speed up by taking advantage of matrix properties of symmetry, triangularity and bandness. Xiaochen Hao, Mingzhe Zhang 0002, Ce Sun 0001, Zhuofu Tao, Hongbo Rong, Yu Zhang 0086, Lei He 0001, Eric Petit 0002, Yun Liang 0001 |
FCCM | 1 |
| 2023 | Monad: Towards Cost-Effective Specialization for Chiplet-Based Spatial AcceleratorsabstractAdvanced packaging offers a new design paradigm in the post-Moore era, where many small chiplets can be assembled into a large system. Based on heterogeneous integration, a chiplet-based accelerator can be highly specialized for a specific workload, demonstrating extreme efficiency and cost reduction. To fully leverage this potential, it is critical to explore both the architectural design space for individual chiplets and different integration options to assemble these chiplets, which have yet to be fully exploited by existing proposals. This paper proposes Monad, a cost-aware specialization approach for chiplet-based spatial accelerators that explores the tradeoffs between PPA and fabrication costs. To evaluate a specialized system, we introduce a modeling framework considering the non-uniformity in dataflow, pipelining, and communications when executing multiple tensor workloads on different chiplets. We propose to combine the architecture and integration design space by uniformly encoding the design aspects for both spaces and exploring them with a systematic ML-based approach. The experiments demonstrate that Monad can achieve an average of 16% and 30% EDP reduction compared with the state-of-the-art chiplet-based accelerators, Simba and NN-Baton, respectively. Xiaochen Hao, Zijian Ding, Jieming Yin, Yuan Wang 0001, Yun Liang 0001 |
ICCAD | 1 |
| 2023 | Improved game algorithm for spectrum resource optimization with variable number of channels
Bai Chen 0003, Xiaochen Hao, Rongrong Yin |
Ad Hoc Networks | 4 |
| 2023 | Elite-guided multi-objective cuckoo search algorithm based on crossover operation and information enhancement
Xunian Yang, Xiaochen Hao, Yonghang Li |
Soft Comput. | 2 |
| 2022 | Predictive control research for cement burning system using two-cycle coupling optimization
Quanwei Sun, Yakun Ji, Qingquan Xu, Xunian Yang, Xiaochen Hao |
Expert Syst. Appl. | 6 |
| 2021 | Real-time semantic segmentation with weighted factorized-depthwise convolution
Xiaochen Hao, Xingjun Hao, Yaru Zhang |
Image Vis. Comput. | 1 |
| 2020 | Distributed resource allocation optimisation algorithm based on particle swarm optimisation in wireless sensor networkabstractThis study concentrates on the optimal resource allocation problem in a wireless sensor network with rare spectrum resources and improper topology structure. To explore the interdependence of various resources and achieve better anti‐interference property, the authors analyse the problem of joint power control and channel allocation based on bit error rate (BER) model and energy consumption model. Specifically, low‐energy consumption and BER are both crucial design objectives for a number of multi‐hop wireless network applications with constrained network resources and battery‐powered sensors. As these two objectives that influenced by power control and channel allocation are conflicting with each other, it becomes important to achieve the trade‐off between them. Aiming at the aforementioned problem, this study formulates a multi‐objective optimisation model to minimise BER and energy consumption under the constraints of link interference, link capacity, and network connectivity. On the basis of this model, they propose a distributed resource allocation optimisation algorithm based on particle swarm optimisation (DRAPSO) to achieve Pareto optimal solutions. Furthermore, the information complexity and time complexity of DRAPSO are theoretically analysed. The simulation results show that DRAPSO can effectively increase network capacity, decrease energy consumption, and BER. Besides, the trade‐off of multi‐performances can be significantly achieved. Xiaochen Hao |
IET Commun. | 1 |
| 2019 | Integrating Cyber-Attack Defense Techniques into Real-Time Cyber-Physical SystemsabstractWith the rapid deployment of Cyber-Physical Systems (CPS), security has become a more critical problem than ever before, as such devices are interconnected and have access to a broad range of critical data. A well-known attack is ReturnOriented Programming (ROP) which can diverge the control flow of a program by exploiting the buffer overflow vulnerability. To protect a program from ROP attacks, a useful method is to instrument code into the protected program to do runtime control flow checking (known as Control Flow Integrity, CFI). However, instrumented code brings extra execution time, which has to be properly handled, as most CPS systems need to behave in a real-time manner. In this paper, we present a technique to efficiently compute an execution plan, which maximizes the number of executions of instrumented code to achieve maximal defense effect, and at the same time guarantees real-time schedulability of the protected task system with a new response time analysis. Simulation-based experimental results show that the proposed method can yield good quality execution plans, but performs orders of magnitude faster than exhaustive search. We also built a prototype in which a small auto-drive car is defended against ROP attacks by the proposed method implemented in FreeRTOS. The prototype demonstrates the effectiveness of our method in real-life scenarios. Xiaochen Hao, Mingsong Lv, Jiesheng Zheng, Zhengkui Zhang, Wang Yi 0001 |
ICCD | 1 |
| 2019 | Prediction of electricity consumption in cement production: a time-varying delay deep belief network prediction method
Xiaochen Hao, Zhaoxu Wang, Zeyu Shan, Yantao Zhao |
Neural Comput. Appl. | 1 |
| 2018 | Topology control game algorithm based on Markov lifetime prediction model for wireless sensor network
Xiaochen Hao, Dehua Geng, Bai Chen 0003 |
Ad Hoc Networks | 1 |
| 2016 | Distributed topology construction algorithm to improve link quality and energy efficiency for wireless sensor networks
Xiaochen Hao, Weijing Liu, Dehua Geng, Xi-Da Li |
J. Netw. Comput. Appl. | 1 |