EDBT 2026 Demo / reviewers in the wild / expert
Ao Hu
dblp:271/4783
· DBLP profile ↗
12ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MRDNet: Multivariable Relational Decomposition Network for Multivariate Time Series Forecasting
Ao Hu, Liangjian Wen, Yong Dai 0001, Dongkai Wang, Jun Wang 0089, Jiang Duan |
Knowl. Based Syst. | 1 |
| 2026 | TimeCNN: Refining inscross-variable interaction on time point for time series forecasting
Ao Hu, Liangjian Wen, Yong Dai 0001, Shiyi Qi, Jun Wang 0089, Xun Zhou 0001, Dongkai Wang, Zenglin Xu, Jiang Duan |
Neural Networks | 1 |
| 2026 | FDNet: High-frequency disentanglement network with information-theoretic guidance for multivariate time series forecasting
Ao Hu, Liangjian Wen, Jiang Duan, Yong Dai 0001, Dongkai Wang, Shudong Huang, Jun Wang 0089, Zenglin Xu |
Pattern Recognit. | 1 |
| 2024 | SpaHet: A Software/Hardware Co-design for Accelerating Heterogeneous-Sparsity based Sparse Matrix MultiplicationabstractSparse general matrix-matrix multiplication is widely used in data mining applications. Its irregular memory access patterns limit the performance of general-purpose processors, thus motivating many FPGA-based hardware innovations in recent years. Nevertheless, existing accelerators fail to efficiently support heterogeneous input matrix sparsity, which is universal in various real-world applications. With in-depth experimental analysis, we observe that their performance is bottlenecked by their fixed tiling mechanisms, which only alleviate the irregularity of one input matrix. Based on the observation, we propose SpaHet, a software/hardware co-design to accelerate heterogeneous-sparsity based sparse matrix multiplication. SpaHet adopts a dual-adaptive sliding window mechanism to cover the reuse characteristics of both input matrices simultaneously. With a specialized exploration algorithm, the window-based mechanism can automatically find the optimal tiling strategy instead of applying a fixed one based on empirical experience. A sparsity-aware merge tree is also proposed to maximize the output matrix reuse via accumulating intermediate results thoroughly. Our results on a Xilinx Alveo U280 accelerator card show that SpaHet outperforms state-of-the-art CPU-, GPU- and FPGA-based solutions by 7.71×, 1.1×, and 2.74× in performance, respectively. Haoqin Huang, Pengcheng Yao, Zhaozeng An, Ao Hu, Peng Xu 0003, Long Zheng 0003, Xiaofei Liao, Hai Jin 0001 |
DAC | 5 |
| 2024 | Live Demonstration: A Reconfigurable, Energy-efficient and High-frame-rate EKF-SLAM Accelerator Based SoC Design for Autonomous Mobile Robot ApplicationsabstractThis demonstration shows a Extend Kalman Filter-Simultaneous Localization And Mapping (EKF-SLAM) accelerator based System On Chip (SoC) design for Autonomous Mobile Robots (AMR). The AMR platform consists of a multi-sensor system with a wheel encoder and LiDAR, and a ZYNQ-7000 FPGA based SoC featuring an EKF-SLAM hardware accelerator. This AMR system achieves real-time SLAM with significant energy efficient improvement against the state-of-the-art designs. Dingcheng Jiang, Bingqiang Liu, Ao Hu, Yequan Zhao, Minjie Bao, Zhendong Fan, Zixuan Shen, Ke Wang 0028, Chao Wang 0096 |
ISCAS | 4 |
| 2024 | Cross-Scale Attention for Long-Term Time Series ForecastingabstractTransformer-based models, especially PatchTST, have demonstrated remarkable success in time series forecasting tasks. However, the unique nature of time series data, which often contains jitter, and noise, and has inherently lower information density compared to images and texts, poses significant challenges. Specifically, the ViT-inspired patching design is suboptimal for time series data due to the sparse semantic relationships in such data. Moreover, modelling these sparse semantic relationships requires more resources and longer processing times. To overcome these limitations, we introduce a novel approach that leverages cross-scale attention interaction via a multi-scale patching technique. Our method initially treats the entire sequence as a single patch, then progressively divides it into increasing patches, doubling each time. This strategy improves the information density within each patch and reduces the total number of patches needed to model the time series to typically just seven effectively. We design a single layer of attention to model the cross-scale relationships among patches. These qualities significantly enhance computational efficiency. Extensive experiments have shown that our method surpasses or closely approaches existing methods in time-series forecasting benchmarks. Additionally, it achieves a speed that is 12x times faster than PatchTST on the large dataset. Liangjian Wen, Quan Hu, Ao Hu |
IEEE Signal Process. Lett. | 4 |
| 2024 | An Efficient GCNs Accelerator Using 3D-Stacked Processing-in-Memory ArchitecturesabstractGraph Convolutional Networks (GCNs) hold great promise in facilitating machine learning on graph-structured data. However, the sparsity of graphs often results in a significant number of irregular memory accesses, leading to inefficient data movement for existing GCNs accelerators. With the advancement of 3D stacked technology, the processing-in-memory (PIM) architecture has emerged as a promising solution for graph processing. Nevertheless, existing PIM accelerators are confronted with the challenges of irregular remote access in the aggregation phase of GCNs and dynamic workload variations between phases. In this paper, we present GCNim, a PIM accelerator based on 3D stacked memory, which features two key innovations in terms of the computation model and hardware designs. First, we present a PIM-based hybrid computation model, which employs a remote merging strategy to achieve the outer product in aggregation and the row-wise product in combination. Second, GCNim builds a three-stage aggregation and combination pipeline and integrates unified processing elements (PEs) supporting these three stages at the bank level, achieving load balance among PEs through a lightweight data placement algorithm. Compared with the state-of-the-art software frameworks running on CPUs and GPUs, GCNim achieves an average speedup of 3,736.06× and 76.56×, respectively. Moreover, GCNim outperforms the state-of-the-art GCN hardware accelerators, I-GCN, PEDAL, FlowGNN, and GCIM, with average speedups of 3.35×, 8.97×, 2.24×, and 5.58×, respectively. Ao Hu, Long Zheng 0003, Qinggang Wang, Jingrui Yuan, Haifeng Liu 0003, Linchen Yu, Xiaofei Liao, Hai Jin 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | PhGraph: A High-Performance ReRAM-Based Accelerator for Hypergraph ApplicationsabstractHypergraph processing has emerged as an effective approach to analyze complex multilateral relationships in real-world scenarios. Existing hypergraph processing solutions based on conventional architectures are severely bottlenecked by off-chip memory accesses. In this paper, we propose the first Processing-In-Memory (PIM)-featured ReRAM-based hypergraph accelerator, dubbed PhGraph, which facilitates performance-and energy-efficient hypergraph processing. On the hardware level, PhGraph integrates analog memristor-based PIM (with high matrix-grained parallelism) and digital memristor-based PIM (for high bipartite-edge-grained efficiency) into one standalone solution. On the software level, an overlap-aware hypergraph partitioning mechanism is proposed to polarize hypergraph workloads into matrix-formatted dense and bipartite-edge-formatted sparse partitions for performance acceleration using analog memristor-based PIM and digital ones, respectively. In addition, PhGraph is equipped with load-balanced partition scheduling and algorithm mapping co-designs to boost hardware utilization and efficiency. Experimental results show that PhGraph outperforms the state-of-the-art CPU-, FPGA-, and ASIC-based solutions by up to 4,309.81×, 547.13×, and 166.76× in terms of performance, and 36,416.11×, 924.12×, and 41.44× in terms of energy-savings, respectively. Long Zheng 0003, Ao Hu, Qinggang Wang, Yu Huang 0013, Haoqin Huang, Pengcheng Yao, Shuyi Xiong, Xiaofei Liao, Hai Jin 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Hardware-Accelerated Hypergraph Processing with Chain-Driven SchedulingabstractBeyond ordinary graphs, hypergraphs are a graph representation to flexibly express complex multilateral relationships between entities. Hypergraph processing can be used to solve many real-world problems, e.g., machine learning, VLSI design, and image retrieval. Existing hypergraph processing systems handle a hypergraph in order of its hyperedge and vertex indices. This makes processing hypergraphs on generalpurpose architectures suffer significantly from excessive offchip memory accesses, most of which however are redundant in frequently accessing overlapped hyperedges and vertices, but the index-ordered scheduling destroys this potential locality.In this paper, we propose a novel Generate-Load-Apply (GLA) execution model to improve locality in hypergraph processing. The key insight of GLA is to use a concept of chain to characterize the overlapped feature of a hypergraph, exposing data reuse opportunities missed in existing hypergraph systems. The precondition of driving GLA model is to generate expected chains on the fly, but the software solution is so expensive that its overheads may outweigh the benefits achieved from the chain-driven scheduling. We further present ChGraph, the first hardware-accelerated hypergraph processing engine near each core. ChGraph is specialized in accelerating the chain generation and the chain-guided data loading (to hide memory access latency) while the general-purpose cores are responsible only for handling the apply operations of GLA. We evaluate ChGraph against a state-of-the-art hypergraph processing system Hygra on six hypergraph algorithms using five large real-world hypergraphs. Results on a simulated 16core system show that ChGraph reduces the number of offchip memory accesses by up to 4.56× and achieves up to 4.73× speedup while introducing only 0.26% area overhead. Qinggang Wang, Long Zheng 0003, Jingrui Yuan, Yu Huang 0013, Pengcheng Yao, Chuangyi Gui, Ao Hu, Xiaofei Liao, Hai Jin 0001 |
HPCA | 7 |
| 2022 | A Data-Centric Accelerator for High-Performance Hypergraph ProcessingabstractHypergraph processing has emerged as a powerful approach for analyzing complex multilateral relationships among multiple entities. Past research on building hypergraph systems suggests that changing the scheduling order of bipartite edge tasks can improve the overlap-induced data locality in hypergraph processing. However, due to the complex intertwined connections between vertices and hyperedges, it is almost impossible to find a locality-optimal scheduling order. Thus, these task-centric hypergraph systems often suffer from substantial off-chip communications. In this paper, we first propose a novel data-centric Load-Trigger-Reduce (LTR) execution model to exploit fully the locality in hypergraph processing. Unlike a task-centric model that loads the required data along with a task, our LTR model invokes tasks as per the data used. Specifically, once the hypergraph data is loaded into the on-chip memory, all of its relevant computation tasks will be triggered simultaneously to output intermediate results, which are finally reduced to update the final results. Our LTR model enables all hypergraph data to be accessed once in each iteration. To fully exploit the LTR performance potential, we further architect an LTR-driven hypergraph accelerator, XuLin, which features with an adaptive data loading mechanism to minimize the loading cost via chunk merging at runtime. XuLin is also equipped with a priority-based differential data reduction scheme to reduce the impact of conflicting updates on performance. We have implemented XuLin both on a Xilinx Alveo U250 FPGA card and using a cycle-accurate simulator. The results show that XuLin outperforms the state-of-the-art hypergraph processing solutions Hygra and ChGraph by $20.47 \times$ and $8.77 \times$ on average, respectively. Qinggang Wang, Long Zheng 0003, Ao Hu, Yu Huang 0013, Pengcheng Yao, Chuangyi Gui, Xiaofei Liao, Hai Jin 0001, Jingling Xue |
MICRO | 3 |
| 2021 | Compound lower limb vibration training rehabilitation robotabstractSummary This paper uses a bionic lower limb rehabilitation mechanism combined with a running frame structure to simulate human gait movement for lower limb rehabilitation training. The vibration excited by the vertical vibration module causes the muscle to oscillate, and the mechanical vibration excites the neuromuscular system to obtain corresponding rehabilitation functions. The robot system is modeled in three dimensions, and the static analysis and modal analysis of the running frame are carried out. The mechanism and application of vertical vibration in the key technology are clarified, and the vibration element in the vertical vibration device is analyzed. The vibration theory is deduced, and the vibration displacement figure is drawn using simulation software. The related cooperative vibration spring has also been analyzed with single‐degree‐of‐freedom damping vibration, and the spring's center of mass movement, momentum, and position changes over time are illustrated. The design of the robot system solves the current situation of single movement of the lower limb rehabilitation robot and unsatisfactory rehabilitation effect, laying a foundation for the practical application of the subsequent lower limb rehabilitation robot system. Qiang Yin 0008, Ao Hu, Hongjun Yang, Beihai Wang, Guoquan Zhang |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | CYSAS-S3: a novel dataset for validating cyber situational awareness related tools for supporting military operationsabstractThe lack of suitable datasets and evaluation processes entails one of the most challenging gaps on the digital transformation era, where data-driven solutions like machine learning algorithms constitute a key pillar of the digitalization, virtualization and analytical on the emerging cyber-physical and ergonomic capabilities. This problem is even greater in the cyber defence domain, where for security or technical reasons, there is not data publicly or on-demand available concerning the role of the cyberspace on military operations. In this context, the expression popularized by the machine learning community "you go to the war with the data you have, not the data you might want" can be literally applied. In order to contribute to overcome this gap, this paper introduces CYSAS-S3, a novel dataset designed and created as the result of a research action that explores the principal needs on datasets by cyber commands, resulting in the generation of a collection of samples that correlated the impact of Advanced Persistent Threat (APT) behaviours and each phase of their cyber kill chain, regarding mission-level operations and goals. Roumen Daton Medenou, Victor Manuel Calzado Mayo, Miriam Garcia Balufo, Miguel Páramo Castrillo, Francisco José González Garrido, Álvaro Luis Martínez, David Nevado Catalán, Ao Hu, David Sandoval Rodríguez-Bermejo, Jorge Maestre Vidal, Gerardo Ramis Pasqual De Riquelme, Antonio Berardi, Paolo De Santis, Francesco Torelli, Salvador Llopis Sánchez |
ARES | 8 |