EDBT 2026 Demo / reviewers in the wild / expert
Ruixin Mao
dblp:287/5402
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-8629-3630ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TSA-P: A Triple-Sparsity-Aware Reconfigurable Few-Spikes-Neuron-Based SNN Processor
Aoyu Shen, Ruixin Mao, Jun Zhou 0017 |
ISCAS | 3 |
| 2025 | CREST: An Efficient Conjointly-trained Spike-driven Framework for Event-based Object Detection Exploiting Spatiotemporal DynamicsabstractEvent-based cameras feature high temporal resolution, wide dynamic range, and low power consumption, which are ideal for high-speed and low-light object detection. Spiking neural networks (SNNs) are promising for event-based object recognition and detection due to their spiking nature but lack efficient training methods, leading to gradient vanishing and high computational complexity, especially in deep SNNs. Additionally, existing SNN frameworks often fail to effectively handle multi-scale spatiotemporal features, leading to increased data redundancy and reduced accuracy. To address these issues, we propose CREST, a novel conjointly trained spike-driven framework to exploit spatiotemporal dynamics in event-based object detection. We introduce the conjoint learning rule to accelerate SNN learning and alleviate gradient vanishing. It also supports dual operation modes for efficient and flexible implementation on different hardware types. Additionally, CREST features a fully spike driven framework with a multi-scale spatiotemporal event integrator (MESTOR) and a spatiotemporal-IoU (ST-IoU) loss. Our approach achieves superior object recognition & detection performance and energy efficiency compared with state of-the-art SNN algorithms on three datasets, providing an efficient solution for event-based object detection algorithms suitable for SNN hardware implementation. Ruixin Mao, Aoyu Shen, Jun Zhou 0017 |
AAAI | 1 |
| 2025 | VSLAM-BA: Algorithm and Hardware Co-Design for High Performance and Energy-Efficient Visual SLAM Backend Hardware AcceleratorabstractVisual Simultaneous Localization and Mapping (VSLAM) is a key localization technology for emerging applications such as autonomous driving and uncrewed aerial vehicles (UAVs). Compared with VSLAM frontend, VSLAM backend plays a more important role as it is employed to improve the localization accuracy. However, the VSLAM backend usually uses Bundle Adjustment (BA) as its core optimization method which is well-known for its large scale of problem construction, high computational complexity and high serialization of data processing, making it difficult to achieve high performance and energy efficiency on platforms such as CPUs or GPUs. Although there are some VSLAM backend accelerators proposed recently for addressing the above issues, they did not well exploit the data regularity and computational characteristics, resulting in limited performance/energy efficiency improvements or degraded accuracy. In this work, we propose VSLAM-BA which is a high performance and energy-efficient VSLAM backend accelerator with algorithm-hardware co-design. On the algorithm level, a keyframe-split-based Schur elimination scheme is proposed to reduce latency, power consumption and memory storage while maintaining accuracy. On the hardware level, a column-folding-based computing architecture is proposed to boost performance and energy efficiency. A loading-sensitive matrix-computing technique with an adaptive task scheduler is proposed to reduce the latency and energy consumption. Further, a recyclable computing technique with point-aware solver is proposed to reduce the memory and energy consumption. The experimental results show that the proposed VSLAM-BA achieves the highest performance (380 fps) and the highest energy efficiency (0.51 mJ per frame) with high accuracy and low memory storage, compared with the SOTA designs. The proposed accelerator can work with different VSLAM frontend for backend optimization of localization accuracy. Ye Liu 0011, Xiuyuan Qi, Shuang Hao 0005, Zili Huang, Neng Zhao, Ruixin Mao, Sixu Li, Ang Hu, Yu Long 0005, Shanshan Liu 0001, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2024 | Stellar: Energy-Efficient and Low-Latency SNN Algorithm and Hardware Co-Design with Spatiotemporal ComputationabstractThe brain-inspired Spiking Neural Network (SNN) has great potential to reduce energy consumption in AI applications. However, the state-of-the-art SNN algorithms focus on high accuracy and large sparsity by constructing complex neuron models with sparse spike generation, leading to low energy efficiency and high latency. The state-of-the-art SNN hardware designs are hard to exploit high data reuse and parallel processing dataflows due to the irregularity and time-dependency of the spikes. To address the above issues, in this work we propose STELLAR, an algorithm-hardware co-design framework exploiting rich spatiotemporal dynamics of the SNN for high energy efficiency and low latency while maintaining high accuracy. Firstly, based on the Few Spikes (FS) neuron, we propose few spikes backpropagation (FSBP) and its training flow with strong hardware awareness to adaptively train the deep SNN for a short time window and few spikes. The resulting sparse SNN enjoys rapid inference with few synaptic operations and competitive accuracy. Secondly, we propose a dedicated SNN architecture and spatiotemporal Row Stationary (stRS) dataflow to exploit large sparsity brought by the proposed algorithm for highly parallel and energy -efficient computation. Several techniques have been proposed to boost energy efficiency and speedup while maintaining accuracy, including the window-based parallel processing technique and the spatiotemporal encoding-based computation architecture. The experimental results show that 1) on the algorithm level, STELLAR outperforms the state-of-the-art SNN models with significantly fewer spikes and shorter time window on both static and neuromorphic datasets with higher or comparable accuracy; 2) on the architecture level, compared with several SOTA SNN hardware designs, STELLAR achieves up to 8.1 × energy efficiency and 7.1 × speedup. Ruixin Mao, Ye Liu 0011, Jun Zhou 0017 |
HPCA | 1 |
| 2021 | A Fast and Energy-Efficient SNN Processor With Adaptive Clock/Event-Driven Computation Scheme and Online LearningabstractIn the recent years, the spiking neural network (SNN) has attracted increasing attention due to its low energy consumption and online learning potential. However, the design of SNN processor has not been thoroughly investigated in the past, resulting in limited performance and energy consumption. In this work, a fast and energy-efficient SNN processor with adaptive clock/event-driven computation scheme and online learning capability has been proposed. Several techniques have been proposed to reduce the computation time and energy consumption, including Adaptive Clock- and Event-Driven Computing Scheme, Neighboring PE Borrowing Technique, Compressed Spike Routing Technique and Reconfigurable PE for Inference and Learning. Implemented on a Virtex-7 FPGA, the proposed design achieves computation time of 3.15 ms/image, inference energy consumption of$0.028~\mu $J/synapse/image and online learning energy consumption of$0.297~\mu $J/synapse/image for the MNIST 10-class dataset, which outperform several state-of-the-art SNN processors. The proposed SNN processor is suitable for real-time and energy-constrained applications. Sixu Li, Zhaomin Zhang, Ruixin Mao, Jianbiao Xiao, Liang Chang 0002, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |