VLDB 2026 Research / reviewers in the wild / expert
Chao Wang 0096
dblp:188/7759-96
· DBLP profile ↗
16ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0002-7460-7628ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Live Demonstration: An Energy-efficient SoC for Correlative Scan Matching Based 2D-LiDAR SLAM
Yulong Tan, Zixuan Shen, Bingqiang Liu, Chao Wang 0096 |
ISCAS | 7 |
| 2026 | D2TriPO-DETR: Dual-Decoder Triple-Parallel-Output Detection TransformerabstractVision-based grasping, though widely employed for industrial and household applications, still struggles with object stacking scenarios. Current methods face three major challenges: limited inter-object relationship understanding; poor grasping adaptation across different viewpoints; and error propagation. To address the above challenges, we propose D2TriPO-DETR, a dual-decoder transformer with three outputs, of which are object detection, manipulation relationship, and grasp detection. Specifically, a distributed attention perception module and a rotation attention invariance module are designed to address limited interobject relationship understanding and poor grasping adaptation across different viewpoints. These two modules are respectively integrated into the two parallel decoders to output the triple results simultaneously, partly eliminating task-level error propagation. Experimental results on the visual manipulation relationship dataset indicate that D2TriPO-DETR outperforms existing state-of-the-art methods across all metrics, e.g., +6.1% object detection recall, +6.7% manipulation relationship image accuracy, and +1.5% grasp detection accuracy. Extensive real-world experiments and quantitative results validate D2TriPO-DETR’s effectiveness. Menghao Pu, Chaoqun Han, Zhiping Chai, Pu Wen, Jihong Zhu 0002, Chao Wang 0096, Han Ding 0002, Xuguang Lan |
IEEE Trans. Ind. Informatics | 7 |
| 2025 | Live Demonstration: An Area and Energy Efficient Reconfigurable Cryptographic Accelerator Based SoC Design for Securing IoT DevicesabstractThis demonstration presents an energy and area efficient Reconfigurable Cryptographic Accelerator (RCA) SoC for secure communication in IoT devices. Built on a ZYNQ-7000 development board, the platform supports multiple block ciphers (DES, AES, SM4) and Hash functions (SHA-1, SHA-2, SM3). Users can follow prompts on the OLED screen to select the cryptographic algorithm via buttons and input data through a keyboard or choose large text files from SD card. The ARM Core and accelerator execute the cryptographic operation simultaneously, and energy efficiency is calculated based on power and computing time, showcasing the improved computing speed and energy efficiency of the proposed accelerator. Xvpeng Zhang, Bingqiang Liu, Lingyun Hu, Zixuan Shen, Zaisheng He, Dengke Xu, Bah-Hwee Gwee, Chao Wang 0096 |
ISCAS | 8 |
| 2025 | An Energy- and Resource-Efficient Parallel-Pipelined Pedestrian Detector With Multiscale Image Computation Scheduling for Always-On Intelligent Edge DevicesabstractHistogram of Oriented Gradients (HOG) and linear Support Vector Machine (SVM) have been widely used for pedestrian detection in applications like video surveillance, automatic driving, and intelligent robots. However, in Internet of Things (IoT) applications relying on intelligent edge devices, it is a big challenge to design a high frame-rate multi-scale pedestrian detector without sacrificing precision under strictly resource-limited and energy-constrained conditions. This paper proposes a HOG-SVM-based pedestrian detector with a novel multi-scale image scheduling method based parallel-pipelined multi-detector architecture to maintain a high frame rate with small hardware overhead, and an optimized inter-module pipeline design to minimize pipeline cycles and on-chip buffer costs. Besides, a fine-grained block-score Multiply-Accumulate (MAC) segmentation and mapping method is proposed for the SVM-classifier MAC array to reduce resource overhead while maintaining the same throughput. FPGA validation shows that as compared to the state-of-the-art design with 12 multi-scale detectors, our proposed design achieves a frame rate of 288 fps using only 2 parallel detectors, which reduces LUT, FF, BRAM, and Digital Signal Processor (DSP) usage by 68.9%, 76.1%, 63.3%, and 94.4%, respectively, while improving energy efficiency by 48.6%. ASIC implementation further improves the energy efficiency by 97% and at the same time increases the frame rate to 400 fps at 200 MHz. Zixuan Shen, Bingqiang Liu, Yulong Tan, Yuanjin Zheng, Chao Wang 0096, Jiang Tang |
IEEE Internet Things J. | 7 |
| 2025 | An Energy-Efficient, High-Frame-Rate, and Reconfigurable EKF-SLAM Processor With Full Acceleration for Autonomous Mobile RobotsabstractIn many intelligent edge applications involving Autonomous Mobile Robots (AMRs), efficient and real-time localization and mapping is a fundamental issue. Extended Kalman Filter Simultaneous Localization and Mapping (EKFSLAM) algorithm is a classic and successful solution to realize localization and mapping, while it is computationally intensive and poses a challenge for real-time tasks in small and micro robots. To address this issue, this work proposes an energy-efficient, highframe-rate, and reconfigurable EKF-SLAM processor. Firstly, a heterogeneous dual-core architecture is proposed to enable full acceleration of both matrix operations and nonlinear calculations in EKF-SLAM at the hardware architecture level. Secondly, a Reconfigurable Matrix Accelerator (RMA) and Reconfigurable Nonlinear Accelerator (RNA) are proposed to maximize data reuse and support diverse nonlinear functions at the data flow level. Thirdly, a data property-aware strategy is proposed at the data property level, which exploits matrix symmetry, sparsity, and dependency to reduce storage significantly and eliminate redundant computations. FPGA validation results show that the proposed design can achieve a frame rate of 774 fps and an energy efficiency of 0.66 mJ/frame, when performing mapping processes involving 60 landmarks at 100 MHz. Bingqiang Liu, Yequan Zhao, Minjie Bao, Zhendong Fan, Dingcheng Jiang, Zixuan Shen, Yulong Tan, Zaisheng He, Dengke Xu, Ke Wang 0028, Chao Wang 0096, Lining Sun |
IEEE Trans. Circuits Syst. I Regul. Pap. | 12 |
| 2025 | A 2.793 μW Near-Threshold Neuronal Population Dynamics Trajectory Filter for Reliable Simultaneous Localization and MappingabstractThis work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for the trajectory error correction task within a simultaneous localization and mapping workflow. A custom discretized procedural algorithm approximating a neuronal population dynamics-based inference operation is developed for mapping onto an ultra-lightweight digital macro featuring massively parallel in-situ processing techniques. Fabricated using a 40nm technology, the test chip features a$22\times 22$neuron array with 0.1358mm2 core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming sub-10-$\mu $W power. Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0096, Tony Tae-Hyoung Kim, Yuanjin Zheng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | Live Demonstration: A High-frame-rate and Energy-efficient SIFT Feature Extraction Accelerator Based SoC Design for AMR ApplicationsabstractThis demonstration presents a high-frame-rate and energy-efficient Scale-Invariant Feature Transform (SIFT) feature extraction accelerator based System on Chip (SoC) design. The platform implementing SIFT-based object recognition consists of an OV5640 camera, a SIFT hardware accelerator based on the ZYNQ-7000 SoC, and a personal computer (PC). The feature points and recognition results are displayed on the monitor in real-time at 60 frames per second (fps) with QVGA resolution for Autonomous Mobile Robot (AMR) applications. Zhenhui Duan, Bingqiang Liu, Zehua Yin, Zixuan Shen, Xupeng Zhang, Zaisheng He, Chao Wang 0096 |
ISCAS | 8 |
| 2024 | Live Demonstration: A Reconfigurable, Energy-efficient and High-frame-rate EKF-SLAM Accelerator Based SoC Design for Autonomous Mobile Robot ApplicationsabstractThis demonstration shows a Extend Kalman Filter-Simultaneous Localization And Mapping (EKF-SLAM) accelerator based System On Chip (SoC) design for Autonomous Mobile Robots (AMR). The AMR platform consists of a multi-sensor system with a wheel encoder and LiDAR, and a ZYNQ-7000 FPGA based SoC featuring an EKF-SLAM hardware accelerator. This AMR system achieves real-time SLAM with significant energy efficient improvement against the state-of-the-art designs. Dingcheng Jiang, Bingqiang Liu, Ao Hu, Yequan Zhao, Minjie Bao, Zhendong Fan, Zixuan Shen, Ke Wang 0028, Chao Wang 0096 |
ISCAS | 10 |
| 2024 | A Low-Latency and High-Accuracy Dual-Mode Neuron Design for Accelerating Neurological Diseases Simulation and AnalysisabstractBiological realism and computational efficiency are crucial for modeling neurological diseases with spiking neural networks (SNNs), as biological realism is necessary for observing neuron ion channel behaviors and computational efficiency is essential for simulating the action potentials of large-scale networks. However, existing spiking neuron models cannot achieve both high biological realism and computational efficiency, resulting in SNN constructed from a single type of neuron to make a compromise between these two attributes, thus reducing the SNN effectiveness in disease simulation and analysis. In this paper, we propose a dual-mode spiking neuron hardware design with an efficient reconfigurable architecture to achieve both biological realism and computational efficiency for diseases modeling. By exploiting the common arithmetic operators in Hodgkin-Huxley neuron and Adaptive Exponential (AdEx) neuron, our design can reuse the computational units including adder, multiplier, and CORDIC to efficiently realize these two neurons. An optimized pipeline design based on data flow dependency and Reconfigurable Fast-Convergence CORDIC is proposed to reduce overall computation latency, while a dynamic bit-width allocation strategy is employed to improve the implementation accuracy. FPGA implementation result shows that our design significantly improves computation latency and accuracy compared to previous neuron designs. Jiatong Guo, Jinxiang Gao, Zixuan Shen, Jingru Jiang, Wenjue Chen, Chao Wang 0096 |
TENCON | 8 |
| 2023 | An Energy-Efficient, Resource-Efficient and High Frame-Rate End-to-End Pedestrian Detector Using HOG-SVM for Intelligent Edge DevicesabstractThis paper proposes a Histogram of Oriented Gradients-Support Vector Machine (HOG-SVM) based pedestrian detector with an end-to-end fully-pipelined architecture to achieve a high frame rate by improving the throughput, and reduce the power consumption by minimizing the data movement. To further improve the energy efficiency under the high frame rate, a bit-width pruning method is used to remove the gray-scale converter's redundant data bit width, and a block-score normalization is employed to significantly reduce the normalizer's required divisions. The reduced computation amount also saves the hardware overhead while maintaining the same calculation accuracy. Besides, a modeling and analysis method of the SVM-classifier-Multiply-ACcumulate (MAC) array is proposed to further improve the energy efficiency and save the logic resources, by optimizing the array size with a hardware utilization of 98.4% while maintaining the same throughput. The FPGA implementation results of$640\times 480$video show a high frame rate of up to 439 fps @143 MHz and a high energy efficiency of 0.76 nJ/pixel with 46.7% fewer LUTs, 22.4% fewer registers, 88.3% fewer DSPs, compared to the state-of-the-art design. The ASIC implementation in 55 nm also confirms a high energy efficiency of 0.35 nJ/pixels at 613 fps and 200 MHz as well as a hardware overhead of 177 k gates and 108 Kbits SRAM. Jianhui Song, Bingqiang Liu, Zixuan Shen, Fengwei An, Chao Wang 0096, Jiang Tang |
IECON | 7 |
| 2023 | Deep convolutional transfer learning-based structural damage detection with domain adaptation
Zuoyi Chen, Chao Wang 0096, Jun Wu 0012 |
Appl. Intell. | 2 |
| 2023 | Deep transfer learning-based damage detection of composite structures by fusing monitoring data with physical mechanism
Cheng Liu 0004, Xuebing Xu, Jun Wu 0012, Haiping Zhu 0001, Chao Wang 0096 |
Eng. Appl. Artif. Intell. | 5 |
| 2022 | Energy-Efficient Intelligent Pulmonary Auscultation for Post COVID-19 Era Wearable Monitoring Enabled by Two-Stage Hybrid Neural NetworkabstractThis paper proposes an energy-efficient intelligent pulmonary auscultation system for post COVID-19 era wearable monitoring. This system consists of a tightly coupled two-stage hybrid neural network (TC-TSHNN) model and a corresponding multi-task training paradigm to improve prediction accuracy and generalization ability based on the fact that the number of COVID-19 patients is far less than that of normal people. At the first stage, two-category coarse classification is performed to identify normal and abnormal lung sounds. If the lung sound is abnormal, the second stage would be triggered to perform a four-category fine-grained classification. Besides, discrete wavelet transform is utilized for feature extraction, denoising and data reduction. In addition, advanced lightweight convolutional neural networks are used to reduce the model’s computation and improve the model’s performance. The hybrid network model can achieve 92% computation reduction and energy saving compared with a direct four-category classification when the input lung sound is normal, which is the majority of cases. Experiment results with inter-patient classification on the COVID-19 lung sound dataset from Tongji Hospital in Wuhan City and the ICBHI’17 dataset show that the proposed TC-TSHNN model can significantly reduce power consumption while maintaining competitive performance against the state-of-the-art work. Bingqiang Liu, Ziyuan Wen, Hongling Zhu, Jinsheng Lai, Jiajun Wu 0006, Heng Ping, Wenqing Liu, Guoyi Yu, Zuozhu Liu, Hesong Zeng, Chao Wang 0096 |
ISCAS | 12 |
| 2022 | An Energy-Efficient SIFT Based Feature Extraction Accelerator for High Frame-Rate Video ApplicationsabstractVisual feature extraction is a key technology of computer vision for intelligent video processing. Efficient feature extraction is a fundamental problem in computer vision applications. Scale-Invariant Feature Transform (SIFT) is one of the most popular feature extraction algorithms because SIFT features are invariant to image scale and rotation and robust to changes in illumination and noise. However, SIFT is a computationally-intensive and power-hungry algorithm, which needs to be accelerated by efficient hardware design to achieve both high-speed feature extraction and high energy efficiency for many high frame-rate video applications at Artificial-intelligent Internet of Things edges. In this work, an energy-efficient SIFT based feature extraction accelerator is proposed. In the Gaussian pyramid and Differences of Gaussian (DoG) pyramid construction process, three design methods are proposed to reduce power consumption and improve information fidelity: a fast and slow dual clock domain design method with a reconfigurable design strategy is proposed to reduce the computation resources; a partial sum reuse design method is proposed to further reduce the computation resources and the amount of computation; a dynamic padding design method is proposed to solve the problem of information loss at image edges and corners after convolution operation. In the keypoint descriptor generation process, an optimized algorithm using circular region and polar coordinates is proposed to parallelize the main orientation assignment and descriptor generation to achieve high-speed processing, while maintaining a comparable matching accuracy with the state-of-the-art designs. The experiment results show that the proposed SIFT hardware accelerator is able to extract features by up to 162 frames per second ($640\times 480$pixels) under 100 MHz, with the power consumption of 364.26 mW and energy efficiency of 2.25 mJ/frame based on 180 nm technology, which is suitable for many high frame-rate AIoT applications including autonomous driving cars and unmanned aerial vehicles. Bingqiang Liu, Zehua Yin, Xvpeng Zhang, Xiaofeng Hu, Guoyi Yu, Yuanjin Zheng, Chao Wang 0096, Xuecheng Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2022 | In Situ Aging-Aware Error Monitoring Scheme for IMPLY-Based Memristive Computing-in-Memory SystemsabstractStateful logic through memristor is a promising technology to build Computing-in-Memory (CIM) systems. However, aging-induced degradation of memristors’ threshold voltage imposes a major challenge to the reliability and guardbands estimation of memristive CIM systems, especially the Material Implication (IMPLY) logic based CIM systems. In this paper, a novel in-situ aging-aware error monitoring scheme for memristor-based IMPLY logic is proposed. The proposed in-situ error monitoring scheme can achieve faster error detection speed and higher detection accuracy than the straightforward program-verify monitoring scheme. Simulation results under Monte-Carlo simulation show that the proposed monitoring scheme can effectively detect the major operation failures existing in IMPLY logic operations with a detection accuracy up to 99.95%. Moreover, a case study of error monitoring design of 4-bit IMPLY-based adder is carried out. The analysis result exhibits that the proposed in-situ monitoring scheme can achieve 75.2% improvement on the detection speed against the program-verify scheme. Further analysis on a convolution filter in VGG-11 based Binarized Neural Network shows that 74% improvement on the detection speed can also be achieved by using the proposed monitoring scheme, which suggests that the proposed in-situ error monitoring scheme is an efficient solution to improve the reliability of IMPLY-based memristive CIM systems. Jiajun Wu 0006, Xinglong Ji, Guoyi Yu, Chao Wang 0096 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2021 | Efficient Design of Spiking Neural Network With STDP Learning Based on Fast CORDICabstractIn emerging Spiking Neural Network (SNN) based neuromorphic hardware design, energy efficiency and on-line learning are attractive advantages mainly contributed by bio-inspired local learning with nonlinear dynamics and at the cost of associated hardware complexity. This paper presents a novel SNN design employing fast COordinate Rotation DIgital Computer (CORDIC) algorithm to achieve fast spike timing–dependent plasticity (STDP) learning with high hardware efficiency. In this study, a system design and evaluation method of CORDIC-based SNN is proposed for finding optimal CORDIC type and precision, from theoretical CORDIC-level error to application-level learning performance. From the proposed design and evaluation method, a reconfigurable SNN design based on fast-convergence CORDIC is designed to achieve high classification accuracy on MNIST, fast on-line learning and good energy efficiency. By utilizing SNN’s fault tolerance and time-division-multiplexing (TDM) strategy, the reconfigurable SNN design employs 8-bit fast-convergence CORDIC and TDM-based hardware accelerator for high efficiency. FPGA implementation results confirm that the proposed fast-convergence CORDIC SNN design outperforms the state-of-the-art CORDIC method by 38.5%−45.3% in terms of learning speed and energy efficiency, with the STDP learning of 30.2 ns/SOP, energy efficiency of 176.6 pJ/SOP, processing speed of 6.1 ms/image, and on-line learning convergence of 21.4 s (time to reach the final accuracy, on average), on MNIST benchmark. Jiajun Wu 0006, Zixuan Peng, Xinglong Ji, Guoyi Yu, Chao Wang 0096 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |