VLDB 2026 Research / reviewers in the wild / expert
Pingdan Xiao
dblp:348/1247
· DBLP profile ↗
19ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-4134-8551ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | In-Sensor Parallel Computing Circuit for High-Speed Image Enhancement Based on Proposed All-Channel Mean Spray Retinex AlgorithmabstractImage enhancement serves as a fundamental step in many advanced computer graphics tasks in the field of Internet of Things. However, existing methods often struggle to meet the demands of high-speed, high-resolution applications due to their high computational complexity, which heavily consumes the processor resources available on edge devices and limits their ability to support subsequent tasks. To address these challenges, this paper proposed the All-Channel Mean Spray Retinex (ACMSR) algorithm for the first time, which offers a more efficient and rapid solution for image enhancement. Furthermore, an in-situ computing ACMSR circuit is designed to eliminate the reliance on processor resources of edge devices. Compared to traditional methods, this circuit not only integrates sensing, storage, and computation in-sensor to execute the ACMSR algorithm, but also fully leverages the analog parallel computing capabilities of the memristor-crossbar-array circuit to perform architecture-level acceleration. Moreover, the proposed method achieves a processing speed of 69.3 FPS for 2K images. The circuit delivers output accuracy exceeding 98% under most types of interference. Haoyou Jiang, Tao Li 0056, Pingdan Xiao, Sichun Du, Qinghui Hong |
IEEE Internet Things J. | 3 |
| 2026 | Design and application of general circuits for solving matrix equation ∑ i = 0 N A i X B i = C
Sichun Du, Bingqian Zhang, Pingdan Xiao, Zhengmiao Wei, Qinghui Hong |
Inf. Sci. | 3 |
| 2026 | An Analog Matrix Computing Scheme for QC-MDPC McEliece Cryptosystem Based on Memristive Array
Pingdan Xiao, Bingqian Zhang, Sichun Du, Qinghui Hong |
IEEE Trans. Computers | 1 |
| 2026 | Memristive Neural Network Circuit Implementation of Model Predictive Control for Trajectory TrackingabstractModel Predictive Control (MPC), a receding-horizon optimal control strategy, predicts system dynamics and optimizes control actions to satisfy performance and constraint requirements, making it widely adopted in control engineering. However, contemporary computing platforms struggle to meet the real-time and energy-efficient demands of MPC’s computationally intensive matrix operations, stemming from high data movement overhead, extensive circuit resource utilization, and frequent data conversions inherent in physical system interfaces. These challenges collectively impose significant latency and power penalties, particularly critical as systems grow in complexity and scale within the big-data era. This article introduces a Zeroing Neural Network (ZNN)-based memristive neural network circuit that directly converges the MPC error function to zero in one step. Theoretical analysis and simulations validate the closed-loop circuit’s stability. For a 32-step prediction horizon, evaluations show that the control output from the proposed circuit matches the ideal digital MPC solution with 96.0% accuracy. The circuit also executes at least an order of magnitude faster and consumes less energy than traditional MPC solvers. Additionally, the circuit successfully accelerates the proposed trajectory tracking algorithm, achieving 98.0% accuracy compared with the theoretical result and 318.2× improvement in computation time compared to CPU. Pingdan Xiao, Yiliu Gu, Haoyou Jiang, Zhen Huan, Sichun Du, Qinghui Hong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | EDCSSM: Edge Detection With Convolutional State Space ModelabstractEdge detection in images is the foundation of many complex tasks in computer graphics. Due to the feature loss caused by multi-layer convolution and pooling architectures, learning-based edge detection models often produce thick edges and struggle to detect the edges of small objects in images. Inspired by state space models, this paper presents an edge detection algorithm which effectively addresses the aforementioned issues. The presented algorithm obtains state space variables of the image from dual-input channels with minimal down-sampling processes and utilizes these state variables for real-time learning and memorization of image patches. To further enhance the processing speed of the algorithm, we have designed parallel computing circuits for the most computationally intensive parts of presented algorithm, significantly improving computational speed and efficiency. Simulation results demonstrate that the proposed algorithm achieves precise thin edge localization and exhibits noise suppression capabilities across various types of images. The parallel computing circuit achieves an output accuracy of over 92.13% under most interference conditions. Accelerated by this circuit, the calculation core of algorithm achieves a processing speed of 30fps on 5K images. Haoyou Jiang, Tao Li 0056, Pingdan Xiao, Sichun Du, Qinghui Hong |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | Analog Solver Design of LU Decomposition Algorithm for Accelerating Public Color WatermarkingabstractAs a key kernel of matrix operation, lower–upper (LU) decomposition plays an important role in linear algebra and various engineering applications, such as deep learning and watermarking. Due to the complexity of large-scale matrix factorization, existing work mainly uses digital circuits and traditional computing architecture to realize LU decomposition, which inevitably faces the constraints of hardware resources and latency. Aiming at the above problem, this article first proposes the analog solver based on memristive arrays to realize LU decomposition of a nonsingular matrix in any dimension, with the advantages of high speed and parallel computing from analog circuits. The closed-loop circuit constructed therein endows it with excellent robustness and convergence characteristics during LU decomposition, and it can be performed in parallel. The simulation results indicate that the circuit designed for even 32nd-order LU matrix decomposition can achieve 95.90% accuracy. The accelerator achieved about$7\times$speed up on latency and$48.4\times$increase in TOPS performance compared to Alveo U50 FPGA, exhibiting superior performance in robustness and programming errors as well. Moreover, the proposed analog circuit serves to accelerate the embedding and extraction of digital watermarking based on LU decomposition, which enhances the invisibility of the watermark, and this application showcases the advantage of high accuracy and low energy consumption. Pingdan Xiao, Meimei Ma, Haoyou Jiang, Zhen Huan, Sichun Du, Qinghui Hong |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | ACIM-QMM: Efficient Analog Computing-in-Memory Accelerator for QC-MDPC McEliece CryptosystemabstractQuasi-cyclic moderate density parity-check McEliece (QMM) cryptosystem is designed to mitigate the security threat posed by quantum computers, and is considered to be a promising candidate for post-quantum cryptography (PQC). However, the growing requirement of data encryption pose severe challenges for QMM implementation in terms of latency and hardware overhead. In this work, we firstly propose ACIM-QMM, an analog computing-in-memory (CIM) accelerator design for QMM cryptosystem. The use of analog circuits and CIM enables the design to efficiently generate key and encrypt ciphertext while breaking the performance bottleneck constrained by digital computing paradigm in PQC. In the experiment, ACIM-QMM can work in low relative error, and it can achieve $31.4 \times \sim 288.1 \times$ speedup compared with SOTA hardware of QMM cryptosystem. Furthermore, the results indicate that ACIM-QMM can achieve a maximum of $3.12 \times$ area efficiency and $20.32 \times$ energy efficiency compared to other PQC hardware for 256-bit security. Pingdan Xiao, Zhengmiao Wei, Sichun Du, Wanli Chang 0001, Qinghui Hong |
DAC | 1 |
| 2025 | Universal Programmable Transfer Function Modeling Circuit Based on Memristors for PID Simulation ApplicationabstractControl systems play a crucial role in Internet of Things applications. However, in the face of increasingly complex physical environments and the sharp increase in data volume in edge computing scenarios, the transfer functions traditionally modeled through software methods tend to be offline processing. This limitation impedes the ability of Internet of Things device terminals to achieve real-time system monitoring and data analysis. In response to the challenge, this paper introduces a novel method for circuit modeling, which enables online simulation of any transfer function for the first time. Utilizing programmable memristors enables effective real-time control of system. This study validates the feasibility of the circuit modeling method by employing five engineering examples. Compared with MATLAB software modeling, the method in processing speed is at least 40 times faster and more prominent in the higher-order transfer function modeling, with a accuracy of 95.31%. Finally, this method is used to simulate an automotive cruise control system based on proportional-integral-derivative algorithm, which successfully realizes the real-time control and analysis of the dynamic behavior of the system, and also provides a new idea and technical path for the efficient and accurate dynamic management of Internet of Things systems in the future. Qinghui Hong, Jiping Kang, Pingdan Xiao, Zhengmiao Wei, Sichun Du |
IEEE Internet Things J. | 3 |
| 2025 | A general analog solver of linear and quadratic programming in one step
Sichun Du, Pingdan Xiao, Zhengmiao Wei, Qinghui Hong |
Neural Networks | 3 |
| 2025 | A Riccati Matrix Equation Solver Design Based Neurodynamics Method and Its ApplicationabstractRiccati matrix equation (RME), a critical nonlinear matrix equation in autonomous driving and deep learning. However, memory-compute separation in traditional solving systems leads to latency and inefficiency when solving nonlinear equations, particularly under real-time requirements. Existing hardware lacks dedicated accelerators for RME, no specialized solvers quick addressing its nonlinear complexity. To address this issue, we propose a novel RME solver based on a memristive array, which leverages the parallel and fast computing advantages of analog circuits to quickly solve any order RME. Inspired by Neurodynamics for non-linear matrix equations, we first introduce an innovative Neurodynamics-based RME solving algorithm specifically designed from an analog circuit perspective. Based on this algorithm, we constructed a pioneering closed-loop analog circuit solver, overcoming bottlenecks in using circuits for such nonlinear matrix equations. Our evaluation demonstrates that the proposed solver achieves over 90% accuracy for a 128th-order parameter Riccati matrix equation. Compared to traditional digital processors, the solver offers significant energy efficiency advantages and is three orders of magnitude faster than CPU. Additionally, the solver successfully accelerates our proposed dung beetle optimizer-linear quadratic control algorithm for vehicle suspension control, achieving high precision while significantly reducing time and energy consumption compared to CPU and GPU. Pingdan Xiao, Junjie Fang, Zhengmiao Wei, Sichun Du, Shiping Wen 0001, Qinghui Hong |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | A Parallel Computing Scheme Utilizing Memristor Crossbars for Fast Corner Detection and Rotation Invariance in the ORB AlgorithmabstractThe Oriented FAST and Rotated BRIEF (ORB) algorithm plays a crucial role in rapidly extracting image keypoints. However, in the domain of high-frame-rate real-time applications, the algorithm faces challenges of the speed and computational efficiency with the increase in both the size and quantity of images. To address this issue, an ORB algorithm accelerator based on a computing-in-memory (CIM) circuit is firstly proposed in this paper, which replaces the iterative calculations in traditional methods with one-step parallel analog computation. The proposed accelerator improves algorithm computational efficiency through CIM technology and enhances algorithm speed through parallel computation. Simulation demonstrate that the proposed method exhibits an average processing speed 22$\boldsymbol{\times}$faster than traditional methods and obtains more uniform corners distribution in large-scale images. Qinghui Hong, Haoyou Jiang, Pingdan Xiao, Sichun Du, Tao Li 0056 |
IEEE Trans. Computers | 3 |
| 2025 | Analog Matrix Inversion Circuit Design for Solving Tridiagonal Linear Systems: A Compact and Decoupled Approach
Sichun Du, Zhengmiao Wei, Pingdan Xiao, Shiping Wen 0001, Qinghui Hong |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | CIM-KF: Efficient Computing-in-memory Circuits for Full-Process Execution of Kalman Filter AlgorithmabstractKalman Filter (KF) algorithm, which can solve the state estimation problem of multi-variable and complex dynamical system, plays a pivotal role in a multitude of engineering scenarios. However, the traditional digital computing architecture represented by the von-Neumann architecture are currently confronted with high overhead challenges in terms of latency, energy, and area when executing KF algorithm. Aiming at the problem, we propose CIM-KF, the first Computing-in-memory (CIM) circuits for KF algorithm. CIM-KF can efficiently carry out the entire process of KF algorithm by capitalizing on the large-scale parallel computation inherent in the CIM architecture. The evaluation shows that CIM-KF has average 96.15% accuracy for 32-th order parameter matrix in KF algorithm, and can be 13.89 × ∼ 55.21 × faster than existing ASIC and FPGA implementations of KF algorithm. The results also demonstrate that CIM-KF can deliver up to 3.52 × energy efficiency improvement and 54.79 × area efficiency improvement, and reduce over 90% latency overhead, compared with the SOTA counterparts. Furthermore, we propose a novel design of ReRAM-CIM architecture, underpinned by the CIM-KF, aimed at precise and rapid calculation of the state of capacity and terminal voltage in Battery Management Systems. By evaluation, our proposed architecture boasts the estimation accuracy within 2% error for the individual battery cell of state of capacity and terminal voltage. With the strengths of the parallel computing intrinsic to CIM architecture and the speed of analog circuit, our architecture facilitates speed up to 70.6 × in accurate estimation when bench-marked against the SOTA work. Pingdan Xiao, Qinghui Hong, Sichun Du, Jiliang Zhang 0002 |
ICPP | 1 |
| 2024 | Memristive neural network circuit design based on locally competitive algorithm for sparse coding application
Qinghui Hong, Pingdan Xiao, Ruijia Fan, Sichun Du |
Neurocomputing | 2 |
| 2024 | Design of Optoelectronic In-Sensor Computing Circuit Based on Memristive Crossbar Array for In Situ Edge ExtractionabstractThe rapid development of artificial intelligence has brought a huge amount of data, and the traditional image processing architecture that separates sensing, storage and computation will face the problems of high power consumption and processing latency. Focusing on these problems, this paper proposed a design scheme of memristor-based optoelectronic sensing circuit, which can integrate image perception, storage, and processing into one entity. Without large-scale data transmission and conversion, the corresponding energy consumption can be avoided effectively. Firstly, an optoelectronic sensing circuit based on memristive crossbar array is proposed, which realizes the acquisition and in situ storage of image information by embedding the photoelectric converters into the memristive array. On this basis, the corresponding peripheral circuit is designed to accomplish the in situ edge feature extraction for the stored image. The extraction process is large-scale parallel computing in the analog domain, and the speed is significantly improved compared with the traditional solution. Moreover, the extraction accuracy of the circuit can reach more than 99%, and it also can withstand a certain degree of programming error and has strong robustness. Jiliang Zhang 0002, Xinjie Li 0005, Pingdan Xiao, Zhengmiao Wei, Qinghui Hong |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Analog In-memory Circuit Design of Polynomial Multiplication for Lattice Cipher Acceleration ApplicationabstractAs the core operation of lattice cipher, large-scale polynomial multiplication is the biggest computational bottleneck in its realization process. How to quickly calculate polynomial multiplication under resource constraints has become an urgent problem to be solved in the hardware implementation of lattice ciphers. Therefore, an analog in-memory circuit for fast polynomial multiplication calculation is proposed. First, an in-memory computing circuit for Discrete Fourier Transform and Inverse Discrete Fourier Transform based on memristor array is designed. On this basis, a fully analog circuit that can realize polynomial multiplication in one step is designed. Compared with traditional hardware implementation, the in-memory calculation method used in this article decreases the calculation time of polynomial multiplication to the microsecond level, which greatly improves the speed of lattice cipher encryption and decryption. For the specific examples in this article, PSPICE simulation shows that the average accuracy of the calculation result is above 99.90%. Sichun Du, Jun Li 0118, Pingdan Xiao, Qinghui Hong, Jiliang Zhang 0002 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2024 | Analog-in-Memory Accelerator Design Based on Memristive Arrays for Opposite Directional Interference Alignment AlgorithmabstractInterference alignment can overcome the shortcomings in traditional interference management. How to quickly and efficiently eliminate interference by using interference alignment is an important question. Aiming at this problem, we propose an analog in-memory circuit based on memristors for accelerating the opposite directional interference alignment algorithm, which is achieved by solving complex-valued matrix equation and multiple complex-valued matrix multiplication. The circuits can adapt to any numbers of antennas condition and adjust the memconductance to map with the channel matrix in communication systems. The evaluation shows that the circuits not only have 99% high accuracy in executing a$2\times 2$channel state but also have good robustness against some nonideal factors from wireless communication that the corresponding accuracy can exceed 95% under the 10% noise impact. Moreover, the circuits accelerate the algorithm which is three orders of magnitude faster than software. Pingdan Xiao, Qinghui Hong, Sichun Du, Jiliang Zhang 0002 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Drift speed adaptive memristor model
Ya Li 0008, Lijun Xie, Pingdan Xiao, Ciyan Zheng, Qinghui Hong |
Neural Comput. Appl. | 3 |
| 2023 | Programmable In-memory Computing Circuit of Fast Hartley TransformabstractDiscrete Hartley transform is a core component of digital signal processing because of its advantages of fast computing speed and less power consumption. Traditional FPGA-based implementation methods have the disadvantage of high latency, which cannot meet the needs of energy-efficient computing in the Internet of Things era. Therefore, A programmable analog memory computing circuit is proposed to accelerate FHT and IFHT calculations for large-scale one-step matrix computation. By adjusting the weight of memristor, different scales of FHT calculation can be achieved. PSPICE simulation results show that the average accuracy of the proposed circuit can reach 99.9%, and the speed can also reach the level of 0.1 μs. The robustness analysis shows that the circuit can tolerate a certain degree of programming error and resistance tolerance. The designed analog circuit is applied to image compression processing, and the image compression accuracy can reach 99.9%. Qinghui Hong, Richeng Huang, Pingdan Xiao, Jun Li 0118, Jingru Sun, Jiliang Zhang 0002 |
ACM Trans. Design Autom. Electr. Syst. | 3 |