Huaxiang Lu

dblp:66/5136 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-5928-9705ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 1 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 A PulseWidth-IN-PulseWidth-Out Universal Nonlinear Processing Element for Time-Domain In-Memory Computing Systems
abstract
Time-Domain In-Memory Computing (TD-IMC) has emerged as a promising analog computing architecture for edge AI applications. However, the lack of developed hardware operators, especially general nonlinear operators, necessitates frequent cross-domain data transmission in practical TD-IMC systems, significantly reducing energy efficiency. In this work, we propose a PulseWidth-IN-PulseWidth-OUT Universal Nonlinear Processing Element (PIPO-UNPE) to address the challenges of nonlinear processing in analog computing. By implementing an RRAM-based two-layer ReLU network, the PIPO-UNPE performs universal nonlinear operations entirely in the time domain. Algorithmically, we introduce Dynamic Loss-Responsive Subset Enhancement (DLRSE) to boost the performance of this low-cost network in function approximation tasks. From a hardware perspective, we design an RRAM-based pulse-driven programmable current source and a low-latency dispersion comparator-based voltage-to-time converter (VTC) to enhance both the energy efficiency and precision of the PIPO-UNPE. Hybrid simulations reveal that the PIPO-UNPE consumes 912 uW of power while delivering a throughput of $\mathbf{1 0 M}$ NOPS (Nonlinear Operations Per Second). Incorporating the PIPO-UNPE into the TD-IMC accelerator can increase energy efficiency by a factor of 9.5 to 25, keeping the accuracy loss below 0.1%.
Pengcheng Feng, Rongxuan Shen, Huaxiang Lu, Xiaoxin Xu
DAC6
2025 CSIA-UIM: A Universal Ising Machine Based on CIM-friendly Spring-Ising Algorithm
abstract
Ising machines are specialized processors designed to solve combinatorial optimization problems through the physical evolution of the Ising graphs. However, conventional Ising machines are restricted to solving problems with specific graph topologies and are further hindered by the inherent memory wall of the von Neumann architecture. In this work, we propose a Computing-In-Memory-friendly Spring-Ising Algorithm (CSIA) to solve Ising models with arbitrary graph topologies. Building on CSIA, we introduce a universal Ising machine (CSIA-UIM) capable of fully parallel spin updates. The CSIA-UIM adopts a Time-Domain Computing-In-Memory architecture, with key modules including the Generalized Momentum Update Module (GMUM), Generalized Coordinate Update Module (GCUM), and Hardware Inelastic Wall (HIW) working collaboratively in a parallel pipeline fashion. Hybrid simulation results show that CSIA-UIM achieves speed improvements of 468×, 64×, 2.5×, compared to GPU, RRAM-based, and CMOS-based universal Ising machines, respectively, when solving a fully connected Ising model with 1,000 spins.
Zhelong Jiang, Pengcheng Feng, Jinke Yu, Rongxuan Shen, Huaxiang Lu
ISCAS8
2025 A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
abstract
FPGA accelerators for lightweight convolutional neural networks (LWCNNs) have recently attracted significant attention. Most existing LWCNN accelerators focus on single-Computing-Engine (CE) architecture with local optimization. However, these designs typically suffer from high on-chip/off-chip memory overhead and low computational efficiency due to their layer-by-layer dataflow and unified resource mapping mechanisms. To tackle these issues, a novel multi-CE-based accelerator with balanced dataflow is proposed to efficiently accelerate LWCNN through memory-oriented and computing-oriented optimizations. Firstly, a streaming architecture with hybrid CEs is designed to minimize off-chip memory access while maintaining a low cost of on-chip buffer size. Secondly, a balanced dataflow strategy is introduced for streaming architectures to enhance computational efficiency by improving efficient resource mapping and mitigating data congestion. Furthermore, a resource-aware memory and parallelism allocation methodology is proposed, based on a performance model, to achieve better performance and scalability. The proposed accelerator is evaluated on Xilinx ZC706 platform using MobileNetV2 and ShuffleNetV2. Implementation results demonstrate that the proposed accelerator can save up to 68.3% of on-chip memory size with reduced off-chip memory access compared to the reference design. It achieves an impressive performance of up to 2092.4 FPS and a state-of-the-art MAC efficiency of up to 94.58%, while maintaining a high DSP utilization of 95%, thus significantly outperforming current LWCNN accelerators.
Pengcheng Feng, Jixing Li, Rongxuan Shen, Huaxiang Lu
IEEE Trans. Circuits Syst. I Regul. Pap.7
2024 An FPGA-Based High-Throughput Dataflow Accelerator for Lightweight Neural Network
abstract
Lightweight neural networks (LWNNs) have drawn significant attention recently for compact architecture and acceptable accuracy. Despite achieving substantial reductions in computation complexity and model size, increased memory access demands are caused by the extensive use of depthwise separable convolutions (DSCs) and skip-connection blocks (SCBs), which makes it difficult to achieve the anticipated performance. To process LWNNs efficiently, an FPGA-based dataflow accelerator is proposed in this paper. Firstly, a pixel-based streaming strategy is introduced to reduce off-chip memory access while minimizing on-chip memory overhead. Furthermore, an adaptive bandwidth computing engine (CE) is designed to increase computational efficiency in multi-CE architecture. Finally, based on the scalable CE, a dynamic parallelism allocation algorithm is proposed to avoid underutilization of on-chip computing resources. Shuf-fleNetV2 is implemented on Xilinx ZC706 platform, and the results show the proposed accelerator can achieve a state-of-the-art performance of 1771.2 FPS and computational efficiency of 0.64 GOPS/DSP, which is 5.3× of the reference design.
Jixing Li, Zhelong Jiang, Ruixiu Qiao, Huaxiang Lu
ISCAS8
2024 ACQ: Improving generative data-free quantization via attention correction
Jixing Li, Xiaozhou Guo, Benzhe Dai, Guoliang Gong, Wenyu Mao, Huaxiang Lu
Pattern Recognit.8
2024 CAT-DUnet: Enhancing Speech Dereverberation via Feature Fusion and Structural Similarity Loss
abstract
Reverberation significantly degrades speech intelligibility, posing a substantial challenge in speech processing. While deep learning advancements offer promising solutions, current methodologies often overlook the effective integration of low-level and high-level feature representations, causing detrimental effects on overall performance. Simultaneously, prior approaches heavily rely on loss functions grounded in quantitative error metrics, which may not fully capture the perceptual intricacies of speech signals. To address these concerns, we introduce CAT-DUnet, a Unet architecture that integrates channel attention, time-frequency attention, and dilated convolution blocks to enhance feature fusion. We innovatively leverage the structural similarity as the training objective to align more closely with human perception, and investigate the effect of applying various reasonable transformations to spectrograms on the performance of the loss function. Through extensive ablation experiments, we demonstrate the effectiveness of our proposed enhancements. Our model outperforms state-of-the-art models on 6 out of 7 metrics, underscoring its exceptional performance.
Bajian Xiang, Wenyu Mao, Kaijun Tan, Huaxiang Lu
IEEE Signal Process. Lett.4
2023 Lightweight real-time stereo matching algorithm for AI chips
Yi Liu 0109, Xintao Xu, Xiaozhou Guo, Guoliang Gong, Huaxiang Lu
Comput. Commun.6
2023 Novel activation function with pixelwise modeling capacity for lightweight neural network design
abstract
Summary The development of lightweight networks makes neural networks more efficient to be widely applied to various tasks. Considering the deployment of hardware like edge devices and mobile phones, we prioritize lightweight networks. However, their accuracy has always lagged far behind SOTA networks. In this article, we present a simple yet effective activation function, called WReLU, to improve the performance of lightweight networks significantly by adding a residual spatial condition. Moreover, we use a strategy to switch activation functions after determining which convolutional layer to use. We perform experiments on ImageNet 2012 classification dataset in CPU, GPU, and edge devices. Experiments demonstrate that WReLU improves the accuracy of classification significantly. Meanwhile, our strategy balances the effect of additional parameters and multiply accumulate. Our method improves the accuracy of SqueezeNet and SqueezeNext by more than 5% without increasing extensive parameters and computation. For the lightweight network with a large number of parameters, such as MobileNet and ShuffleNet, there is also a significant improvement. Additionally, the inference speed of most lightweight networks using our WReLU strategy is almost the same as the baseline model on different platforms. Our approach not only ensures the practicability of the lightweight network but also improves its performance.
Yi Liu 0109, Xiaozhou Guo, Kaijun Tan, Guoliang Gong, Huaxiang Lu
Concurr. Comput. Pract. Exp.5
2023 Bird-Count: a multi-modality benchmark and system for bird population counting in the wild
Hongchang Wang, Huaxiang Lu, Huimin Guo, Haifang Jian, Chuang Gan 0001, Wu Liu 0005
Multim. Tools Appl.2
2022 Blind separation of noncooperative paired carrier multiple access signals based on improved quantum-inspired evolutionary algorithm and receding horizon optimization
abstract
Abstract The single‐channel blind source separation of paired carrier multiple access (PCMA) signal is a key technology in satellite communications. Due to the high‐order complexity of existing separation methods and the uncertainty of channel parameters, the processing of noncooperative PCMA signals with long memory remains a great challenge. In this article, the blind separation was solved as the joint channel estimation and sequence detection, and a novel blind separation algorithm based on the improved quantum‐inspired evolutionary algorithm (IQEA) and receding horizon optimization was proposed. Considering the practical scenarios, a mixed PCMA signal model with noninteger period sampling was presented in advance. Combined with the proposed signal model, the IQEA was applied to recover the uplink signals from the mixed PCMA signals by its iterative estimation. Meanwhile, to reduce the computational complexity caused by long memory of PCMA signals, the receding horizon optimization was introduced to process the whole signal sequence as a dynamic programming problem. The simulation results verified that the proposed RH‐IQEA outperforms existing state‐of‐the‐art separation algorithms in the blind separation of noncooperative PCMA signals, with higher accuracy, lower‐order complexity and stronger robustness to parameter estimation errors.
Huaxiang Lu
Concurr. Comput. Pract. Exp.4
2020 Deep Matching Network for Handwritten Chinese Character Recognition
Huaxiang Lu
Pattern Recognit.5
2018 Tracking the multi-well surface dynamometer card state for a sucker-rod pump by using a particle filter
abstract
For a non‐linear sucker‐rod pumping system, a surface dynamometer card estimation algorithm based on a particle filter is presented. The dynamometer card is a plot of the polished rod load at various positions of a pump stroke. Since the polished rod load measured by a load sensor is frequently affected by drift problems, a local characteristic correlation method is proposed while building the state‐space model for the pumping unit. The local characteristic correlation method makes the system insensitive to load drift problems. Moreover, the prior data recorded from different wells are used to construct the importance density. To make the k ‐time importance density closer to the real posterior distribution, current measurement information is used. The performance of the proposed algorithm is evaluated on the actual operating data of a Xinjiang oil field containing typical daily production activities that can cause sudden system state changes. The results show that the proposed algorithm can adapt to sudden changes of the underground environment caused by various human factors, and it can provide robust estimation for multi‐well long‐term state tracking.
Guoliang Gong, Rongxuan Shen, Wenyu Mao, Huaxiang Lu
IET Commun.5
2018 Building efficient CNN architecture for offline handwritten Chinese character recognition
Nanjun Teng, Min Jin 0001, Huaxiang Lu
Int. J. Document Anal. Recognit.4
2008 Single-electron tunneling depressing synapse for cellular neural networks
Huaxiang Lu
Neural Comput. Appl.2
2006 Application of ICA in On-Line Verification of the Phase Difference of the Current Sensor
Huaxiang Lu
ICONIP (3)2