EDBT 2026 Demo / reviewers in the wild / expert
Yuxiang Huan
dblp:167/8478
· DBLP profile ↗
17ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-9155-1451ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WaferBRAIN: Whole-Brain Scale Neuromorphic Architecture Based on Wafer-Scale Integration
Yukun Feng, Liangyu Gan, Haoming Chu, Yufan He, Jiaxin Yin, Lirong Zheng 0001, Yuxiang Huan |
ISCA | 9 |
| 2026 | PicoSleepNet: An Ultra Lightweight Sleep Stage Classification by Spike Neural Network Using Single-Channel EEG SignalabstractThis study introduces PicoSleepNet, an ultra-lightweight sleep stage classification method that utilizes a spiking neural network (SNN) with single-channel electroencephalogram (EEG) signals. Traditional methods use multi-bit Nyquist sampling and dense computing, which result in high complexity and power consumption, hindering their deployment on wearable devices. To address these limitations, we propose an innovative pipeline combining single-bit sub-Nyquist level-crossing sampling (LCS) and sparse computing based on SNN. First, LCS adaptively encodes EEG signals into event-driven spike sequences, reducing data volume by 6.98× while preserving essential signal characteristics compared to Nyquist sampling. Second, a sparse recurrent spiking neural network (RSNN) architecture, optimized by the masked backpropagation and sparse regularization (Masked-BPSR) technique, improves performance and reduces computational costs. Third, quantization-aware training (QAT) ensures that the model maintains high accuracy with low-bitwidth quantization, significantly reducing computational power consumption and enabling hardware-friendly deployment. Compared with current state-of-the-art sleep staging approaches, PicoSleepNet achieves competitive performance on three public datasets (Sleep-EDF-20, Sleep-EDF-78, and ISRUC-Sleep) with accuracies of 83.5%, 77.9%, 79.4% and macro-F1 scores of 75.2%, 68.1%, 77.2%, respectively. Meanwhile, by leveraging the computational sparsity design of RSNN and the joint optimization of Masked-BPSR and QAT, PicoSleepNet achieves an ultra-lightweight model with only 14.0-25.8 K parameters (reduced by nearly 2×) and 681.4-842.0 K operations (reduced by 27×), reducing computational power consumption by 1480×. This approach demonstrates the feasibility of deploying ultra-lightweight sleep staging systems in wearable devices and neuromorphic hardware, paving the way for broader applications in real-time health monitoring. Shengnan Liu, Haoming Chu, Yukun Feng, Yulong Yan, Yuxiang Huan |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | MCU-Enabled Epileptic Seizure Detection System With Compressed LearningabstractEpilepsy is one of the most common neurological disorder diseases all over the world, which gives patients a huge burden in seizure-related disabilities. For epileptic seizure detection, encephalography (EEG) is a commonly used clinical approach. Recently, several Internet of Things (IoT)-based wearable monitoring systems using machine learning (ML) approaches have been proposed to assist real-time detection of epileptic seizure attack scenarios. Among these approaches, convolutional neural networks (CNNs) provide superior accuracy, at the expense of high computational complexity that is not friendly to resource-constrained wearable devices. In this work, we propose a compressed learning (CL)-based epileptic seizure detection system which enables implementing CNN on microcontroller units (MCUs). The proposed CL approach combines the measurement matrix of compressed sensing (CS) and a 1-D CNN. As a result, the input data size and parameter number of CNN can be significantly reduced while eliminating the complex reconstruction process of the traditional CS approach. Evaluated on the Bonn dataset, our proposed system ensures a 96.44% accuracy under a 0.1 compressed ratio (CR) corresponding to a$33.5\times $multiply-accumulate operations (MACs) reduction and a$21.9\times $decrease in model size compared with the baseline. Mapping such a model on an MSP432 MCU, the memory requirement is 50.13 kB and the power consumption is 13.4727 mW at 40 MHz frequency, corresponding to an energy consumption of$269.4~\mu \text{J}$/Classification, with a classification latency of 21.07 ms. Liyu Qian, Yuxiang Huan, Yaojie Sun, Lirong Zheng 0001, Zhuo Zou |
IEEE Internet Things J. | 4 |
| 2023 | A Low-Power Hybrid-Precision Neuromorphic Processor With INT8 Inference and INT16 Online Learning in 40-nm CMOSabstractIn this work, we present a neuromorphic processor for artificial intelligence of things (AIoT) applications featuring low-power consumption, a small footprint, STDP-based online learning, and the ability to adapt to multiple applications. Hybrid precision, i.e., INT8 for inference and INT16 for training, is suggested to achieve balanced accuracy and energy efficiency. A precision-configurable leaky integrate-and-fire(LIF) neuron unit and a unified memory architecture are designed to maximize datapath reuse. A dynamic pruning technique is proposed to exploit the temporal sparsity, yielding synaptic operations reduction by 3.68x in training and 1.63x in inference, respectively. The design is implemented and fabricated in a 40-nm CMOS process, with a core area of 0.87 mm2. It is measured to consume a minimal power of$680~\mu \text{W}$at 70 MHz under a 0.75 V power supply, corresponding to 9.9 pJ per synaptic operation. Evaluated with typical spatial, temporal, and spatiotemporal datasets (MNIST, MIT-BIH, and N-MNIST), the proposed design achieve energy efficiency comparable to the best-in-class solutions with handcrafted training and customized ASICs, while demonstrating improved versatility across multiple applications with balanced accuracy, power consumption, and model adaptability. Congyang Liu, Ziyi Yang 0014, Zikai Zhu, Haoming Chu, Yuxiang Huan, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | ASLog: An Area-Efficient CNN Accelerator for Per-Channel Logarithmic Post-Training QuantizationabstractPost-training quantization (PTQ) has been proven an efficient model compression technique for Convolution Neural Networks (CNNs), without re-training or access to labeled datasets. However, it remains challenging for a CNN accelerator to fulfill the efficiency potential of PTQ methods. A large number of PTQ techniques blindly pursue high theoretic compression effect and accuracy, ignoring their impact on the actual hardware implementation, which causes more hardware overhead than benefit. This paper introduces ASLog, a PTQ-friendly CNN accelerator that explores four key designs in an algorithm-hardware co-optimizing manner: the first practical 4-bit logarithmic PTQ pipeline SLogII, the multiplier-free arithmetic element (AE) design, the energy-efficient bias correction element (BCE) design, and the per-channel quantization friendly (PCF) architecture and dataflow. The proposed SLogII PTQ pipeline can push the limit of logarithmic PTQ to 4-bit with40% lower in power and area consumption compared with a common 8-bit multiplier. The BCE and PCF design proposed in this paper are the first to consider the hardware impact of the widely-used per-channel quantization and bias correction technique, enabling an efficient PTQ-friendly implementation with a small hardware overhead. The ASLog is validated in a UMC 40-nm process, with 12.2 TOPS/W energy efficiency and 0.80 mm2 core area. The ASLog can achieve 336.3 GOPS/mm2 area efficiency and >500 OPs/Byte operational intensity, which map to over$1.85\times $and$1.12\times $improvement compared with the previous related works. Jiawei Xu 0002, Jiangshan Fan, Baolin Nan, Chen Ding 0010, Lirong Zheng 0001, Zhuo Zou, Yuxiang Huan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | A Domain-Specific Accelerator for Ultralow Latency Market Data Distribution SystemabstractUltralow latency parsing of financial data is gaining significance in the high-frequency trading of the security exchange market. Hardware-aided systems exhibit superior improvement of latency over traditional software solutions, but flexibility may suffer when processing the financial protocol. This article presents a domain-specific accelerator for the market data distribution system, which integrates a financial information exchange adapted for streaming (FAST) decoder, a 10-Gbps network interface, and a high-speed PCIe host interface into a single field programmable gate array (FPGA) acceleration card. The proposed FAST decoder adopts the finite state machines-coordinated sequence mapping table to achieve run-time reconfigurability over fine-grained FPGA programming, and 16 fields can be decoded simultaneously in a pipelined manner, resulting in a decoding latency of only 33 ns. Evaluated on the Xilinx Alveo U200 acceleration card, this work improves the latency of decoding a FAST message by 26%–72% than state-of-the-art FPGA designs, and outperforms the software baseline by$>27\times$in terms of latency covering both decoding and communication. Yuxiang Huan, Chen Ding 0010, Yulong Yan, Jianjun Cui, Jiachen Wang 0008, Chuhuang Cai, Zhuo Zou, Lirong Zheng 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | A Hybrid-Mode On-Chip Router for the Large-Scale FPGA-Based Neuromorphic PlatformabstractLarge-scale neuromorphic computing requires the multi-chip network to provide high computing power. Efficient routing schemes and on-chip router design are necessary for handling various inter-chip transmission patterns. In this paper, we propose a hybrid-mode on-chip router that supports both multicast and unicast routing for the large-scale neuromorphic simulation. Two routing schemes, namely Cache-like Spike Weight Indexing and General Unicast Flow Control, are proposed to accommodate the chip-to-chip transmission of spike and non-spike data. This work is evaluated on a neuromorphic platform built with an$8\times 8$FPGA chips array. Running a simulation of 1M neurons at 200MHz, the proposed router achieves a processing latency of 25ns and a chip-to-chip latency of 287ns. Working in the unicast mode, the router can synchronize status flags of all chips within$5 ~\mu \text{s}$. Moreover, it reduces the peak spike traffic by 25.65% with the help of Load-aware Multicast Routing, compared with other multicast routing strategies. Chen Ding 0010, Yuxiang Huan, Yulong Yan, Fanxi Yang, Lizheng Liu, Meigen Shen, Zhuo Zou, Lirong Zheng 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Edge-Based Collaborative Training System for Artificial Intelligence-of-ThingsabstractThe descending of intelligence from the cloud to the heterogeneous and low-power edge in the Artificial Intelligence-of-Things prevents uploading user-sensitive information to the cloud. It brings an urgent demand for deploying training tasks collaboratively in industrial scenarios to manage data locally. This article proposes an edge-based collaborative training system for the smart factory which harnesses the intelligence of edge devices by balancing the computational and communicational resources and improving system dependability. Two typical scenarios of parts recognition and defect inspection are evaluated as a case study with our system. The feasibility and dependability of the presented system are verified with a platform composed of eight high-performance (Nvidia Jetson Nano) and eight low-performance edge devices (Raspberry Pi 4B). The efficiency under tradeoff between computational resource and network condition constraints in a cluster is tested to simulate real-case performance in smart factory scenarios. Our platform reaches the peak performance of 1167 images/s training efficiency on ResNet32 under a 125 MB/s bandwidth. Experimental results demonstrate that the proposed design can collaboratively perform training tasks with optimized efficiency and provide dependable collaborations for system fault detection and cluster extension. Yi Jin 0007, Yulong Yan, Yuxiang Huan, Jiawei Xu 0002, Shancang Li, Prosanta Gope, Zhuo Zou, Lirong Zheng 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | Self-aware distributed deep learning framework for heterogeneous IoT edge devices
Yi Jin 0007, Jiawei Cai, Jiawei Xu 0002, Yuxiang Huan, Yulong Yan, Yongliang Guo, Lirong Zheng 0001, Zhuo Zou |
Future Gener. Comput. Syst. | 4 |
| 2021 | IECA: An In-Execution Configuration CNN Accelerator With 30.55 GOPS/mm² Area EfficiencyabstractIt remains challenging for a Convolutional Neural Network (CNN) accelerator to maintain high hardware utilization and low processing latency with restricted on-chip memory. This paper presents an In-Execution Configuration Accelerator (IECA) that realizes an efficient control scheme, exploring architectural data reuse, unified in-execution controlling, and pipelined latency hiding to minimize configuration overhead out of the computation scope. The proposed IECA achieves row-wise convolution with tiny distributed buffers and reduces the size of total on-chip memory by removing 40% of redundant memory storage with shared delay chains. By exploiting a reconfigurable Sequence Mapping Table (SMT) and Finite State Machine (FSM) control, the chip realizes cycle-accurate Processing Element (PE) control, automatic loop tiling and latency hiding without extra time slots for pre-configuration. Evaluated on AlexNet and VGG-16, the IECA retains over 97.3% PE utilization and over 95.6% memory access time hiding on average. The chip is designed and fabricated in a UMC 55-nm process running at a frequency of 250 MHz and achieves an area efficiency of 30.55 GOPS/mm2and 0.244 GOPS/KGE (kilo-gate-equivalent), which makes an over$2.0\times $and$2.1\times $improvement, respectively, compared with that of previous related works. Implementation of the IEC control scheme uses only a 0.55% area of the 2.75 mm2core. Boming Huang, Yuxiang Huan, Haoming Chu, Jiawei Xu 0002, Lizheng Liu, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | An Autonomous Error-Tolerant Architecture Featuring Self-reparation for Convolutional Neural NetworksabstractConvolutional neural networks are widely used in artificial intelligence and Internet of Things area. As the scale of convolutional neural network expands, more and more processing units are provided for it. The systems are easy prone to error, and any computing problems in any layer of the network will lead to wrong output results. Traditional multimode redundancy methods make the systems more complex, and increase power consumption. This paper proposes an autonomous error-tolerant architecture for convolutional neural networks. Taking the LeNet-5 as an example, the network layers of CNN are mapped on the AET architecture, an error-tolerant synapse is designed to discover the errors, an active evolution scheme is designed to handle unrecoverable errors and implement network reconfiguration. This design is implemented on FPGA, and the experimental results show that this architecture can realize effective error tolerance for convolutional neural network and has fast error recovery ability under the premise of ensuring the same recognition accuracy. Lizheng Liu, Yuxiang Huan, Zhuo Zou, Xiaoming Hu 0001, Lirong Zheng 0001 |
VTC Spring | 2 |
| 2020 | A Smart Dental Health-IoT Platform Based on Intelligent Hardware, Deep Learning, and Mobile TerminalabstractThe dental disease is a common disease for a human. Screening and visual diagnosis that are currently performed in clinics possibly cost a lot in various manners. Along with the progress of the Internet of Things (IoT) and artificial intelligence, the internet-based intelligent system have shown great potential in applying home-based healthcare. Therefore, a smart dental health-IoT system based on intelligent hardware, deep learning, and mobile terminal is proposed in this paper, aiming at exploring the feasibility of its application on in-home dental healthcare. Moreover, a smart dental device is designed and developed in this study to perform the image acquisition of teeth. Based on the data set of 12 600 clinical images collected by the proposed device from 10 private dental clinics, an automatic diagnosis model trained by MASK R-CNN is developed for the detection and classification of 7 different dental diseases including decayed tooth, dental plaque, uorosis, and periodontal disease, with the diagnosis accuracy of them reaching up to 90%, along with high sensitivity and high specificity. Following the one-month test in ten clinics, compared with that last month when the platform was not used, the mean diagnosis time reduces by 37.5% for each patient, helping explain the increase in the number of treated patients by 18.4%. Furthermore, application software (APPs) on mobile terminal for client side and for dentist side are implemented to provide service of pre-examination, consultation, appointment, and evaluation. Lizheng Liu, Jiawei Xu 0002, Yuxiang Huan, Zhuo Zou, Shih-Ching Yeh, Lirong Zheng 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2018 | A 3D Tiled Low Power Accelerator for Convolutional Neural NetworkabstractIt remains a challenge to run Deep Learning in devices with stringent power budget in the Internet-of-Things. This paper presents a low-power accelerator for processing Convolutional Neural Networks on the embedded devices. The power reduction is realized by exploring data reuse in three different aspects, with regards to convolution, filter and input features. A systolic-like data flow is proposed and applied to rows of Processing Elements (PEs), which facilitate reusing the data during convolution. Reuse of input features and filters is achieved by arranging the PE array in a 3D tiled architecture, whose dimension is 3 × 14 × 4. Local storage within PEs is therefore reduced and only cost 17.75 kB, which is 20% of the state-of-the-art. With dedicated delay chains in each PE, this accelerator is reconfigurable to suit various parameter settings of convolutional layers. Evaluated in UMC 65 nm low leakage process, the accelerator can reach a peak performance of 84 GOPS and consume only 136 mW at 250 Mhz. Yuxiang Huan, Jiawei Xu 0002, Lirong Zheng 0001, Hannu Tenhunen, Zhuo Zou |
ISCAS | 1 |
| 2018 | TMR Group Coding Method for Optimized SEU and MBU Tolerant Memory DesignabstractThis work proposes a fault tolerant memory design using the method of Triple Module Redundancy (TMR) group coding to tolerant the Single-Event Upset (SEU) and Multi-Bit Upset (MBU) influence on memory devices in space environment. The group coding method uses different models to partition and code each word line in memory with Hamming code to achieve best performance. TMR group coding method further increases the capability of self-correction for the errors occurred in parity bits. The evaluation results show that the suggested approach can obtain improved correctness for the memory output with optimized tradeoff between reliability and cost. At 5% error rate, the probability of correct output reaches 70.78% with small cost increment. To achieve 90% reliability, the accuracy improvement is 31.9% compared to TMR with 9% increased area. This solution proposed is evaluated on the memory rich micro-coded processor, but can be further extended to other memory-based processors that need high reliability for the SEU and MBU influence in aerospace applications. Yi Jin 0007, Yuxiang Huan, Haoming Chu, Zhuo Zou, Lirong Zheng 0001 |
ISCAS | 2 |
| 2018 | A Design of Autonomous Error-Tolerant Architectures for Massively Parallel Computing
Lizheng Liu, Yi Jin 0007, Yi Liu 0027, Yuxiang Huan, Zhuo Zou, Lirong Zheng 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2017 | Smart energy efficient gateway for Internet of mobile thingsabstractInternet of Things (IoT) is a fast developing vision in which physical quantities are digitized, processed and analyzed. Internet of Mobile Things (IoMT) as one of new domains of IoT, due to mobility, requires a more demanding and rigorous solution in many aspects, especially in terms of energy efficiency. We propose a solution consisting of energy efficient and fast hardware platform for building IoMT Fog layer facilities. Experimental results are presented to prove superiority of the proposed hardware in several aspects to popular general purpose platforms. Igor Tcarenko, Yuxiang Huan, David Juhasz, Amir-Mohammad Rahmani, Zhuo Zou, Tomi Westerlund, Pasi Liljeberg, Lirong Zheng 0001, Hannu Tenhunen |
CCNC | 2 |
| 2016 | A 101.4 GOPS/W reconfigurable and scalable control-centric embedded processor for domain-specific applicationsabstractIncreasing the energy efficiency and performance while providing the customizability and scalability is vital for embedded processors adapting to domain-specific applications such as Internet of Things. In this paper, we proposed a reconfigurable and scalable control-centric architecture, and implemented the design consisting of two cores and an on-chip multi-mode router in 65 nm technology. The reconfigurability is enabled by the restructurable sequence mapping table (SMT) thus the reorganizable functional units. Owing to the integration of the multi-mode router, on-chip or inter-chip network for multi-/many-core computing can be composed for performance extension on demand even in the post-fabrication stage. Control-centric design simplifies the control logic, shrinks the non-functional units and orchestrates the operations to increase the hard are utilization and reduce the excessive data movement for high energy efficiency. As a result, the processor can both conduct general-purpose processing with 29% smaller code size and application-specific processing with over 10 times performance improvement when implementing AES by SMT. The dual-core processor consumes 19.7 μW/MHz with die size of 3.5 mm2. The achieved energy efficiency is 101.4GOPS/W. Zhuo Zou, Zhonghai Lu, Lirong Zheng 0001, Yuxiang Huan, Stefan Blixt |
ISCAS | 5 |