VLDB 2026 Research / reviewers in the wild / expert
Qinyu Chen
dblp:91/5007
· DBLP profile ↗
33ranked-venue papers
8as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 7 first-author · 20 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsabstractXin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen, Bingxuan Wang, Zhenda Xie, Kezhao Huang, Xingkai Yu, Zhewen Hao, Han Zhang, Yu-Kun Li, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xin Cheng 0002, Wangding Zeng, Damai Dai, Qinyu Chen, Bingxuan Wang, Zhenda Xie, Kezhao Huang, Xingkai Yu, Zhewen Hao, Huishuai Zhang, Dongyan Zhao 0001, Wenfeng Liang |
ACL (1) | 4 |
| 2026 | JaneEye: A 12-nm 2K-FPS 18.9-μJ/Frame Event-based Eye Tracking AcceleratorabstractEye tracking has become a key technology for gaze-based interactions in Extended Reality (XR). However, conventional frame-based eye-tracking systems often fall short of XR’s stringent requirements for high accuracy, low latency, and energy efficiency. Event cameras present a compelling alternative, offering ultra-high temporal resolution and low power consumption. In this paper, we present JaneEye, an energy-efficient event-based eye-tracking hardware accelerator designed specifically for wearable devices, leveraging sparse, high-temporal-resolution event data. We introduce an ultra-lightweight neural network architecture featuring a novel ConvJANET layer, which simplifies the traditional ConvLSTM by retaining only the forget gate, thereby halving computational complexity without sacrificing temporal modeling capability. Our proposed model achieves high accuracy with a pixel error of 2.45 on the 3ET+ dataset, using only 17.6 K parameters, with up to 1250 Hz event frame rate. To further enhance hardware efficiency, we employ custom linear approximations of activation functions (HardSigmoid and Hard-Tanh) and fixed-point quantization. Through software-hardware co-design, our 12-nm ASIC implementation operates at 400 MHz, delivering an end-to-end latency of 0.5 ms (equivalent to 2000 Frames Per Second (FPS)) at an energy efficiency of 18.9 μJ/frame. JaneEye sets a new benchmark in low-power, high-performance eye-tracking solutions suitable for integration into next-generation XR wearables. Qinyu Chen, Chang Gao 0002 |
ASP-DAC | 3 |
| 2026 | SHAP-AAD: DeepSHAP-Guided Channel Reduction for EEG Auditory Attention DetectionabstractElectroencephalography (EEG)-based auditory attention detection (AAD) offers a non-invasive way to enhance hearing aids, but conventional methods rely on too many electrodes, limiting wearability and comfort. This paper presents SHAP-AAD, a two-stage framework that combines DeepSHAP-based channel selection with a lightweight temporal convolutional network (TCN) for efficient AAD using fewer channels.DeepSHAP, an explainable AI technique, is applied to a Convolutional Neural Network (CNN) trained on topographic alpha-power maps to rank channel importance, and the top-k EEG channels are used to train a compact TCN. Experiments on the DTU dataset show that using 32 channels yields comparable accuracy to the full 64-channel setup (79.21% vs. 81.06%) on average. In some cases, even 8 channels can deliver satisfactory accuracy. These results demonstrate the effectiveness of SHAP-AAD in reducing complexity while preserving high detection performance. Rayan Salmi, Guorui Lu, Qinyu Chen |
ISCAS | 3 |
| 2026 | A 1.1 μJ/Inference Binary Spiking Neural Network Accelerator for DVS Gesture RecognitionabstractDynamic vision sensors (DVS) are bioinspired sensors that can generate sparse data streams with low latency, low power consumption and high dynamic range. Spiking neural networks (SNNs), which are inspired by biological brains, are event-based models, and therefore they can be used to process the binary data streams produced by such sensors naturally. However, SNN accelerators usually require more memory and longer time for inference, due to the extra time dimension in SNNs. In this paper, an energy efficient binary spiking neural network (BSNN) accelerator for DVS gesture recognition is proposed with algorithm and hardware codesign. We integrate the binary neural network (BNN) training method into SNN to directly train a BSNN model, which significantly reduces memory consumption. A temporal pooling (TP) layer is further proposed to reduce the time steps in SNNs while maintaining competitive accuracy. The proposed BSNN accelerator can achieve high parallelism with high resource utilization, and the sparsity of input spikes is utilized to further reduce power consumption. The proposed BSNN model achieves an accuracy of 95.49% on IBM DVS Gesture dataset. The implementation results show that the BSNN accelerator can achieve 38.2k inference per second with$1.1~\mu $J/inference energy consumption and 216.9 TOPS/W energy efficiency. Congyi Sun, Xusen Zeng, Qiang Tao, Heng Zhang 0025, Qinyu Chen, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2026 | QSNNA: An Energy-Efficient Quaternary Spiking Neural Network Accelerator for Seizure Detection
Heng Zhang 0025, Linfeng Wu, Linxiang Wang, Youbin Luo, Haochuan Pan, Xinyu Wang 0027, Guoqiang He, Qinyu Chen, Li Li 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2025 | WIKIGENBENCH: Exploring Full-length Wikipedia Generation under Real-World ScenarioabstractIt presents significant challenges to generate comprehensive and accurate Wikipedia articles for newly emerging events under real-world scenario. Existing attempts fall short either by focusing only on short snippets or by using metrics that are insufficient to evaluate real-world scenarios. In this paper, we construct WIKIGENBENCH, a new benchmark consisting of 1,320 entries, designed to align with real-world scenarios in both generation and evaluation. For generation, we explore a real-world scenario where structured, full-length Wikipedia articles with citations are generated for new events using input documents from web sources. For evaluation, we integrate systematic metrics and LLM-based metrics to assess the verifiability, organization, and other aspects aligned with real-world scenarios. Based on this benchmark, we conduct extensive experiments using various models within three commonly used frameworks: direct RAG, hierarchical structure-based RAG, and RAG with fine-tuned generation model. Experimental results show that hierarchical-based methods can generate more comprehensive content, while fine-tuned methods achieve better verifiability. However, even the best methods still show a significant gap compared to existing Wikipedia content, indicating that further research is necessary. Jiebin Zhang, Eugene J. Yu, Qinyu Chen, Chenhao Xiong, Han Qian, Mingbo Song, Weimin Xiong, Qun Liu 0001, Sujian Li |
COLING | 3 |
| 2025 | FACET: Fast and Accurate Event-Based Eye Tracking Using Ellipse Modeling for Extended RealityabstractEye tracking is a key technology for gaze-based interactions in Extended Reality (XR), but traditional frame-based systems struggle to meet XR's demands for high accuracy, low latency, and power efficiency. Event cameras offer a promising alternative due to their high temporal resolution and low power consumption. In this paper, we present FACET (Fast and Accurate Event-based Eye Tracking), an end-to-end neural network that directly outputs pupil ellipse parameters from event data, optimized for real-time XR applications. The ellipse output can be directly used in subsequent ellipse-based pupil trackers. We enhance the EV-Eye dataset by expanding annotated data and converting original mask labels to ellipse-based annotations to train the model. Besides, a novel trigonometric loss is adopted to address angle discontinuities and a fast causal event volume event representation method is put forward. On the enhanced EV-Eye test set, FACET achieves an average pupil center error of$\mathbf{0. 2 0}$pixels and an inference time of 0.53 ms, reducing pixel error and inference time by$1.6 \times$and$1.8 \times$compared to the prior art, EV-Eye, with$4.4 \times$and$11.7 \times$less parameters and arithmetic operations. The code is available at https://github.com/DeanJY/FACET. Junyuan Ding, Chang Gao 0002, Qinyu Chen |
ICRA | 5 |
| 2025 | CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel PruningabstractThis paper presents CleanUMamba, a time-domain neural network architecture designed for real-time causal audio denoising directly applied to raw waveforms. CleanUMamba leverages a U-Net encoder-decoder structure, incorporating the Mamba state-space model in the bottleneck layer. By replacing conventional self-attention and LSTM mechanisms with Mamba, our architecture offers superior denoising performance while maintaining a constant memory footprint, enabling streaming operation. To enhance efficiency, we applied structured channel pruning, achieving an 8X reduction in model size without compromising audio quality. Our model demonstrates strong results in the Interspeech 2020 Deep Noise Suppression challenge. Specifically, CleanUMamba achieves a PESQ score of 2.42 and STOI of 95.1% with only 442K parameters and 468M MACs, matching or outperforming larger models in real-time performance. Code will be available at: https://github.com/lab-emi/CleanUMamba Sjoerd Groot, Qinyu Chen, Jan C. van Gemert, Chang Gao 0002 |
ISCAS | 2 |
| 2025 | DPD-NeuralEngine: A 22-nm 6.6-TOPS/W/mm2 Recurrent Neural Network Accelerator for Wideband Power Amplifier Digital Pre-DistortionabstractThe increasing adoption of Deep Neural Network (DNN)-based Digital Pre-distortion (DPD) in modern communication systems necessitates efficient hardware implementations. This paper presents DPD-NeuralEngine, an ultra-fast, tiny-area, and power-efficient DPD accelerator based on a Gated Recurrent Unit (GRU) neural network (NN). Leveraging a co-designed software and hardware approach, our 22 nm CMOS implementation operates at 2 GHz, capable of processing I/Q signals up to 250 MSps. Experimental results demonstrate a throughput of 256.5 GOPS and power efficiency of 1.32 TOPS/W with DPD linearization performance measured in Adjacent Channel Power Ratio (ACPR) of -45.3 dBc and Error Vector Magnitude (EVM) of -39.8 dB. To our knowledge, this work represents the first AI-based DPD application-specific integrated circuit (ASIC) accelerator, achieving a power-area efficiency (PAE) of 6.6 TOPS/W/mm2. Yizhuo Wu, Qinyu Chen, Leo C. N. de Vreede, Chang Gao 0002 |
ISCAS | 4 |
| 2025 | SlimSeiz: Efficient Channel-Adaptive Seizure Prediction Using a Mamba-Enhanced NetworkabstractEpileptic seizures cause abnormal brain activity, and their unpredictability can lead to accidents, underscoring the need for long-term seizure prediction. Although seizures can be predicted by analyzing electroencephalogram (EEG) signals, existing methods often require too many channels or larger models, limiting mobile usability. This paper introduces a SlimSeiz framework that utilizes adaptive channel selection with a lightweight neural network model. SlimSeiz operates in two states: the first stage selects the optimal channel set for seizure prediction using machine learning algorithms, and the second stage employs a lightweight neural network based on convolution and Mamba for prediction. On the Children’s Hospital Boston-MIT (CHB-MIT) EEG dataset, SlimSeiz can reduce channels from 22 to 8 while claiming a satisfactory result of 94.8% accuracy, 95.5% sensitivity, and 94.0% specificity with only 21.2 K model parameters, matching or outperforming larger models’ performance. We also validate SlimSeiz on a new EEG dataset, SRH-LEI, collected from Shanghai Renji Hospital, demonstrating its effectiveness across different patients. The code and SRH-LEI dataset are available at https://github.com/guoruilu/SlimSeiz. Guorui Lu, Bingyuan Huang, Chang Gao 0002, Todor Stefanov, Qinyu Chen |
ISCAS | 7 |
| 2025 | HengNet: An Ultra-lightweight Model with Two-level Reuse Algorithm for Seizure Detection and PredictionabstractTraditional models based on electroencephalographic (EEG) signals for seizure monitoring encounter difficulties in simultaneously optimizing accuracy, response latency, and computational load. These challenges hinder their deployment in edge computing environments, where real-time local inference is critical. To address these issues, we introduce a novel network architecture, designated as HengNet. This architecture integrates a Two-level Reuse Algorithm (TRA), which strategically reutilizes outputs from intermediate layers, considerably reducing the average computational load per inference—vital for scenarios requiring frequent inferences. When tested on the CHB-MIT dataset, this patient-specific model attains classification accuracies of 95.67% and 99.60% for seizure prediction and detection, respectively. Notably, it maintains an average computational load of merely 0.05 million multiply-accumulate operations (MACs) per inference and has a compact model size of 6.87 K parameters. These results represent a significant advancement compared with existing methods. Operating at a rate of 32 inferences per second, the computational load of the model for seizure prediction has been reduced by more than 19.4 times, and for seizure detection, by more than 6.4 times. Heng Zhang 0025, Linxiang Wang, Wenjie Fan 0004, Zhenglin Gu, Youbin Luo, Xingjie Zou, Chang Gao 0002, Qinyu Chen, Li Li 0003 |
ISCAS | 8 |
| 2025 | An Energy Efficient Residual Spiking Neural Network Accelerator With Ternary SpikesabstractSpiking neural networks (SNNs) use discrete binary spikes to transfer information between neurons, which is different from artificial neural networks (ANNs). Although event-based characteristics bring potential computation power and efficiency to SNNs, the long processing time window of discrete spikes leads to high latency. In this brief, a spike version of the residual network using ternary spikes is proposed. A shorter time window is required to achieve competitive performance because the ability to transfer information of the ternary spikes is strengthened. An SNN accelerator based on the proposed residual network with ternary spikes is designed and implemented with 28 nm CMOS technology, and the core area is 0.63 mm2. The proposed SNN accelerator achieves the classification accuracy of 92.07% on CIFAR-10 dataset with SResNet20 and only 6 time steps. The accelerator achieves 0.39 mJ energy consumption per frame with a throughput of 165.7 FPS when running at 500 MHz. Congyi Sun, Wenqing Song, Qinyu Chen, Chenyang Dai, Li Li 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Toward Efficient Eye Tracking in AR/VR Devices: A Near-Eye DVS-Based Processor for Real-Time Gaze EstimationabstractThis paper presents an efficient near-eye dynamic vision sensor (DVS)-based processor for real-time eye tracking in augmented reality/virtual reality (AR/VR) devices. The processor takes advantage of the sparse event data with fine time resolution from the DVS, addressing the need for high frame-rate, low-power, and accurate eye tracking on wearable devices with extended battery life. Exploiting the inherent sparsity of event data, we propose an event-density-based region of interest (ROI) determination method that operates directly on event stream, which requires$47\times $fewer operations than the traditional methods, effectively overcoming the latency problem caused by the heavy computational loads. To eliminate the issue of decreasing accuracy at the edges of the field of view (FoV), we customized and fine-tuned a neural network for gaze estimation, ensuring uniformly distributed sub-degree accuracy. An estimator with a streamlined output mapping strategy and an adaptive window-sliding convolution scheme is implemented for gaze estimation acceleration. The processor is designed and fabricated in UMC 40-nm LP technology with a core area of 2.52 mm2 and performs end-to-end eye tracking exclusively with the raw event stream from DVS, achieving an average accuracy of 0.91° within a$96^{\circ } \times 64^{\circ }$FoV. Operating at 200 MHz, it achieves a dynamic frame rate of up to 1.2 kHz and requires only$12.7~\mu $J of energy per gaze estimation. By integrating the DVS, the processor enables real-time, low-power, and accurate eye tracking, enhancing the immersive experience on AR/VR devices and offering intuitive and seamless interactions. Shihang Tan, Jinqiao Yang, Ziyi Yang 0014, Qinyu Chen, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | A 1D-AE-PINN Crack Quantification Network Inspired by a Novel Physical Feature of ACFMabstractAlternating current field measurement (ACFM) is widely used in the quantitative detection of crack due to its advantages of noncontact measurement and high accuracy. However, the noncontact measurement introduces signal interference including constant and random lift-off. Both lift-offs bring challenges to the accurate quantification of cracks. It is difficult to obtain bothBxandBzsignals effectively. In this article, a new 1D-AE-PINN framework to accurately quantify the crack under the lift-off interference is proposed. A novel insight feature ofBxsignal with physical information about the crack size is studied and integrated into loss functions of the 1D-AE-PINN. The features encoded by 1D-AE-PINN are used as input to the quantization network. The advantages of 1D-AE-PINN in accuracy are proved by comparative experiments. The results show that the length and depth of the crack can be measured by onlyBxsignal. The mean squared errors of length and depth are 0.66 and 0.39 mm2. Jianxi Ding, Xin'an Yuan, Wei Li 0072, Baoping Cai, Xiaokang Yin 0001, Xiao Li 0032, Jianchao Zhao, Qinyu Chen, Zichen Nie, Qiyue Yin, Jianming Zhao |
IEEE Trans. Ind. Informatics | 8 |
| 2024 | Selecting Large Language Model to Fine-tune via Rectified Scaling LawabstractThe ever-growing ecosystem of LLMs has posed a challenge in selecting the most appropriate pre-trained model to fine-tune amidst a sea of options. Given constrained resources, fine-tuning all models and making selections afterward is unrealistic. In this work, we formulate this resource-constrained selection task into predicting fine-tuning performance and illustrate its natural connection with Scaling Law. Unlike pre-training, we find that the fine-tuning scaling curve includes not just the well-known "power phase" but also the previously unobserved "pre-power phase". We also explain why existing Scaling Law fails to capture this phase transition phenomenon both theoretically and empirically. To address this, we introduce the concept of "pre-learned data size" into our Rectified Scaling Law, which overcomes theoretical limitations and fits experimental results much better. By leveraging our law, we propose a novel LLM selection algorithm that selects the near-optimal model with hundreds of times less resource consumption, while other methods may provide negatively correlated selection. The project page is available at rectified-scaling-law.github.io. Haowei Lin, Baizhou Huang, Haotian Ye, Qinyu Chen, Sujian Li, Jianzhu Ma, Xiaojun Wan 0001, James Zou 0001, Yitao Liang |
ICML | 4 |
| 2024 | Epilepsy Seizure Detection and Prediction using an Approximate Spiking Convolutional TransformerabstractEpilepsy is a common disease of the nervous system. Timely prediction of seizures and intervention treatment can significantly reduce the accidental injury of patients and protect the life and health of patients. This paper presents a tiny neuromorphic Spiking Convolutional Transformer, named Spiking Conformer, to detect and predict epileptic seizure segments from scalped long-term electroencephalogram (EEG) recordings. We report evaluation results from the Spiking Conformer model using the Boston Children’s Hospital-MIT (CHB-MIT) EEG dataset. By leveraging spike-based addition operations, the Spiking Conformer significantly reduces the classification computational cost compared to the non-spiking model. Additionally, we introduce an approximate spiking neuron layer to further reduce spike-triggered neuron updates by nearly 38% without sacrificing accuracy. Using raw EEG data as input, the proposed Spiking Conformer achieved an average sensitivity rate of 94.9% and a specificity rate of 99.3% for the seizure detection task, and 96.8%, 89.5% for the seizure prediction task, and needs >10x fewer operations compared to the non-spiking equivalent model. Qinyu Chen, Congyi Sun, Chang Gao 0002, Shih-Chii Liu |
ISCAS | 1 |
| 2024 | HAS-RL: A Hierarchical Approximate Scheme Optimized With Reinforcement Learning for NoC-Based NN AcceleratorsabstractNetwork-on-Chip (NoC) is a scalable on-chip communication architecture for the NN accelerator, but with the increase in the number of nodes, the communication delay becomes higher. Applications such as machine learning have a certain resilience to noisy/erroneous transmitted data. Therefore, approximate communication becomes a promising solution to improving performance by reducing traffic loads under the constraint of the acceptable maximum accuracy loss of neural networks. It is a key issue to balance the result quality and the communication delay for approximate NoC systems. The traditional approximate NoC only considers the node-to-node approximation-based dynamic traffic regulation. However, the dynamically changing traffic patterns across different nodes, different times, and different applications lead to a huge search space, which makes it hard to explore an optimal global approximation solution. In this paper, we propose a quality model for different neural networks, which presents the relationship between the quality loss and the data approximate rate. Then, a hierarchical approximate scheme optimized with reinforcement learning (HAS-RL) is proposed and we reduce the complexity of the HAS-RL by reducing the state space and action space, which will reduce the resource overhead as well. After that, we embed a global approximate controller in the NoC system, in which we deploy a policy network trained with the offline reinforcement learning algorithm to adjust the data approximate rates of each node at run time. Compared with the state-of-the-art method, the proposed scheme reduces the average network delay by 13.5% while their accuracies are similar. The proposed HAS-RL only causes an additional area overhead of 1.24% and power consumption of 0.77% compared with the traditional router design. Shize Zhou, Yongqi Xue, Wenjie Fan 0004, Tong Cheng, Jinlun Ji, Chenyang Dai, Wenqing Song, Qinyu Chen, Chang Gao 0002, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2023 | An Area-Efficient Ultra-Low-Power Time-Domain Feature Extractor for Edge Keyword SpottingabstractKeyword spotting (KWS) is an important task on edge low-power audio devices. A typical edge KWS system consists of a front-end feature extractor which outputs mel-scale frequency cepstral coefficients (MFCC) features followed by a back-end neural network classifier. KWS edge designs aim for the best power-performance-area metrics. This work proposes an area-efficient ultra-low-power time-domain infinite impulse response (IIR) filter-based feature extractor for a KWS system. It uses a serial architecture, and the architecture is further optimized for a low-cost computing structure and mixed-precision bit selection of the IIR coefficients while maintaining good KWS accuracy. Using a 65 nm process technology and a back-end neural network classifier, this simulated feature extractor has an area of 0.02 mm2and achieves$\mathbf{3.3}\mu \mathbf{W}$@ 1.2 V, and achieves 92.5% accuracy on a 10-keyword, 12-class KWS task using the GSCD dataset. Qinyu Chen, Yaoxing Chang, Kwantae Kim, Chang Gao 0002, Shih-Chii Liu |
ISCAS | 1 |
| 2023 | An End-to-End Physics-Informed Neural Network for Defect Identification and 3-D Reconstruction Using Rotating Alternating Current Field MeasurementabstractThe alternating current field measurement (ACFM) technique has been widely used in the defect detection of metal structures. However, the identification and reconstruction of defects depend on human experience or simple empirical formulas, which leads to misjudgment of defects and large quantization errors. In this article, we propose an end-to-end physics-informed neural network for defect identification and 3-D reconstruction. The high-precision automatic detection system with a specially designed probe is established to detect defects in any direction. The faster RCNN network is used to identify and classify defects. The physics-informed Pix2Pix network is constructed to realize 3-D reconstruction of defects with different types. The results show that the established end-to-end physics-informed neural network can realize the identification and 3-D reconstruction of defects in which the mean average precision is 0.9982, the average length error of cracks is 0.9249 mm, the average depth error of cracks is 0.3402 mm, the average volume error of corrosion is 0.0667, and the average maximum depth error of corrosion is 0.3464 mm. Jianming Zhao, Wei Li 0072, Xin'an Yuan, Xiaokang Yin 0001, Xiao Li 0032, Qinyu Chen, Jianxi Ding |
IEEE Trans. Ind. Informatics | 6 |
| 2022 | Domain Adaptation via Maximizing Surrogate Mutual InformationabstractUnsupervised domain adaptation (UDA), which is an important topic in transfer learning, aims to predict unlabeled data from target domain with access to labeled data from the source domain. In this work, we propose a novel framework called SIDA (Surrogate Mutual Information Maximization Domain Adaptation) with strong theoretical guarantees. To be specific, SIDA implements adaptation by maximizing mutual information (MI) between features. In the framework, a surrogate joint distribution models the underlying joint distribution of the unlabeled target domain. Our theoretical analysis validates SIDA by bounding the expected risk on target domain with MI and surrogate distribution bias. Experiments show that our approach is comparable with state-of-the-art unsupervised adaptation methods on standard UDA tasks. Haiteng Zhao, Qinyu Chen, Zhi-Hong Deng 0001 |
IJCAI | 3 |
| 2022 | Unsupervised Learning Based on Temporal Coding Using STDP in Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have been recognized as one of the next generation of Neural Networks (NNs), showing a great potential in a variety of applications. Spiking-Timing Dependent Plasticity (STDP) underlies the brain’s learning mechanisms, and trains SNNs with great energy efficiency. In this paper, we propose a low-cost spike-time based unsupervised learning method. It constructs a SNN with one fully-connected excitatory layer structure without inhibitory layer, and trains the SNN with STDP using a first-spike-based temporal coding scheme where input information is directly encoded into spike times. It only updates the synaptic weights connected to the neuron that first generates a spike in a forward propagation step, which reduces the frequency of the synaptic weight updates significantly. The forward propagation process can be stopped once a neuron fires whether in the training mode or the inference mode, by which many unnecessary computations are just avoided and the latency in the inference mode is reduced. The method was used to train on the classification task on MNIST dataset and achieved an accuracy of 90.4% with 800 excitatory neurons. Congyi Sun, Qinyu Chen, Kai Chen 0034, Guoqiang He, Li Li 0003 |
ISCAS | 2 |
| 2022 | Skydiver: A Spiking Neural Network Accelerator Exploiting Spatio-Temporal Workload BalanceabstractSpiking neural networks (SNNs) are developed as a promising alternative to artificial neural networks (ANNs) due to their more realistic brain-inspired computing models. SNNs have sparse neuron firing over time, i.e., spatio-temporal sparsity; thus, they are useful to enable energy-efficient hardware inference. However, exploiting spatio-temporal sparsity of SNNs in hardware leads to unpredictable and unbalanced workloads, degrading the energy efficiency. In this work, we propose an FPGA-based convolutional SNN accelerator called Skydiver that exploits spatio-temporal workload balance. We propose the approximate proportional relation construction (APRC) method that can predict the relative workload channel-wisely and a channel-balanced workload schedule (CBWS) method to increase the hardware workload balance ratio to over 90%. Skydiver was implemented on a Xilinx XC7Z045 FPGA and verified on image segmentation and MNIST classification tasks. Results show improved throughput by$1.4\times $and$1.2\times $for the two tasks. Skydiver achieved 22.6KFPS throughput, and$42.4~\mu \text{J}$/image prediction energy on the classification task with 98.5% accuracy. Qinyu Chen, Chang Gao 0002, Xinyuan Fang, Haitao Luan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | An Energy Efficient STDP-Based SNN Architecture With On-Chip LearningabstractIn this paper, we propose a spike-time based unsupervised learning method using spiking-timing dependent plasticity (STDP). A simplified linear STDP learning rule is proposed for the energy efficient weight updates. To reduce unnecessary computations for the input spike values, a stop mechanism of the forward pass is introduced in the forward pass. In addition, a hardware-friendly input quantization scheme is used to reduce the computational complexities in both the encoding phase and the forward pass. We construct a two-layer fully-connected spiking neuron network (SNN) based on the proposed method. Compared to general rate-based SNNs trained by STDP, the proposed method reduces the complexity of network architecture (an extra inhibitory layer is not needed) and the computations of synaptic weight updates. According to the fixed-point simulation with 9-bit synaptic weights, the proposed SNN with 6144 excitatory neurons achieves 96% of recognition accuracy on MNIST dataset without any supervision. An SNN processor that contains 384 excitatory neurons with on-chip learning capability is designed and implemented with 28 nm CMOS technology based on the proposed low complexity methods. The SNN processor achieves an accuracy of 93% on MNIST dataset. The implementation results show that the SNN processor achieves a throughput of 277.78k FPS with$0.50~\mu \text{J}$/inference energy consuming in inference mode, and a throughput of 211.77k FPS with$0.66~\mu \text{J}$/learning energy consuming in learning mode. Congyi Sun, Haohan Sun, Jianing Han, Xinyuan Wang 0007, Xinyu Wang 0027, Qinyu Chen, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | Cerebron: A Reconfigurable Architecture for Spatiotemporal Sparse Spiking Neural NetworksabstractSpiking neural networks (SNNs) are promising alternatives to artificial neural networks (ANNs) since they are more realistic brain-inspired computing models. SNNs have sparse neuron firing over time, i.e., spatiotemporal sparsity; thus, they are helpful in enabling energy-efficient hardware inference. However, exploiting the spatiotemporal sparsity of SNNs in hardware leads to unpredictable and unbalanced workloads, degrading the energy efficiency. Compared to SNNs with simple fully connected structures, those extensive structures (e.g., standard convolutions, depthwise convolutions, and pointwise convolutions) can deal with more complicated tasks but lead to difficulties in hardware mapping. In this work, we propose a novel reconfigurable architecture, Cerebron, which can fully exploit the spatiotemporal sparsity in SNNs with maximized data reuse and propose optimization techniques to improve the efficiency and flexibility of the hardware. To achieve flexibility, the reconfigurable compute engine is compatible with a variety of spiking layers and supports inter-computing-unit (CU) and intra-CU reconfiguration. The compute engine can exploit data reuse and guarantee parallel data access when processing different convolutions to achieve memory efficiency. A two-step data sparsity exploitation method is introduced to leverage the sparsity of discrete spikes and reduce the computation time. Besides, an online channelwise workload scheduling strategy is designed to reduce the latency further. Cerebron is verified on image segmentation and classification tasks using a variety of state-of-the-art spiking network structures. Experimental results show that Cerebron has achieved at least$17.5\times $prediction energy reduction and$20\times $speedup compared with state-of-the-art field-programmable gate array (FPGA)-based accelerators. Qinyu Chen, Chang Gao 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Reducing Latency in a Converted Spiking Video Segmentation NetworkabstractSpiking Neural Networks (SNNs) can be configured to produce almost-equivalent accurate Analog Neural Networks (ANNs) by various ANN-SNN conversion methods. Most of these methods are applied to classification and object detection networks tested on frame-based datasets. In this work, we demonstrate a converted SNN for image segmentation and applied to a natural video dataset. Instead of resetting the network state with each input frame, we capitalize on the temporal redundancy between adjacent frames in a natural scene, and propose an interval reset method where the network state is reset after a fixed number of frames. We studied the trade-off between accuracy and latency with the number of interval reset frames. We also applied layer-specific normalization and early stopping to speed up network convergence and to reduce the latency. Our results show that the SNN achieved a 35.7x increase in convergence speed with only 1.5% accuracy drop using an interval reset of 20 frames. Qinyu Chen, Bodo Rueckauer, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 1 |
| 2021 | Optimizing Vertical Link Placement and Congestion Aware Dynamic Elevator Assignment for Partially Connected 3D-NoCsabstractThe fully connected 3D-NoCs in which all routers are vertically connected with their neighbors above and below need a lot of Through-Silicon-Vias (TSVs), and they will occupy a large silicon area and reduce the fabrication yield. Thus, the idea of partially connected 3D-NoCs has emerged. The optimal number and placement of the vertical links (elevators) must be determined at the chip design stage, which is a multiobjective optimization problem of the performance and the cost. However, optimizing the static elevator placement needs a great amount of calculation and we can not examine all possible solutions at design time. Therefore, we propose a hybrid heuristic strategy for the static elevator placement and assignment, in which the genetic algorithm and the tabu search are combined. The dynamic assignment method is essential for the partially connected 3D-NoCs, and it leads to different traffic distributions and therefore has a huge impact on performance. Many previous static assignment methods can not dynamically change the elevator assignment according to the real-time states of the network, thus it may lead to network congestion. A congestion-aware dynamic assignment (CDA) scheme is proposed in this article, which considers the impact of the distance factor and the congestion factor on the network performance. Experiments show that the proposed CDA method can improve the network performance by 67%-86% compared with the random selection algorithm and can improve the reliability of the partially connected 3D-NoC as well. The key component for the CDA method, the path selection module (PSM), is implemented in FPGA, and the results show that its area cost is negligible compared with a router. Chuan Zhang 0001, Wenqing Song, Qinyu Chen, Hui Chen 0015, Li Li 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Exploring How Game Genre in Student-Designed Games Influences Computational Thinking DevelopmentabstractGame design is increasingly used in modern education to foster Computational Thinking (CT). Yet, it is unclear how and if the game genre of student-designed games impact CT and programming. We explore how game genre impacts CT development and programming routines in Scratch games designed by 8th-grade students using a metrics-based approach (i.e., Dr. Scratch). Our findings show that designing particular games (e.g., action, storytelling) impact CT and programming development. We observe, for instance, that CT skills develop and consolidate fast, after which students can focus on aspects more specific to game design. Based on the results, we suggest that researchers and educators in constructionist learning consider the impact of game genre when designing game-based curricula for the learning of programming and CT. Giovanni Maria Troiano, Qinyu Chen, Ángela Vargas-Alba, Gregorio Robles, Gillian Smith 0001, Michael P. Cassidy, Eli Tucker-Raymond, Gillian Puttick, Casper Harteveld |
CHI | 2 |
| 2020 | An Efficient Accelerator for Multiple Convolutions From the Sparsity PerspectiveabstractConvolutional neural networks (CNNs) have emerged as one of the most popular ways applied in many fields. These networks deliver better performance when going deeper and larger. However, the complicated computation and huge storage impede hardware implementation. To address the problem, quantized networks are proposed. Besides, various convolutional structures are designed to meet the requirements of different applications. For example, compared with the traditional convolutions (CONVs) for image classification, CONVs for image generation are usually composed of traditional CONVs, dilated CONVs, and transposed CONVs, leading to a difficult hardware mapping problem. In this brief, we translate the difficult mapping problem into the sparsity problem and propose an efficient hardware architecture for sparse binary and ternary CNNs by exploiting the sparsity and low bit-width characteristics. To this end, we propose an ineffectual data removing (IDR) mechanism to remove both the regular and irregular sparsity based on dual-channel processing elements (PEs). Besides, a flexible layered load balance (LLB) mechanism is introduced to alleviate the load imbalance. The accelerator is implemented with 65-nm technology with a core size of 2.56 mm2. It can achieve 3.72-TOPS/W energy efficiency at 50.1 mW, which makes it a promising design for embedded devices. Qinyu Chen, Wenqing Song, Zhonghai Lu, Li Li 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Smilodon: An Efficient Accelerator for Low Bit-Width CNNs with Task PartitioningabstractConvolutional Neural Networks (CNNs) have been widely applied in various fields such as image and video recognition, recommender systems, and natural language processing. However, the massive size and intensive computation loads prevent its feasible deployment in practice, especially on the embedded systems. As a highly competitive candidate, low bit-width CNNs are proposed to enable efficient implementation. In this paper, we propose Smilodon, a scalable, efficient accelerator for low bit-width CNNs based on a parallel streaming architecture, optimized with a task partitioning strategy. We also present the 3D systolic-like computing arrays fitting for convolutional layers. Our design is implemented on Zynq XC7Z020 FPGA, which can satisfy the needs of real-time with a frame rate of 1, 622 FPS throughput, while consuming 2 1 Watt. To the best of our knowledge, our accelerator is superior to the state-of-the-art works in the tradeoff among throughput, power efficiency, and area efficiency. Qinyu Chen, Kaifeng Cheng, Wenqing Song, Zhonghai Lu, Li Li 0003, Chuan Zhang 0001 |
ISCAS | 1 |
| 2019 | Congestion-Aware Dynamic Elevator Assignment for Partially Connected 3D-NoCsabstractThe combination of Network-on-Chips (NoCs) and 3D IC technology, 3D NoCs, has been proven to be able to achieve a great improvement in both network performance and power consumption compared to 2D NoCs. In the traditional 3D NoC, all routers are vertically connected. Due to the large overhead of Through-Silicon-Via (TSV, e.g., low fabrication yield and the occupied silicon area), the partially connected 3D NoC has emerged. The assignment method determines the traffic loads of the vertical links (elevators), thus has a great impact on 3D-NoCs' performance. In this paper, we propose a congestion-aware dynamic elevator assignment (CDA) scheme, which takes both the distance factors and network congestion information into account. Experiments show that the performance of the proposed CDA scheme is improved by 67% to 87% compared to the random selection scheme, 8% to 25% compared to SelByDis-1, and 13% to 18% compared to SelByDis-2. Qinyu Chen, Guoqiang He, Kai Chen 0034, Zhonghai Lu, Chuan Zhang 0001, Li Li 0003 |
ISCAS | 2 |
| 2019 | Thermal Sensor Placement and Thermal Reconstruction Under Gaussian and Non-Gaussian Sensor Noises for 3-D NoCabstractOn-chip thermal sensors are essential for temperature management in 3-D network-on-chip (NoC) systems. However, due to the physical (area and power) or economical constraints, the number of sensors is limited. Therefore, the two critical issues we face are: 1) how to figure out an efficient thermal sensor placement with the limited number of sensors and 2) how to reconstruct the entire thermal profile based on sensor observations. Another major issue for the thermal reconstruction is the sensor measurement accuracy. Thus, online accurate full-chip thermal reconstruction under Gaussian and non-Gaussian noises is another great challenge. In this paper, a greedy thermal sensor placement algorithm maximizing the rank of the observability Gramian is proposed. A good placement algorithm always relies on a specific reconstruction method. The proposed placement algorithm is designed for the state-space-based thermal model, thus the combination of the proposed placement algorithm and the Kalman filter-based reconstruction method provides a high reconstruction accuracy under Gaussian noise. For accurate temperature reconstruction under non-Gaussian noise, the Gaussian-Sum filter is applied to 3-D NoC. Compared with the Kalman filter, the Gaussian-Sum filter can reduce the root-mean-squared-error and the max error by 29.27%–35% and 33.26%–40.6%, respectively. A reusable architecture for the Kalman filter and the Gaussian-Sum filter has been proposed. Its hardware implementation details are presented in this paper. Besides, the performance and the area are evaluated as well. Li Li 0003, Hongbing Pan, Kun Wang 0005, Qinyu Chen, Chuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2004 | NUAT B-spline curves
Guozhao Wang, Qinyu Chen, Minghua Zhou |
Comput. Aided Geom. Des. | 2 |
| 2003 | A class of Bézier-like curves
Qinyu Chen, Guozhao Wang |
Comput. Aided Geom. Des. | 1 |