EDBT 2026 Demo / reviewers in the wild / expert
Zhe Chen 0030
dblp:06/4240-30
· DBLP profile ↗
9ranked-venue papers
7as first author
2since 2021 · last 2022
0000-0002-5371-2058ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Electronic design automation · 47% Hardware accelerators and domain-specific architectures · 41% Reconfigurable computing and FPGAs · 12% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › signal processing accelerator
biomedical signal processing accelerator |
1.1 | 3 | 2020 | CANSEE: Customized Accelerator for Neural Signal Enhancement and Extraction from the Calcium Image in Real Time · FPGA 2020 LANMC: LSTM-Assisted Non-Rigid Motion Correction on FPGA for Calcium Image Stabilization · FPGA 2019 FPGA-based LSTM Acceleration for Real-Time EEG Signal Processing: (Abstract Only) · FPGA 2018 |
Electronic design automation
high-level synthesis |
0.9 | 2 | 2020 | Analysis and Optimization of the Implicit Broadcasts in FPGA HLS to Improve Maximum Frequency · FPGA 2020 Analysis and Optimization of the Implicit Broadcasts in FPGA HLS to Improve Maximum Frequency · DAC 2020 |
Electronic design automation › physical design
timing optimization |
0.4 | 1 | 2020 | Analysis and Optimization of the Implicit Broadcasts in FPGA HLS to Improve Maximum Frequency · FPGA 2020 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.3 | 3 | 2020 | CANSEE: Customized Accelerator for Neural Signal Enhancement and Extraction from the Calcium Image in Real Time · FPGA 2020 LANMC: LSTM-Assisted Non-Rigid Motion Correction on FPGA for Calcium Image Stabilization · FPGA 2019 FPGA-based LSTM Acceleration for Real-Time EEG Signal Processing: (Abstract Only) · FPGA 2018 |
Methods — techniques the papers use, named apart from their topics
LSTM inference · 0.8synchronization pruning · 0.4skid-buffer-based pipeline control · 0.4skid-buffer flow control · 0.4motion correction · 0.4dynamic programming · 0.4broadcast-aware scheduling · 0.4non-rigid motion correction · 0.4causal filtering · 0.3LSTM · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Efficient Kernels for Real-Time Position Decoding from In Vivo Calcium ImagesabstractRecent studies have found that the position of mice or rats can be decoded from calcium imaging of brain activity offline. However, given the complex analysis pipeline, real-time position decoding remains a challenging task, especially considering strict requirements on hardware usage and energy cost for closed-loop feedback applications. In this paper, we propose two neural network based methods and corresponding hardware designs for real-time position decoding from calcium images. Our methods are based on: 1) convolutional neural network (CNN), 2) spiking neural network (SNN) converted from the CNN. We implemented quantized CNN and SNN models on FPGA. Evaluation results show that the CNN and the SNN methods achieve 56.3%/83.1% and 56.0%/82.8% Hit-1/Hit-3 accuracy for the position decoding across different rats, respectively. We also observed an accuracy-latency tradeoff of the SNN method in decoding positions under various time steps. Finally, we present our SNN implementation on the neuromorphic chip Loihi. Zhe Chen 0030, Jim Zhou, Garrett J. Blair, Hugh T. Blair, Jason Cong |
ISCAS | 1 |
| 2022 | Energy-Efficient LSTM Inference Accelerator for Real-Time Causal PredictionabstractEver-growing edge applications often require short processing latency and high energy efficiency to meet strict timing and power budget. In this work, we propose that the compact long short-term memory (LSTM) model can approximate conventional acausal algorithms with reduced latency and improved efficiency for real-time causal prediction, especially for the neural signal processing in closed-loop feedback applications. We design an LSTM inference accelerator by taking advantage of the fine-grained parallelism and pipelined feedforward and recurrent updates. We also propose a bit-sparse quantization method that can reduce the circuit area and power consumption by replacing the multipliers with the bit-shift operators. We explore different combinations of pruning and quantization methods for energy-efficient LSTM inference on datasets collected from the electroencephalogram (EEG) and calcium image processing applications. Evaluation results show that our proposed LSTM inference accelerator can achieve 1.19 GOPS/mW energy efficiency. The LSTM accelerator with 2-sbit/16-bit sparse quantization and 60% sparsity can reduce the circuit area and power consumption by 54.1% and 56.3%, respectively, compared with a 16-bit baseline implementation. Zhe Chen 0030, Hugh T. Blair, Jason Cong |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2020 | Analysis and Optimization of the Implicit Broadcasts in FPGA HLS to Improve Maximum FrequencyabstractDesigns generated by high-level synthesis (HLS) tools typically achieve a lower frequency compared to manual RTL designs. In this work, we study the timing issues in a diverse set of realistic and complex FPGA HLS designs. (1) We observe that in almost all cases the frequency degradation is caused by the broadcast structures generated by the HLS compiler. (2) We classify three major types of broadcasts in HLS-generated designs, including high-fanout data signals, pipeline flow control signals and synchronization signals for concurrent modules. (3) We reveal a number of limitations of the current HLS tools that result in those broadcast-related timing issues. (4) We propose a set of effective yet easy-to-implement approaches, including broadcast-aware scheduling, synchronization pruning, and skid-buffer-based flow control. Our experimental results show that our methods can improve the maximum frequency of a set of nine representative HLS benchmarks by 53% on average. In some cases, the frequency gain is more than 100 MHz. Licheng Guo, Jason Lau, Yuze Chi, Jie Wang 0022, Cody Hao Yu, Zhe Chen 0030, Zhiru Zhang, Jason Cong |
DAC | 6 |
| 2020 | CANSEE: Customized Accelerator for Neural Signal Enhancement and Extraction from the Calcium Image in Real TimeabstractMiniaturized fluorescent calcium imaging miniscope has become a prominent technique in monitoring the activity of a large population of neurons in vivo. However, existing calcium image processing algorithms are developed for off-line analysis, and their implementations on general-purpose processors are difficult to meet the real-time processing requirement under constrained energy budget for closed-loop applications. In this paper, we propose the CANSEE, a customized accelerator for neural signal enhancement and extraction from calcium image in real time. The accelerator can perform the motion correction, the calcium image enhancement, and the fluorescence tracing from up to 512 cells with less than 1-ms processing latency. We also designed the hardware that can detect new cells based on the long short-term memory (LSTM) inference. We implemented the accelerator on a Xilinx Ultra96 FPGA. The implementation achieves 15.8x speedup and over 2 orders of magnitude improvement in energy efficiency compared to the evaluation on the multi-core CPU. Zhe Chen 0030, Garrett J. Blair, Hugh T. Blair, Jason Cong |
FPGA | 1 |
| 2020 | Analysis and Optimization of the Implicit Broadcasts in FPGA HLS to Improve Maximum FrequencyabstractDesigns generated by high-level synthesis (HLS) tools typically achieve a lower frequency compared to manual RTL designs. We study the timing issues in a diverse set of nine realistic HLS designs and observe that in most cases the frequency degradation is related to the signal broadcast structures. In this work, we classify the common broadcast types in HLS designs, including the data signal broadcast and two types of control signal broadcast: the pipeline control broadcast and the synchronization signal broadcast. We further identify several common limitations of the current HLS tools, which lead to improper handling of the broadcasts. First, the HLS delay model does not consider the extra delay caused by broadcasts, thus the scheduling results will be suboptimal. To solve the issue, we implement a set of comprehensive synthetic designs and benchmark the extra delay to calibrate the HLS delay model. Second, the HLS adopts back-pressure signals for pipeline control, which will lead to large broadcasts. Instead, we propose to use the skid-buffer-based pipeline control, where the back-pressure signal is removed, and an extra skid-buffer is used for flow-control. We use dynamic programming to minimize the area of the extra FIFO. Third, there exist redundant synchronizations among concurrent modules that may lead to huge broadcasts. We propose methods to identify and prune unnecessary synchronization signals. Our solutions boost the frequency of nine real-world HLS benchmarks by 53% on average and with marginal area and latency overhead. In some cases, the gain is more than 100 MHz. Licheng Guo, Jason Lau, Yuze Chi, Jie Wang 0022, Cody Hao Yu, Zhe Chen 0030, Zhiru Zhang, Jason Cong |
FPGA | 6 |
| 2020 | BLINK: bit-sparse LSTM inference kernel enabling efficient calcium trace extraction for neurofeedback devicesabstractMiniaturized fluorescent calcium imaging microscopes are widely used for monitoring the activity of a large population of neurons in freely behaving animals in vivo. Conventional calcium image analyses extract calcium traces by iterative and bulk image processing and they are hard to meet the power and latency requirements for neurofeedback devices. In this paper, we propose the calcium image processing pipeline based on a bit-sparse long short-term memory (LSTM) inference kernel (BLINK) for efficient calcium trace extraction. It largely reduces the power and latency while remaining the trace extraction accuracy. We implemented the customized pipeline on the Ultra96 platform. It can extract calcium traces from up to 1024 cells with sub-ms latency on a single FPGA device. We designed the BLINK circuits in a 28-nm technology. Evaluation shows that the proposed bit-sparse representation can reduce the circuit area by 38.7% and save the power consumption by 38.4% without accuracy loss. The BLINK circuits achieve 410 pJ/inference, which has 6293x and 52.4x gains in energy efficiency compared to the evaluation on the high performance CPU and GPU, respectively. Zhe Chen 0030, Garrett J. Blair, Hugh T. Blair, Jason Cong |
ISLPED | 1 |
| 2019 | LANMC: LSTM-Assisted Non-Rigid Motion Correction on FPGA for Calcium Image StabilizationabstractCalcium imaging is an emerging technique for visualizing and recording neural population activity at large scale in vivo. Non-rigid motion correction is a critical step in the calcium image analysis pipeline due to non-uniform deformations of the brain tissue during the data collection. However, existing non-rigid motion correction algorithms are costly in computation time and energy, and it is hard to implement such algorithm in real time on an embedded device. In this paper, we propose LANMC, an LSTM-assisted non-rigid motion correction method for real-time calcium image stabilization. This method reduces the computational cost by using the LSTM inference to predict the non-rigid motion. Based on this method, we demonstrate a non-rigid motion correction implementation for real-time calcium image stabilization on FPGA. Experimental results show that the non-rigid motion correction can be accomplished within 80 µs on the Ultra96 under 300 MHz frequency, and the latency outperforms that on a 12-thread CPU by 82x. Zhe Chen 0030, Hugh T. Blair, Jason Cong |
FPGA | 1 |
| 2018 | FPGA-based LSTM Acceleration for Real-Time EEG Signal Processing: (Abstract Only)abstractClosed-loop neurofeedback is a growing area of research and development for novel therapies to treat brain disorders. A neurofeedback device can detect disease symptoms (such as motor tremors or seizures) in real time from electroencephalogram (EEG) signals, and respond by rapidly delivering neurofeedback stimulation that relieves these symptoms. Conventional EEG processing algorithms rely on acausal filters, which impose delays that can exceed the short feedback latency required for closed-loop stimulation. In this paper, we first introduce a method for causal filtering using long short-term memory (LSTM) networks, which radically reduces the filtering latency. We then propose a reconfigurable architecture that supports time-division multiplexing of LSTM inference engines on a prototype neurofeedback device. We implemented a 128-channel EEG signal processing design on a Zynq-7030 device, and demonstrated its feasibility. Then, we further scaled up the design onto Zynq-7045 and Virtex-690t devices to achieve high performance and energy efficient implementations for massively parallel brain signal processing. We evaluated the performance against optimized implementations on CPU and GPU at the same CMOS technology node. Experiment results show that the Virtex-690t can achieve 1.32x and 11x speed-up against the K40c GPU and the multi-thread Xeon E5-2860 CPU, respectively, while FPGA achieves 6.1x and 26.6x energy efficiency compared to the GPU and CPU. Zhe Chen 0030, Andrew Howe, Hugh T. Blair, Jason Cong |
FPGA | 1 |
| 2018 | CLINK: Compact LSTM Inference Kernel for Energy Efficient Neurofeedback DevicesabstractNeurofeedback device measures brain wave and generates feedback signal in real time and can be employed as treatments for various neurological diseases. Such devices require high energy efficiency because they need to be worn or surgically implanted into patients and support long battery life time. In this paper, we propose CLINK, a compact LSTM inference kernel, to achieve high energy efficient EEG signal processing for neurofeedback devices. The LSTM kernel can approximate conventional filtering functions while saving 84% computational operations. Based on this method, we propose energy efficient customizable circuits for realizing CLINK function. We demonstrated a 128-channel EEG processing engine on Zynq-7030 with 0.8 W, and the scaled up 2048-channel evaluation on Virtex-VU9P shows that our design can achieve 215x and 7.9x energy efficiency compared to highly optimized implementations on E5-2620 CPU and K80 GPU, respectively. We carried out the CLINK design in a 15-nm technology, and synthesis results show that it can achieve 272.8 pJ/inference energy efficiency, which further outperforms our design on the Virtex-VU9P by 99x. Zhe Chen 0030, Andrew Howe, Hugh T. Blair, Jason Cong |
ISLPED | 1 |