EDBT 2026 Demo / reviewers in the wild / expert
Chuping Qu
dblp:257/5381
· DBLP profile ↗
7ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-9398-6037ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimizing Neural Networks with Learnable Non-Linear Activation Functions via Lookup-Based FPGA AccelerationabstractLearned activation functions in models like Kolmogorov-Arnold Networks (KANs) outperform fixed-activation architectures in terms of accuracy and interpretability; however, their computational complexity poses critical challenges for energy-constrained edge AI deployments. Conventional CPUs/GPUs incur prohibitive latency and power costs when evaluating higher order activations, limiting deployability under ultra-tight energy budgets. We address this via a reconfigurable lookup architecture with edge FPGAs. By coupling fine-grained quantization with adaptive lookup tables, our design minimizes energy-intensive arithmetic operations while preserving activation fidelity. FPGA reconfigurability enables dynamic hardware specialization for learned functions, a key advantage for edge systems that require post-deployment adaptability. Evaluations using KANs - where unique activation functions play a critical role—demonstrate that our FPGA-based design achieves superior computational speed and over 104times higher energy efficiency compared to edge CPUs and GPUs, while maintaining matching accuracy and minimal footprint overhead. This breakthrough positions our approach as a practical enabler for energy-critical edge AI, where computational intensity and power constraints traditionally preclude the use of adaptive activation networks. Mengyuan Yin, Benjamin Chen Ming Choong, Chuping Qu, Rick Siow Mong Goh, Weng-Fai Wong, Tao Luo 0014 |
ICCAD | 3 |
| 2023 | DeepFire2: A Convolutional Spiking Neural Network Accelerator on FPGAsabstractBrain-inspired spiking neural networks (SNNs) replace the multiply-accumulate operations of traditional neural networks by integrate-and-fire neurons, with the goal of achieving greater energy efficiency. Specialized hardware implementations of those neurons clearly have advantages over general-purpose devices in terms of power and performance, but exhibit poor scalability when it comes to accelerating large neural networks. DeepFire2 introduces a hardware architecture which can map large network layers efficiently across multiple super logic regions in a multi-die FPGA. That gives more control over resource allocation and parallelism, benefiting both throughput and energy consumption. Avoiding the use of lookup tables to implement theANDoperations of an SNN, prevents the layer size to be limited by logic resources. A deep pipeline does not only lead to an increased clock speed of up to 600 MHz. We double the throughput and power efficiency compared to our previous version of DeepFire, which equates to an almost 10-fold improvement over other previous implementations. Importantly, we are able to deploy a large ImageNet model, while maintaining a throughput of over 1500 frames per second. Myat Thu Linn Aung, Daniel Gerlinghoff, Chuping Qu, Tian Huang, Rick Siow Mong Goh, Tao Luo 0014, Weng-Fai Wong |
IEEE Trans. Computers | 3 |
| 2022 | Coreset: Hierarchical neuromorphic computing supporting large-scale neural networks with improved resource efficiency
Huaipeng Zhang, Tao Luo 0014, Chuping Qu, Myat Thu Linn Aung, Yingnan Cui, Jun Zhou 0014, Ming Ming Wong, Junran Pu, Anh-Tuan Do, Rick Siow Mong Goh, Weng-Fai Wong |
Neurocomputing | 4 |
| 2022 | Corrigendum to "Coreset: Hierarchical neuromorphic computing supporting large-scale neural networks with improved resource efficiency" [Neurocomputing (2022) 128-140]
Huaipeng Zhang, Tao Luo 0014, Chuping Qu, Myat Thu Linn Aung, Yingnan Cui, Jun Zhou 0014, Ming Ming Wong, Junran Pu, Anh-Tuan Do, Rick Siow Mong Goh, Weng-Fai Wong |
Neurocomputing | 4 |
| 2022 | NC-Net: Efficient Neuromorphic Computing Using Aggregated Subnets on a Crossbar-Based Architecture With Nonvolatile MemoryabstractNeuromorphic computing chips consisting of crossbar arrays of emergent nonvolatile memory (NVM) have the potential of achieving both high energy efficiency and throughput as the low-power implementation of convolutional neural network (CNN) inference engines. However, such hardware has design constraints, such as its limited fan-in/fan-out and resource-inefficient mapping, that make the design and deployment of CNN on them challenging. As a result, the user has to design the CNN model with intricate knowledge of the hardware architecture and even cannot fit the models in the hardware for CNN with high resolution image input. In this article, we propose the use ofaggregated subnets, NC-net, which is a constrained form of the traditional layer structure, to solve these issues. With our method, we put forward an energy-efficient buffer- and analogue-to-digital converter and digital-to-analogue converter (ADC/DAC)-free architecture and a scalable end-to-end solution that automatically satisfies the hardware constraints of crossbar architectures, while optimizing the resource usage. In our solution, the exploration and deployment of a CNN for a neuromorphic crossbar hardware start with a design front end based onTensorFlow. Our automated design flow maps the NC-net network fromTensorFlowto the crossbar architecture. We tested our designs on both a simulator and a field-programmable gate array (FPGA) emulator with various benchmarks. In addition to general benchmarks, including MNIST, SVHN, CIFAR-10, and CIFAR-100, we tested our system on a real-world application, human detection with high resolution (224$\times $224) images as the input. Our system achieves the state-of-the-art accuracy for these benchmarks on the crossbar-based neuromorphic hardware, with an accuracy of more than 90% for the latter. It also yielded up to$4.25\times $improvement in the efficiency of spiking core usage compared to TrueNorth. Tao Luo 0014, Huaipeng Zhang, Chuping Qu, Yingnan Cui, Weng-Fai Wong, Rick Siow Mong Goh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | DeepFire: Acceleration of Convolutional Spiking Neural Network on Modern Field Programmable Gate ArraysabstractSpiking neural networks (SNN) with their ‘integrate and fire’ (I&F) neurons replace the hardware-intensive multiply-accumulate (MAC) operations in convolutional neural networks (CNN) with accumulate operations — not only making it easy to implement on FPGAs but also opening up the opportunities for energy-efficient hardware acceleration. In this paper, we propose DeepFire — the high-performance RTL IP — for accelerating convolutional SNN inference. The IP exploits various resources available on modern FPGAs, and it outperforms existing SNN implementations by more than 10× in terms of both frame per second (FPS) and performance per watt (FPS/Watt). Our design achieves up to 40.1kFPS and 28.3kFPS on MNIST and CIFAR-10/SVHN datasets with 99.14% and 81.8%/93.1% accuracies respectively. IP was evaluated with 7-series and Ultrascale+ FPGAs from Xilinx achieving Fmax of 375MHz and 500MHz respectively. Myat Thu Linn Aung, Chuping Qu, Tao Luo 0014, Rick Siow Mong Goh, Weng-Fai Wong |
FPL | 2 |
| 2020 | An FPGA-Based Hardware Emulator for Neuromorphic Chip With RRAMabstractNeuromorphic chip with RRAM devices has been demonstrated as a promising computing platform for neural network-based applications. By directly mapping the weight matrices of neural networks onto RRAM-based crossbar arrays, high energy, and area efficiency can be achieved. However, the design of an RRAM-based neuromorphic chip faces many constraints due to the variability and limitations of RRAM. Simulation and emulation can help in the design of a neuromorphic chip prior to fabrication. However, software-based chip simulation on CPU is slow, especially for large-scale network-on-chip (NoC)-based chip design. In this paper, we present a hardware emulator on field-programmable gate array (FPGA) for an RRAM-based neuromorphic chip. Our emulator supports the emulation of static and dynamic variation of the RRAM-based crossbars used in the neural cores of a neuromorphic chip. Furthermore, an NoC is also implemented on FPGA to emulate the communication between the neural cores. Using the emulator, we show that effects, such as RRAM write and read noise and stuck-at faults affect the accuracy of an application on a neuromorphic chip. We also demonstrate the utility of the emulator in investigating NoC topologies, routing buffer depths, and neural core mappings. Tao Luo 0014, Chuping Qu, Matthew Kay Fei Lee, Wai Teng Tang, Weng-Fai Wong, Rick Siow Mong Goh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |