EDBT 2026 Demo / reviewers in the wild / expert
Yuncheng Lu
dblp:237/3613
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-0465-942XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hardware-Efficient Union-Find Decoder Towards Scalable Topological Quantum Codes
Shuang Liang 0012, Jubo Xu, Yuncheng Lu, Hao Mark Chen, Hongxiang Fan |
ASP-DAC | 3 |
| 2026 | Advancing Full-Stack Acceleration for SchröDinger-Style Quantum SimulationabstractRecent developments in quantum hardware, including the scaling of physical qubits and advanced quantum error correction techniques, have increased the number of reliable logical qubits. However, this progress has introduced new challenges for quantum algorithm developers. Limited access to physical quantum machines and the insufficient performance of classical quantum simulators for near-term scales ($\sim 30$logical qubits) hinder the simulation and validation of quantum algorithms. To address this urgent need for improving simulation performance, we propose a novel end-to-end full-stack solution for Schrödingerstyle simulation that jointly explores algorithm, software, and hardware optimizations. At the algorithmic level, by identifying the inefficiency in executing complex signed permutations and complex unitary permutation gates, we introduce index redirection and pre-compute merging that significantly reduce data movement and computational complexity. At the hardware level, we propose a reconfigurable dataflow architecture with adaptive memory scheduling and swapping optimizations. At the software level, an end-to-end toolchain is introduced to jointly explore both algorithmic and hardware optimizations. A comprehensive evaluation across a large suite of quantum circuits demonstrates that our work achieves a maximum speedup exceeding$50 \times$over the GPU-based Qiskit baseline. Shuang Liang 0012, Yuncheng Lu, Ce Guo 0002, Paul H. J. Kelly, Wayne Luk, Hongxiang Fan |
HPCA | 2 |
| 2026 | Coset Ensemble Decoder for Quantum Error Correction with Algorithm-Hardware Co-Design
Shuang Liang 0012, Jubo Xu, Giulio Bassanino, Qianzhou Wang, Yuncheng Lu, Zhiwen Mo, Paul H. J. Kelly, Wayne Luk, Hongxiang Fan |
ISCA | 6 |
| 2026 | A Real-Time End-to-End Event-Based Tactile Sensing SystemabstractEvent-based tactile sensing offers a promising alternative to frame-based approaches by reducing data redundancy, yet existing systems often lack end-to-end hardware support and remain power-inefficient. This work presents an energy-efficient event-based tactile sensing system that codesigns algorithms and hardware for real-time perception. At the front end, a leakage-compensated event-driven readout circuit integrates multiple tactile sensors into a single node, minimizing wiring and static power. At the algorithm level, a 2-D convolutional neural network (CNN) reconstructs event frames for handwritten digit recognition with strong robustness to varying interaction durations, while a graph neural network (GNN) processes irregular tactile layouts for object classification. A customized event-based tactile processor (ETP) supports end-to-end tactile processing, incorporating a dual-mode event frame builder (EFB) and a neural network processing unit (NPU) with an optimized data-reuse scheme. Evaluated using 28-nm CMOS, the ASIC performs handwritten digit recognition and object classification in 0.44 and 0.37 ms, respectively. The system consumes the total power of 5.2 mW with only 0.11-mW dynamic power, demonstrating an efficient solution for always-on tactile perception. Yuncheng Lu, Kiho Seong, Si En Timothy Ng, Shibi Varku, Arindam Basu, Nripan Mathews, Tony Tae-Hyoung Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | Versatile Cross-platform Compilation Toolchain for Schrödinger-style Quantum Circuit SimulationabstractWhile existing quantum hardware resources have limited availability and reliability, there is a growing demand for exploring and verifying quantum algorithms. Efficient classical simulators for high-performance quantum simulation are critical to meeting this demand. However, due to the vastly varied characteristics of classical hardware, implementing hardware-specific optimizations for different hardware platforms is challenging. To address such needs, we propose CAST (Cross-platform Adaptive Schrödinger-style Simulation Toolchain), a novel compilation toolchain with cross-platform (CPU and Nvidia GPU) optimization and high-performance backend supports. CAST exploits a novel sparsity-aware gate fusion algorithm that automatically selects the best fusion strategy and backend configuration for targeted hardware platforms. CAST also aims to offer versatile and high-performance backend for different hardware platforms. To this end, CAST provides an LLVM IR-based vectorization optimization for various CPU architectures and instruction sets, and a PTX-based code generator for Nvidia GPU support. We benchmark CAST against IBM Qiskit, Google QSimCirq, Nvidia cuQuantum backend, and other high-performance simulators. On various 32-qubit CPU-based benchmarks, CAST achieves up to 8.03x speedup than Qiskit. On various 30-qubit GPU-based benchmarks, CAST achieves up to 39.3x speedup than Nvidia cuQuantum backend. Yuncheng Lu, Shuang Liang 0012, Hongxiang Fan, Ce Guo 0002, Wayne Luk, Paul H. J. Kelly |
DAC | 1 |
| 2025 | QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear OperationsabstractTransformer-based models have revolutionized computer vision (CV) and natural language processing (NLP) by achieving state-of-the-art performance across a range of benchmarks. However, nonlinear operations in models significantly contribute to inference latency, presenting unique challenges for efficient hardware acceleration. To this end, we propose QUARK, a quantization-enabled FPGA acceleration framework that leverages common patterns in nonlinear operations to enable efficient circuit sharing, thereby reducing hardware resource requirements. QUARK targets all nonlinear operations within Transformer-based models, achieving high-performance approximation through a novel circuit-sharing design tailored to accelerate these operations. Our evaluation demonstrates that QUARK significantly reduces the computational overhead of nonlinear operators in mainstream Transformer architectures, achieving up to a 1.96× end-to-end speedup over GPU implementations. Moreover, QUARK lowers the hardware overhead of nonlinear modules by more than 50% compared to prior approaches, all while maintaining high model accuracy—and even substantially boosting accuracy under ultra-low-bit quantization. Zhixiong Zhao, Haomin Li 0002, Fangxin Liu, Yuncheng Lu, Zongwu Wang, Tao Yang 0031, Li Jiang 0002, Haibing Guan |
ICCAD | 4 |
| 2025 | A 2.793 μW Near-Threshold Neuronal Population Dynamics Trajectory Filter for Reliable Simultaneous Localization and MappingabstractThis work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for the trajectory error correction task within a simultaneous localization and mapping workflow. A custom discretized procedural algorithm approximating a neuronal population dynamics-based inference operation is developed for mapping onto an ultra-lightweight digital macro featuring massively parallel in-situ processing techniques. Fabricated using a 40nm technology, the test chip features a$22\times 22$neuron array with 0.1358mm2 core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming sub-10-$\mu $W power. Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0096, Tony Tae-Hyoung Kim, Yuanjin Zheng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | PCQ: Parallel Compact Quantum Circuit SimulationabstractSince quantum computers are not readily available, much quantum computing research such as quantum algorithm verification has to be conducted on classical computer platforms. While many quantum circuit simulators have been developed on CPUs and GPUs, the potential of FPGAs as a platform with parallel computing capabilities and high energy efficiency has not been fully explored. This paper describes a novel approach with two modes of data movement optimization for an FPGA-based parallel pipelined dataflow architecture targeting a compact computation format. A data decoupling method is adapted to partition computing tasks and data into non-interacting sub- sets, significantly reducing external data interaction overhead. The proposed approach shows significant promise in improving performance and energy efficiency compared with existing state vector based CPU, GPU, and FPGA implementations. Shuang Liang 0012, Yuncheng Lu, Ce Guo 0002, Wayne Luk, Paul H. J. Kelly |
FCCM | 2 |
| 2024 | Live Demonstration: Real-Time Object Detection & Classification System in IoT with Dynamic Neuromorphic Vision SensorsabstractIn this paper, we demonstrate an energy-efficient real-time object detection and classification system featuring a hybrid event-based frame generation pipeline and a background-removal region proposal algorithm. The event-based frame is generated by aggregating active events within a programmable time interval, generating an event-based binary image (EBBI). This approach enables the utilization of low-complexity algorithms for denoising and object detection. The background-removal region proposal algorithm reduces memory requirements and removes dynamic backgrounds, leading to better detection performance. The proposed system is demonstrated on Zynq-7000 FPGA device with a DAVIS346 sensor. Experimental results show that the proposed system achieves comparable detection accuracy while requiring significantly less computation than existing event-based trackers. Wenhao Lu, Yuncheng Lu, Junying Li, Yucen Shi, Yuanjin Zheng, Tony Tae-Hyoung Kim |
ISCAS | 3 |
| 2024 | An Energy-Efficient Object Detection System in IoT with Dynamic Neuromorphic Vision SensorsabstractNeuromorphic vision sensors (NVSs) mimic the function of the human visual system, with significant energy-saving potential in IoT-based object detection systems. Unlike conventional sensors, NVSs only generate asynchronous spiking events in response to changes in light intensity. However, the inherent noise generated by NVSs causes a degradation of detection performance. Moreover, an interested object usually occupies only a portion of the entire image frame. Therefore, a real-time, accurate event-based object detection system is needed to identify the region of interest (Rol) and leverage this spatial redundancy to reduce computational load in subsequent recognition modules. In this article, we present an energy-efficient real-time object detection system featuring a hybrid event-based frame generation pipeline and a background-removal region proposal algorithm. The event-based frame is generated by aggregating active events within a programmable time interval, generating an event-based binary image (EBBI). This approach enables the utilization of low-complexity algorithms for denoising and object detection. The background-removal region proposal algorithm reduces memory requirements and removes dynamic backgrounds, leading to better detection performance. The proposed system is demonstrated on Zynq-7000 FPGA device with a DAVIS346 sensor. Experimental results show that the proposed system achieves comparable accuracy while requiring significantly less computation than existing event-based trackers. Wenhao Lu, Yuncheng Lu, Junying Li, Yucen Shi, Yuanjin Zheng, Tony Tae-Hyoung Kim |
ISCAS | 3 |
| 2024 | A Memory-Efficient High-Speed Event-based Object Tracking SystemabstractDynamic vision sensors (DVS) have become prevalent in edge vision applications due to their low power and short latency attributes. However, current DVS-based object tracking systems suffer from high power consumption or long processing latency due to high computing intensity of the object detection algorithms. This paper proposes an energy-efficient object detection system through algorithm and hardware co-optimization. We design hardware-efficient denoising and region proposal (RP) algorithms to reduce on-chip memory usage and power consumption. Besides, the processing latency is dramatically reduced thanks to the less computing complexity. The devised algorithm is executed on a heterogeneous platform, with segments particularly sensitive to latency being accelerated via FPGA. An RP processor, supporting both parallel and systolic computing modes, is developed to facilitate the computing-intensive RP generation. Remarkably, the proposed system reduces the on-chip memory by 95.3% in contrast to traditional methods that employ connected component labeling. Moreover, the processing time per frame stands at 92.2 ms, marking a reduction of 82.4% compared to CPU-only operations. Yuncheng Lu, Kaixiang Cui, Yucen Shi, Junying Li, Wenhao Lu, Yuanjin Zheng, Tony Tae-Hyoung Kim |
ISCAS | 1 |
| 2024 | A 2.793µW Near-Threshold Neuronal Population Dynamics Simulator for Reliable Simultaneous Localization and MappingabstractThis work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for a component within the back-end of simultaneous localization and mapping. A custom discretized procedural algorithm including injection, finite difference update, activation, and inhibition to approximate neuronal population dynamics is developed for digital implementation. Fabricated using a 40nm technology, the test chip features a scalable neuron 22 × 22 array with 0.1358mm2core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming 2.793µW of power. Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0016, Tony Tae-Hyoung Kim, Yuanjin Zheng |
ISCAS | 6 |
| 2023 | Analysis on the inherent noise tolerance of feedforward network and one noise-resilient structure
Wenhao Lu, Zhengyuan Zhang 0002, Yuncheng Lu, Yuanjin Zheng |
Neural Networks | 5 |
| 2021 | A Multi-Functional 4T2R ReRAM Macro Enabling 2-Dimensional Access and Computing In-MemoryabstractThis paper presents a multi-functional resistive random access memory (ReRAM) macro using a novel 4T2R bit-cell. The proposed 4T2R ReRAM enables 2-dimensional (2D) memory access, which offers significant latency and energy reductions for many applications such as matrix operations. Besides non-volatile storage, the proposed 4T2R ReRAM can support two types of computing in-memory operations: ternary content-addressable memory (TCAM) and logic in-memory (LIM). Evaluations on various matrix operations show that the proposed 4T2R ReRAM with 2-D access capability can reduce memory access latency and energy by up to 88% and 82%, respectively compared with conventional 1T1R ReRAM. For TCAM, the proposed 4T2R bit-cell takes a smaller area than SRAM-based TCAM cell, while achieves a comparable search speed. For LIM, we propose an optimized LIM full adder (LIM- FA) that improves the delay and the power by 3.2* and 1.6*, respectively compared with prior LIM-FAs. Yuzong Chen 0001, Lu Lu 0013, Yuncheng Lu, Tony Tae-Hyoung Kim |
ISCAS | 3 |