VLDB 2026 Research / reviewers in the wild / expert
Cheng Chu
dblp:274/0580
· DBLP profile ↗
15ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0002-3226-0750ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QNBAD: Quantum Noise-induced Backdoor Attacks against Zero Noise Extrapolation
Cheng Chu, Qian Lou, Fan Chen 0001, Lei Jiang 0001 |
NDSS | 1 |
| 2025 | LSTM-QGAN: Scalable NISQ Generative Adversarial NetworkabstractCurrent quantum generative adversarial networks (QGANs) still struggle with practical-sized data. First, many QGANs use principal component analysis (PCA) for dimension reduction, which, as our studies reveal, can diminish the QGAN’s effectiveness. Second, methods that segment inputs into smaller patches processed by multiple generators face scalability issues. In this work, we propose LSTM-QGAN, a QGAN architecture that eliminates PCA preprocessing and integrates quantum long short-term memory (QLSTM) to ensure scalable performance. Our experiments show that LSTM-QGAN significantly enhances both performance and scalability over state-of-the-art QGAN models, with visual data improvements, reduced Fréchet Inception Distance scores, and reductions of 5× in qubit counts, 5× in single-qubit gates, and 12× in two-qubit gates. Cheng Chu, Aishwarya Hastak, Fan Chen 0001 |
ICASSP | 1 |
| 2024 | TITAN: A Fast and Distributed Large-Scale Trapped-Ion NISQ ComputerabstractTrapped-Ion (TI) technology offers potential breakthroughs for Noisy Intermediate Scale Quantum (NISQ) computing. TI qubits offer extended coherence times and high gate fidelity, making them appealing for large-scale NISQ computers. Constructing such computers demands a distributed architecture connecting Quantum Charge Coupled Devices (QCCDs) via quantum matter-links and photonic switches. However, current distributed TI NISQ computers face hardware and system challenges. Entangling qubits across a photonic switch introduces significant latency, while existing compilers generate suboptimal mappings due to their unawareness of the interconnection topology. In this paper, we introduce TITAN, a large-scale distributed TI NISQ computer, which employs an innovative photonic interconnection design to reduce entanglement latency and an advanced partitioning and mapping algorithm to optimize matter-link communications. Our evaluations show that TITAN greatly enhances quantum application performance by 56.6% and fidelity by 19.7% compared to existing systems. Cheng Chu, Zhenxiao Fu, Hausi A. Müller, Fan Chen 0001, Lei Jiang 0001 |
DAC | 1 |
| 2024 | QuantumLeak: Stealing Quantum Neural Networks from Cloud-based NISQ MachinesabstractVariational quantum circuits (VQCs) have become a powerful tool for implementing Quantum Neural Networks (QNNs), addressing a wide range of complex problems. Well-trained VQCs serve as valuable intellectual assets hosted on cloud-based Noisy Intermediate Scale Quantum (NISQ) computers, making them susceptible to malicious VQC stealing attacks. However, traditional model extraction techniques designed for classical machine learning models encounter challenges when applied to NISQ computers due to significant noise in current devices. In this paper, we introduce QuantumLeak, an effective and accurate QNN model extraction technique from cloud-based NISQ machines. Compared to existing classical model stealing techniques, QuantumLeak improves local VQC accuracy by 4.99%~7.35% across diverse datasets and VQC architectures. Zhenxiao Fu, Cheng Chu, Fan Chen 0001 |
IJCNN | 3 |
| 2024 | OFHE: An Electro-Optical Accelerator for Discretized TFHEabstractThis paper presents OFHE, an electro-optical accelerator designed to process Discretized TFHE (DTFHE) operations, which encrypt multi-bit messages and support homomorphic multiplications, lookup table operations and full-domain functional bootstrappings. While DTFHE is more efficient and versatile than other fully homomorphic encryption schemes, it requires 32-, 64-, and 128-bit polynomial multiplications, which can be time-consuming. Existing TFHE accelerators are not easily upgradable to support DTFHE operations due to limited datapaths, a lack of datapath bit-width reconfigurability, and power inefficiencies when processing FFT and inverse FFT (IFFT) kernels. Compared to prior TFHE accelerators, OFHE addresses these challenges by improving the DTFHE operation latency by 8.7%, the DTFHE operation throughput by 57%, and the DTFHE operation throughput per Watt by 94%. Mengxin Zheng, Cheng Chu, Qian Lou, Nathan Youngblood, Sajjad Moazeni, Lei Jiang 0001 |
ISLPED | 2 |
| 2023 | QTROJAN: A Circuit Backdoor Against Quantum Neural NetworksabstractWe propose a circuit-level backdoor attack, QTrojan, against Quantum Neural Networks (QNNs) in this paper. QTrojan is implemented by a few quantum gates inserted into the variational quantum circuit of the victim QNN. QTrojan is much stealthier than a prior Data-Poisoning-based Backdoor Attack (DPBA) since it does not embed any trigger in the inputs of the victim QNN or require access to original training datasets. Compared to a DPBA, QTrojan improves the clean data accuracy by 21% and the attack success rate by 19.9%. Cheng Chu, Lei Jiang 0001, D. Martin Swany, Fan Chen 0001 |
ICASSP | 1 |
| 2023 | IQGAN: Robust Quantum Generative Adversarial Network for Image Synthesis On NISQ DevicesabstractIn this work, we propose IQGAN, a quantum Generative Adversarial Network (GAN) framework for multiqubit image synthesis that can be efficiently implemented on Noisy Intermediate Scale Quantum (NISQ) devices. We investigate the reasons for the inferior generative performance of current quantum GANs in our preliminary study and conclude that an adjustable input encoder is the key to ensuring high-quality data synthesis. We then propose the IQGAN architecture featuring a trainable multiqubit quantum encoder that effectively embeds classical data into quantum states. Furthermore, we propose a compact quantum generator that significantly reduces the design cost and circuit depth on NISQ devices. Experimental results on both IBM quantum processors and quantum simulators demonstrated that IQGAN outperforms state-of-the-art quantum GANs in qualitative and quantitative evaluation of the generated samples, model convergence, and quantum computing cost. Cheng Chu, Grant Skipper, D. Martin Swany, Fan Chen 0001 |
ICASSP | 1 |
| 2023 | Accelerating Deformable Convolution Networks with Dynamic and Irregular Memory AccessesabstractDeformable convolution networks (DCNs) proposed to address image recognition with geometric or photometric variations typically involve deformable convolution that convolves on arbitrary locations of input features. The locations change with different inputs and induce considerable dynamic and irregular memory accesses that cannot be handled by classic neural network accelerators (NNAs). Moreover, bilinear interpolation (BLI) operation, which is required to obtain deformed features in DCNs, also cannot be deployed on existing NNAs directly. Although a general purposed processor (GPP) seated along with classic NNAs can process the deformable convolution, the processing on GPP can be extremely slow due to the limited parallel computing capability and massive additional data movement. To address the problem, we develop a DCN accelerator on existing NNAs to support both the standard convolution and deformable convolution. Specifically, for the dynamic and irregular accesses in DCNs, we have both the input and output features divided into tiles and build a tile dependency table (TDT) to track the irregular tile dependency at runtime. With the TDT, we further develop an on-chip tile scheduler to handle the dynamic and irregular accesses efficiently. In addition, we propose a novel mapping strategy to enable parallel BLI processing on NNAs and apply layer fusion techniques for more energy-efficient DCN processing. According to our experiments, the proposed accelerator achieves orders of magnitude higher performance and energy efficiency compared to the typical computing architectures including ARM, ARM+TPU, and GPU with 6.6% chip area penalty to a classic NNA. Cheng Chu, Cheng Liu 0008, Dawen Xu 0002, Ying Wang 0001, Tao Luo 0014, Huawei Li 0001, Xiaowei Li 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2022 | MOCCA: A Process Variation Tolerant Systolic DNN Accelerator using CNFETs in Monolithic 3DabstractHardware accelerators based on systolic arrays have become the dominant method for efficient processing of deep neural networks (DNNs). Although such designs provide significant performance improvement compared to its contemporary CPUs or GPUs, their power efficiency and area efficiency are greatly limited by the large computing array and on-chip memory. In this work, we demonstrate that we can further improve the efficiency of systolic accelerators using emerging carbon nanotube field-effect transistors (CNFETs) by stacking the computing logic and on-chip memory on multiple layers and utilizing monolithic 3D (M3D) vias for low-latency communication. We comprehensively explore the design space and present MOCCA, the first process variation tolerable CNFET-based systolic DNN accelerator. We validate MOCCA against previous 2D accelerators on state-of-the-arts DNN models. On average, MOCCA achieves the same throughput with 6.12× and 2.12× improvement respectively on performance and power efficiency in a 2× reduced chip footprint. Samuel J. Engers, Cheng Chu, Dawen Xu 0002, Ying Wang 0001, Fan Chen 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | Canopy: A CNFET-based Process Variation Aware Systolic DNN AcceleratorabstractAlthough systolic accelerators have become the dominant method for executing Deep Neural Networks (DNNs), their performance efficiency (quantified as Energy-Delay Product or EDP) is limited by the capabilities of silicon Field-Effect Transistors (FETs). FETs constructed from Carbon Nanotubes (CNTs) have demonstrated > 10 × EDP benefits, however, the processing variations inherent in carbon nanotube FETs (CNFETs) fabrication compromise the EDP benefits, resulting > 40% performance degradation. In this work, we study the impact of CNT process variations and present Canopy, a process variation aware systolic DNN accelerator by leveraging the spatial correlation in CNT variations. Canopy co-optimizes the architecture and dataflow to allow computing engines in a systolic array run at their best performance with non-uniform latency, minimizing the performance degradation incurred by CNT variations. Furthermore, we devise Canopy with dynamic reconfigurability such that the microarchitectural capability and its associated flexibility achieves an extra degree of adaptability with regard to the DNN topology and processing hyper-parameters (e.g., batch size). Experimental results show that Canopy improves the performance by 5.85 × (4.66 ×) and reduces the energy by 34% (90%) when inferencing a single (a batch of) input compared to the baseline design under an iso-area comparison across seven DNN workloads. Cheng Chu, Dawen Xu 0002, Ying Wang 0001, Fan Chen 0001 |
ISLPED | 1 |
| 2022 | QMLP: An Error-Tolerant Nonlinear Quantum MLP Architecture using Parameterized Two-Qubit GatesabstractDespite potential quantum supremacy, state-of-the-art quantum neural networks (QNNs) suffer from low inference accuracy. First, the current Noisy Intermediate-Scale Quantum (NISQ) devices with high error rates of 10− 3 to 10− 2 significantly degrade the accuracy of a QNN. Second, although recently proposed Re-Uploading Units (RUUs) introduce some non-linearity into the QNN circuits, the theory behind it is not fully understood. Furthermore, previous RUUs that repeatedly upload original data can only provide marginal accuracy improvements. Third, current QNN circuit ansatz uses fixed two-qubit gates to enforce maximum entanglement capability, making task-specific entanglement tuning impossible, resulting in poor overall performance. In this paper, we propose a Quantum Multilayer Perceptron (QMLP) architecture featured by error-tolerant input embedding, rich nonlinearity, and enhanced variational circuit ansatz with parameterized two-qubit entangling gates. Compared to prior arts, QMLP increases the inference accuracy on the 10-class MNIST dataset by 10% with 2 × fewer quantum gates and 3 × reduced parameters. Our source code is available and can be found in https://github.com/chuchengc/QMLP/. Cheng Chu, Nai-Hui Chia, Lei Jiang 0001, Fan Chen 0001 |
ISLPED | 1 |
| 2022 | HyCA: A Hybrid Computing Architecture for Fault-Tolerant Deep LearningabstractHardware faults on the regular 2-D computing array of a typical deep learning accelerator (DLA) can lead to dramatic prediction accuracy loss. Prior redundancy design approaches typically have each homogeneous redundant processing element (PE) to mitigate faulty PEs for a limited region of the 2-D computing array rather than the entire computing array to avoid the excessive hardware overhead. However, they fail to recover the computing array when the number of faulty PEs in any region exceeds the number of redundant PEs in the same region. The mismatch problem deteriorates when the fault injection rate rises and the faults are unevenly distributed. To address the problem, we propose a hybrid computing architecture (HyCA) for fault-tolerant DLAs. It has a set of dot-production processing units (DPPUs) to recompute all the operations that are mapped to the faulty PEs despite the faulty PE locations. According to our experiments, HyCA shows significantly higher reliability, scalability, and performance with less chip area penalty when compared to the conventional redundancy approaches. Moreover, by taking advantage of the flexible recomputing, HyCA can also be utilized to scan the entire 2-D computing array and detect the faulty PEs effectively at runtime. Cheng Liu 0008, Cheng Chu, Dawen Xu 0002, Ying Wang 0001, Huawei Li 0001, Xiaowei Li 0001, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | RECOIN: A Low-Power Processing-in-ReRAM Architecture for Deformable ConvolutionabstractThe recent proposed Deformable Convolutional Networks (DCNs)greatly enhance the performance of conventional Convolutional Neural Networks (CNNs) on vision recognition tasks by allowing flexible input sampling during inference runtime. DCNs introduce an additional convolutional layer for adaptive sampling offset generation, followed by a bilinear interpolation (BLI) algorithm to integerize the generated non-integer offset values. Finally, a regular convolution is performed on the loaded input pixels. Compared with conventional CNNs, DCN demonstrated significantly increased computational complexity and irregular input-dependentmemory access patterns, making it a great challenge for deploying DCNs onto edge devices for real-time computer vision tasks. In this work, we propose RECOIN, a processing-in-memory (PIM) architecture, which supports DCN inference on resistive memory (ReRAM)crossbars, thus making the first DCN inference accelerator possible. We present a novel BLI processing engine that leverage both row-and column-oriented computation for in-situ BLI calculation. Amapping scheme and an address converter are particular designed to accommodate the intensive computation and irregular data access. We implement the DCN inference in a 4-stage pipeline and evaluate the effectiveness of RECOIN on six DCN models. Experimental results show RECOIN achieves respectively 225×and 17.4×improvement in energy efficiency compared to general-purpose CPU and GPU. Compared to two state-of-the-art ASIC accelerators, RECOIN achieve 26.8× and 20.4× speedup respectively. Cheng Chu, Fan Chen 0001, Dawen Xu 0002, Ying Wang 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | Multi-task Scheduling for PIM-based Heterogeneous Computing SystemabstractProcessing-in-Memory (PIM) or Near-Data Processing has been recognized as the most potential solution to resolve the ever-aggravating memory wall especially as the thrive of memory-intensive scale-out workloads such as graph computing and data analytics. However, when the future computing system becomes more and more likely to adopt PIM architectures as a type of the storage and processing component, there is a lack of literature and research work on the general scheduling framework with the emerging heterogeneous system except for some ad-hoc task partitioning methods with specialized PIM designs. This work is the first to propose a formalized model to quantitatively describe the multi-task scheduling problem in PIM+CPU platform without loss of generality, and also an optimized task mapping-and-scheduling algorithm to boost the hardware utility for these novel heterogeneous systems. The proposed scheduling framework is fully aware of the data access bandwidth and processing capability distinction between the CPU and PIM devices, and also the implications of task mapping on the bandwidth contention, data communication intensity and hardware utility for the concurrent workloads. Experimental results show that, compared to the traditional scheduling algorithm for heterogeneous system, the proposed method is able to improve the system performance by over 10% and the energy efficiency by almost 10% for multi-core scale-out applications. Dawen Xu 0002, Cheng Chu, Cheng Liu 0008, Ying Wang 0001, Xianzhong Zhou, Lei Zhang 0008, Huaguo Liang, Huawei Li 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2020 | A Hybrid Computing Architecture for Fault-tolerant Deep Learning AcceleratorsabstractRegular 2D computing array is widely utilized for the processing of the major neural network operations in many deep learning accelerators (DLAs). Hardware failures on the array can lead to considerable computing errors and prediction accuracy loss. Prior works proposed to add homogeneous redundant PEs to each row or column of the regular computing array to mitigate faulty PEs, but they may fail to recover the computing array from faults when the number of faulty PEs in a row or column exceeds the number of redundant PEs in the corresponding row or column. The problem gets worse when the faults are not evenly distributed across the computing array. To address the problem, we propose a hybrid computing architecture (HCA) for fault-tolerant DLAs. Instead of adding homogeneous redundant PEs to the regular computing array of DLAs, it has a dot-production processing unit (DPPU) to recompute the operations that are mapped to the faulty PEs concurrently without performance penalty under moderate fault injection. Even under high fault injection, HCA can be degraded smoothly and remains functional. In addition, DPPU exploits the parallelism within each operation and processes the network operations sequentially, so it can tolerate faulty PEs in arbitrary locations and ensures steady performance under distinct fault distributions. According to our experiments, HCA shows significantly higher reliability and performance under various fault injection with comparable chip area penalty compared to the conventional redundancy approaches. Dawen Xu 0002, Cheng Chu, Cheng Liu 0008, Ying Wang 0001, Lei Zhang 0008, Huaguo Liang, Kwang-Ting Cheng |
ICCD | 2 |