EDBT 2026 Demo / reviewers in the wild / expert
Qianyu Cheng
dblp:295/9597
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0001-8321-6920ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Semantic Modeling for Glass Surface Detection in the WildabstractGlass surfaces challenge object detection models as they mix the transmitted background with the reflected surrounding, creating confusing visual patterns. Previous methods relying on low-level cues (e.g., reflections and boundaries) or surrounding semantics are often unreliable in complex real-world scenarios. A glass image inherently comprises three distinct semantic components: semantics of the transmitted content, semantics of the reflected content, and semantics of the surrounding content. In this work, we observe that there is a relationship among these three types of semantics, where reflection semantics closely resembles surrounding semantics, while these two types of semantics tend to be different from the transmission semantics. For example, when on a street, we may see into a cafeteria through a glass wall, intermixed with reflection of the street, while the glass is surrounded by other street contents like shops and pedestrians, thereby creating a unique multi-semantic signature. Based on this observation, we propose the Multi-Semantic Net, MSNet, which identifies transmission, reflection, and surrounding semantics from glass images and exploits their relationships for glass surface detection. MSNet consists of two novel modules: (1) A Semantic Decomposition Module (SDM) containing Dual-Semantics Extraction Block to extract original image and reflection semantics and Semantic Elimination Block to progressively derive transmission and surrounding semantics, and (2) An Adaptive Semantic Fusion Module (ASFM) to fuse these semantic components and adaptively learn their relationships to handle varying reflection conditions. Extensive experiments demonstrate that MSNet surpasses SOTA methods on public glass detection benchmarks. Qianyu Cheng, Huankang Guan, Rynson W. H. Lau |
AAAI | 1 |
| 2026 | HE-DeepFM: An FHE Inference System for CTR Prediction with Efficient FM InteractionsabstractScoring models such as click-through rate (CTR) prediction underpin recommendation and advertising systems, but their features are highly sensitive, making plaintext cloud inference risky. Fully homomorphic encryption (FHE) enables inference directly on ciphertexts, yet homomorphic computation is expensive and bootstrapping often dominates end-to-end latency. We present HE-DeepFM, a FHE inference system for CTR prediction. We first design HE-FM, a homomorphic-friendly Factorization Machine (FM) operator that exploits CKKS SIMD packing to compute second-order interactions efficiently, thereby avoiding the naive O(F2) cost of pairwise feature interactions. Building on HE-FM, HE-DeepFM reduces bootstrapping under the same FHE budget, and can eliminate it for small model configurations. We implement HE-DeepFM with Orion and Lattigo and evaluate it on real-world datasets. On Criteo, HE-DeepFM reduces bootstrapping from 12 to 4 and achieves up to 2.85× end-to-end speedup while maintaining prediction quality comparable to the baseline. Qiyue Su, Hang Gu, Zhiguang Wang, Zhendong Zheng, Qianyu Cheng, Lei Gong 0003, Chao Wang 0003 |
SIGIR | 6 |
| 2026 | Out-of-Memory Graph Processing Acceleration via Algorithmic-Hardware Codesign on FPGAsabstractEmerging FPGA devices are widely used in graph processing and achieve high performance by exploiting fine-grained parallelism and high-bandwidth memory (HBM) subsystems with dozens of channels. However, the graph size increases rapidly and often exceeds the capacity of the FPGA device or one of the memory channels, incurring out-of-memory (OoM) faults and performance degradation. Existing approaches focus on improving the performance of on-device full graph processing, ignoring the support for hyper-scale graphs and cross-iteration acceleration. In this work, we propose SoGraph, a graph processing architecture for large graph acceleration on HBM-equipped FPGAs, cooperating with the host. To minimize PCIe traffic during the processing, we adopt three hardware-oriented optimizations: 1) activeness-driven subgraph scheduler, which minimizes on-device graph resident set; 2) lightweight state-aware processing extension in edge-centric paradigm; 3) bit-level update analyzer for to-be-transfer subgraph reduction. The results show that the prototype on Alveo U280 FPGA achieves an average of 22.3x performance speedup over the modified state-of-the-art FPGA design and 1.3x device energy efficiency over the GPU solution. Qianyu Cheng, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Huaping Chen 0001, Xuehai Zhou |
IEEE Trans. Computers | 1 |
| 2026 | UniCoX: A Unified Cost Model for Tensorized Program Tuning Across Ubiquitous AcceleratorsabstractTensorized programs leverage hardware intrinsics on accelerators to boost tensor computation performance. With the rise of hardware customization, massive accelerators and intrinsics have emerged, posing engineering challenges for manually coding efficient tensorized programs. Tensorized program tuning in deep learning compilers (DLCs) is considered an effective approach. At the core of program tuning relies the design of the cost model to predict performance. However, there is currently a lack of cost models specifically for tensorized programs, which severely hinders the co-optimization of deep learning compilers and hardware accelerators.In this paper, we propose UniCoX, a unified cost model specifically for tensorized program tuning across ubiquitous accelerators. We systematically analyze the design challenges introduced by tensorized programs from the perspectives of feature representation and transfer prediction. For feature representation, we leverage attention mechanisms to mine key software schedule features and design corresponding aligned hardware features, resulting in a unified cross-accelerator feature representation. For transfer prediction, by integrating lifelong learning and transfer learning with data sampling strategies, we propose a unified transfer prediction strategy to keep pace with the rapid development of accelerators. To meet training and testing demands, we construct TensorizeSetX, a dataset dedicated to tensorized program tuning. Results show that UniCoX achieves the state-of-the-art accuracy while supporting low-cost and flexible transfer prediction. It can accelerate search time by 11.3’ and improve inference speed by 1.9’ within the state-of-the-art tensorized program tuning framework, TVM MetaSchedule. Lei Gong 0003, Wenqi Lou, Qianyu Cheng, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Computers | 5 |
| 2026 | UniSparTa: A Unified Sparse Tensor Program Tuning FrameworkabstractSparse tensor computation is widely used in deep learning and scientific computing. However, diverse sparse data and algorithmic characteristics at the application level, combined with the diversity of hardware platforms, pose significant challenges for efficient sparse tensor program optimization. Manually crafted operator libraries are time-consuming to develop and lack portability. To address this, we propose UniSparTa, a unified sparse tensor program tuning framework that automatically generates high-performance programs. First, we extract unified optimization principles for high-performance sparse tensor programs and propose a domain-specific language (DSL) to automatically generate a high-quality design space without manual intervention. Second, by analyzing the general distribution of the design space, we introduce an adaptive search strategy combining Deep Q-Networks (DQN) and Simulated Annealing (SA). Finally, to avoid the unacceptable time cost of real measurement during tuning, we propose a unified cost model based on multimodal fusion to accurately predict program performance. Furthermore, by leveraging data augmentation and transfer learning, we enable low-cost transfer prediction across different sparse data patterns, algorithms, and hardware platforms. Results show that, compared to the state-of-the-art operator library MKL, the manually optimized scheme ASpT, the tensor compiler TVM, and the sparse tensor tuning framework WACO, UniSparTa achieves average speedups of 1.98×, 2.75×, 6.13×, and 1.75×, respectively. Moreover, UniSparTa significantly accelerates the tuning process. Lei Gong 0003, Xiangjun Qu, Cheng Tang 0004, Wenqi Lou, Qianyu Cheng, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2026 | LORA: A Latency-Oriented Recurrent Architecture for Large Language Model on Multi-FPGA Platform With Communication OptimizationabstractThe remarkable performance of Large Language Models (LLMs) has driven their widespread deployment in data centers to support diverse user-facing applications. However, the rapidly growing computational and storage demands of these models have made single-device deployment increasingly impractical. Prior research on LLM inference has primarily addressed this challenge through algorithmic optimizations such as quantization or by integrating customized hardware acceleration frameworks. As model parameters continue to scale, multi-device deployment has become a necessary approach for enabling efficient LLM inference. Nevertheless, constructing low-latency multi-device platforms for LLMs inference using available FPGA or GPU accelerators remains constrained by inefficient synchronization schemes or limited compute intensity in current architectures. Furthermore, existing solutions often lack co-optimized designs that effectively integrate communication with computation. To address these limitations, this paper proposes LORA, a low-latency end-to-end LLMs acceleration platform utilizing multiple FPGAs. Firstly, we optimize the synchronization timing within the LLMs to minimize storage, computation, and BRAM overhead. Secondly, we tightly couple communication and computation through techniques such as pipeline overlapping and input data packing. Next, we deploy homogeneous accelerators on each FPGA device, leveraging a recurrent architecture to further reduce inference latency. Finally, we apply FPGA-specific optimizations and conduct performance modeling and analysis of the acceleration framework to select optimal deployment parameters for various computational tasks. Implemented on Xilinx Alveo U280 FPGAs, LORA-F and LORA-Q achieve average speedups of 14.4× and 32.6×, respectively, compared to NVIDIA V100 GPUs when running modern LLMs. Compared with existing multi-FPGA accelerator platforms, LORA-F and LORA-Q demonstrate average performance improvements of up to 2.6× and 4.3×, respectively. Zhendong Zheng, Qianyu Cheng, Wenqi Lou, Lei Gong 0003, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Hermes: An FPGA-based NTT Accelerator Supporting Various Lengths for HHEabstractHybrid Homomorphic Encryption (HHE) scheme integrates two types of Fully Homomorphic Encryption (FHE), arithmetic FHE and logic FHE to enhance the performance and scalability of privacy-preserving computations. However, the performance of HHE mainly depends on the efficiency of the Number Theoretic Transform (NTT). Accordingly, this paper introduces Hermes, an FPGA-based NTT accelerator for HHE. We have designed a cross-scheme-friendly NTT architecture that supports NTT of varying lengths through the reuse of NTT units. Experimental results demonstrate that our proposed architecture achieves high hardware utilization and increases throughput by 1.3× compared to existing state-of-the-art approaches across various NTT lengths. Hang Gu, Qianyu Cheng, Jinao Li, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
CODES+ISSS | 3 |
| 2025 | Late Breaking Results: Source-Aware Adaptive Cache Management for CXL-enabled Disaggregated Memory SharingabstractDynamic workloads running on multiple hosts will bring changing access patterns on CXL-enabled shared disaggregated memory. Existing works often un-traceably cache multi-source accesses, making it hard to exploit each host’s access behavior and assure service quality. Our solution Alchemy jointly optimizes cache replacement and bypassing and runs as an online reinforcement learning agent with source-aware adaptivity. It gives rewards derived from sampling-based action effectiveness and per-host macro performance. The multi-host prototype-based results on FPGAs show $8.71 \%-14.56 \%$ reduction in average access latency over LRU policy and 44x faster than the hardware-efficient ICGMM method in decision-making with comparable overhead. Qianyu Cheng, Jiajun Ji, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
DAC | 1 |
| 2025 | Late Breaking Results: A Fast Nearest Neighbor Search Acceleration for 3D Point CloudabstractThis paper presents FastNN, a novel accelerator architecture for efficient K-Nearest Neighbors (KNN) search in point clouds. FastNN leverages a locality-sensitive E2LSH partitioning method and a precomparator module to significantly reduce the candidate search space and minimize the number of Euclidean distance calculations. Compared to octree-based partitioning methods, our approach reduces candidate points by 58.57% to 86.17% and achieves a $10.04 \times$ acceleration in processing throughput relative to the BitNN comparator subsystem. The proposed design effectively enhances search throughput, resource utilization, and precision, highlighting its potential for accelerating KNN search on FPGA platforms. Jinao Li, Qianyu Cheng, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
DAC | 3 |
| 2025 | Spectral Enhanced Tuning: An Efficient Plug-and-Play Framework for Frequency-Aware DehazingabstractImage dehazing aims to recover clear images from hazy counterparts; however, current deep learning methods tend to rely on low-frequency information while neglecting multi-frequency integration, limiting their performance. To address this, we introduce Spectral Enhanced Tuning (SET), a modular framework composed of lightweight, plug-and-play components that can be easily integrated into existing dehazing networks to improve performance via efficient multi-frequency feature fusion. The framework features a Frequency Spectrum Decoder (FSD) that utilizes localized windows for high-frequency details and average pooling for low-frequency global features, improving efficiency over traditional global processing. Additionally, the Dual-Domain Guided Attention (DDGA) computes importance maps across frequency and spatial domains, surpassing simple concatenation or addition, while the Direction-Aware Convolution Module (DAConv) captures spatial details for better reconstruction. Experimental results demonstrate that SET achieves significant dehazing performance gains with negligible increases in parameters and computational cost, validating its efficiency and effectiveness. Cheng Tang 0004, Wenqi Lou, Qianyu Cheng, Jiayi Tuo, Tianhao Jiang, Chao Wang 0003, Xuehai Zhou |
ICME | 3 |
| 2024 | Ph.D. Project: Optimizing the Data Traffic for Large Graph Processing on FPGA via a Stateful ApproachabstractReal-world graph size increases rapidly and often exceeds the capacity of the FPGA device, incurring memory over- subscription and performance degradation. Existing approaches ignore the support for hyper-scale graphs and cross-iteration acceleration. In this work, we propose SoGraph, a processing architecture for out-of-FPGA-memory graph acceleration on FPGAs, cooperating with the host. To minimize data transfer during the processing,$w$e adopt a lightweight stateful hardware extension and an activeness-driven subgraph scheduler in the edge-centric processing pipeline for graph working set pruning. A preliminary experiment shows that our prototype achieves 5.63x performance speedup over the modified state-of-the-art design. Qianyu Cheng, Chao Wang 0003 |
FCCM | 1 |
| 2024 | SoGraph: A State-Aware Architecture for Out-of-Memory Graph Processing on HBM-Equipped FPGAsabstractEmerging FPGA devices are widely used in graph processing and achieve high performance by exploiting fine-grained parallelism and high-bandwidth memory (HBM) sub-systems with dozens of channels. However, the graph size increases rapidly and often exceeds the capacity of the FPGA device or one of the memory channels, incurring out-of-memory (OoM) faults and performance degradation. Existing approaches focus on improving the performance of on-device full graph processing, ignoring the support for hyper-scale graphs and cross-iteration acceleration. In this work, we propose SoGraph, a graph processing architecture for large graph acceleration on HBM-equipped FPGAs, cooperating with the host. To minimize PCIe traffic during the processing, we adopt three hardware-oriented optimizations: 1) activeness-driven subgraph scheduler, which minimizes on-device graph resident set; 2) lightweight state-aware processing extension in edge-centric paradigm; 3) bit-level update analyzer for to-be-transfer subgraph reduction. The results show that the prototype on Alveo U280 FPGA achieves an average of 3.18x performance speedup over the modified state-of-the-art FPGA design and 1.3x energy efficiency over the GPU solution. Qianyu Cheng, Zhendong Zheng, Tianhao Jiang, Cheng Tang 0004, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
FPL | 1 |
| 2024 | LORA: A Latency-Oriented Recurrent Architecture for GPT Model on Multi-FPGA Platform with Communication OptimizationabstractLarge Language Models (LLMs) have been widely deployed in data centers to provide various services, among which the most representative is the Generative Pre-trained Transformer (GPT). The GPT model has heavy memory and computing overhead, and its inference process has two stages with distinct computing characteristics: Prefill and Decode. Utilizing existing GPUs and FPGA accelerators to construct a platform for deploying GPT in data centers faces the challenges of needing more effective synchronization schemes or structures with higher computational intensity. This paper proposes LORA, a low latency end-to-end GPT acceleration platform utilizing multiple FPGAs. Firstly, we optimize the synchronization timing of the GPT model to reduce the computation and communication overhead. Secondly, we devise some efficient synchronization steps for specific layers of the GPT model that overlap part of the computation and communication delay to improve the latency of our platform. Finally, we deploy recurrent structures on each FPGA to accelerate the different stages of the GPT model. Implemented on the Xilinx Alveo U280 FPGAs, LORA achieves an average $11.1 \times$ speedup over NVIDIA V100 GPUs on the modern GPT-2 model. Compared to the existing multi-FPGA accelerator appliance, LORA shows performance improvements of up to $4 \times$ and $2.7 \times$ in the Prefill and Decode stages. Zhendong Zheng, Qianyu Cheng, Lei Gong 0003, Xianglan Chen, Cheng Tang 0004, Chao Wang 0003, Xuehai Zhou |
FPL | 2 |
| 2024 | AutoSparse: A Source-to-Source Format and Schedule Auto- Tuning Framework for Sparse Tensor ProgramabstractSparse tensor computation plays a crucial role in modern deep learning workloads, and its expensive computational cost leads to a strong demand for high-performance oper-ators. However, developing high-performance sparse operators is exceptionally challenging and tedious. Existing vendor operator libraries fail to keep pace with the evolving trends in new algorithms. Sparse tensor compilers simplify the development and optimization of operator, but existing work either requires significant engineering effort for tuning or suffers from limitations in search space and search strategies, which creates unavoidable cost and efficiency issues. In this paper, we propose AutoSparse, a source-to-source auto-tuning framework that targets sparse for-mat and schedule for sparse tensor program. Firstly, AutoSparse designs a sparse tensor DSL based on dynamic computational graph at the front-end, and proposes a sparse tensor program computational pattern extraction and automatic design space generation scheme based on it. Second, AutoSparse's back-end designs an adaptive exploration strategy based on reinforcement learning and heuristic algorithm to find the optimal format and schedule configuration in a large-scale design space. Compared to prior work, developers using AutoSparse do not need to specify tuning design space relied on any compilation or hardware knowledge. We use the SuiteS parse dataset to compare with four state-of-the-art baselines, namely, the high-performance operator library MKL, the manually-based optimisation scheme ASpT, the auto-tuning-based framework TVM-S and WACO. The results demonstrate that AutoSparse achieves average speedups of 1.92-$2.48 \times. 1.19-6.34 \times$. and$1.47-2.23\times$for the SpMV, SpMM, and SDDMM operators, respectively. We will open-source AutoSparse at https://github.com/Qu-Xiangjun/AutoSparse. Xiangjun Qu, Lei Gong 0003, Wenqi Lou, Qianyu Cheng, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
ICCD | 4 |
| 2024 | UniCoMo: A Unified Learning-Based Cost Model for Tensorized Program TuningabstractTensorized programs use hardware intrinsics on accelerators to significantly improve tensor computation performance. The trend of hardware customization has led to the emergence of massive hardware accelerators and intrinsics, which poses a significant engineering challenge for manually coding efficient tensorized programs. Tensorized program tuning in deep learning compilers (DLCs) is considered an effective approach to address the challenge. At the core of program tuning relies the design of the cost model, but currently there is still a lack of cost models specifically designed for tensorized programs, which hampers the co-optimization of DLCs and hardware accelerators. In this paper, we propose UniCoMo, a unified cost model for tensorized program tuning on various hardware platforms and hardware intrinsics. We first analyze the design challenges introduced by tensorized programs for cost models in terms of feature representation and transfer prediction. And then, for feature representation, we propose a unified feature representation for tensorized programs by using program behavior as a template, mining program features with the schedule attention matrix, and incorporating hardware intrinsic abstraction. For transfer prediction, we propose a unified transfer prediction strategy for tensorized program cost model based on lifelong learning and transfer learning. To meet training and testing requirements, we constructed a dataset dedicated to tensorized program tuning. Results show that UniCoMo maintains the state-of-the-art accuracy while significantly improving adaptability to diverse execution environments and enabling flexible transfer prediction. It can speed up search time by 9.8× and improve inference speed by 1.9× open-sourced at https://github.com/ZhW-loop/UniCoMo. Lei Gong 0003, Wenqi Lou, Qianyu Cheng, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
ICCD | 4 |