EDBT 2026 Demo / reviewers in the wild / expert
Cheng Tang 0004
dblp:14/11534-4
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CloserToMe: A Unified Framework for Accurate and Transferable Latency Prediction Across Heterogeneous DevicesabstractHardware accelerators such as GPUs, NPUs, and FPGAs are essential to meeting AI’s computational demands. With the proliferation of heterogeneous devices across cloud and edge, various model optimization techniques adapt to diverse hardware characteristics through operator transformations and structural modifications. Accurate, efficient latency prediction enables rapid selection of optimal strategies across hardware backends. Many existing methods treat hardware as a black-box executor, directly regressing latency without explicitly modeling the intricate interactions between neural network (NN) structures and device-specific execution behaviors. To address these challenges, we introduce a new modeling perspective that captures the interaction between neural architectures and hardware execution. To capture device-specific characteristics, we propose two complementary modeling strategies. The Device Behavior Signature Selector (DBSel) characterizes hardware execution behavior by selectively probing a small set of representative architectures, forming a compact, workload-driven profile. In parallel, we construct capability vectors that capture the hierarchical memory of each device and compute characteristics, providing a structured abstraction of its architectural capacity. To unify both behavioral and structural views, we introduce the Hardware–Operation Dialogue Module (HODM), which models fine-grained interactions between neural operators and hardware properties. Together, these components empower CloserToMe to deliver accurate and transferable latency predictions across unseen and diverse platforms. Cheng Tang 0004, Guochong Sui, Wenqi Lou, Jiayi Tuo, Wenqian Xie, Yinkang Gao, Yixuan Zhu, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
AAAI | 1 |
| 2026 | MessToClean: Evidence-Grounded Structure-Preserving Reconstruction for Real-World Degraded Exam Paper ImagesabstractJiayi Tuo, Cheng Tang, Zihan Wang, Chenyue Zhou, Yao Li, Yanbiao Ma, Chao Wang, Wei Dai, Mingxuan Wang, Shitong Qin, Ziwei Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiayi Tuo, Cheng Tang 0004, Chenyue Zhou, Yanbiao Ma, Chao Wang 0003, Wei Dai 0015, Mingxuan Wang, Shitong Qin |
ACL (1) | 2 |
| 2026 | UniSparTa: A Unified Sparse Tensor Program Tuning FrameworkabstractSparse tensor computation is widely used in deep learning and scientific computing. However, diverse sparse data and algorithmic characteristics at the application level, combined with the diversity of hardware platforms, pose significant challenges for efficient sparse tensor program optimization. Manually crafted operator libraries are time-consuming to develop and lack portability. To address this, we propose UniSparTa, a unified sparse tensor program tuning framework that automatically generates high-performance programs. First, we extract unified optimization principles for high-performance sparse tensor programs and propose a domain-specific language (DSL) to automatically generate a high-quality design space without manual intervention. Second, by analyzing the general distribution of the design space, we introduce an adaptive search strategy combining Deep Q-Networks (DQN) and Simulated Annealing (SA). Finally, to avoid the unacceptable time cost of real measurement during tuning, we propose a unified cost model based on multimodal fusion to accurately predict program performance. Furthermore, by leveraging data augmentation and transfer learning, we enable low-cost transfer prediction across different sparse data patterns, algorithms, and hardware platforms. Results show that, compared to the state-of-the-art operator library MKL, the manually optimized scheme ASpT, the tensor compiler TVM, and the sparse tensor tuning framework WACO, UniSparTa achieves average speedups of 1.98×, 2.75×, 6.13×, and 1.75×, respectively. Moreover, UniSparTa significantly accelerates the tuning process. Lei Gong 0003, Xiangjun Qu, Cheng Tang 0004, Wenqi Lou, Qianyu Cheng, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | UniCoS: A Unified Neural and Accelerator Co-Search Framework for CNNs and ViTsabstractCurrent algorithm-hardware co-search works often suffer from lengthy training times and inadequate exploration of hardware design spaces, leading to suboptimal performance. This work introduces UniCoS, a unified framework for co-optimizing neural networks and accelerators for CNNs and Vision Transformers (ViTs). By introducing a novel training-free proxy that evaluates accuracy within seconds and a clustering-based algorithm for exploring heterogeneous dataflows, UniCoS efficiently navigates the design spaces of both architectures. Experimental results demonstrate that the solutions generated by UniCoS consistently surpass state-of-the-art (SOTA) methods (e.g., $3.54 \times$ energy-delay product (EDP) improvement with a $1.76 \%$ higher accuracy on ImageNet) while requiring notably reduced search time (up to $48 \times, \sim 3$ hours). The code is available at https://github.com/mine7777/Unicos.git. Wenqi Lou, Cheng Tang 0004, Hongbing Wen, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
DAC | 3 |
| 2025 | Spectral Enhanced Tuning: An Efficient Plug-and-Play Framework for Frequency-Aware DehazingabstractImage dehazing aims to recover clear images from hazy counterparts; however, current deep learning methods tend to rely on low-frequency information while neglecting multi-frequency integration, limiting their performance. To address this, we introduce Spectral Enhanced Tuning (SET), a modular framework composed of lightweight, plug-and-play components that can be easily integrated into existing dehazing networks to improve performance via efficient multi-frequency feature fusion. The framework features a Frequency Spectrum Decoder (FSD) that utilizes localized windows for high-frequency details and average pooling for low-frequency global features, improving efficiency over traditional global processing. Additionally, the Dual-Domain Guided Attention (DDGA) computes importance maps across frequency and spatial domains, surpassing simple concatenation or addition, while the Direction-Aware Convolution Module (DAConv) captures spatial details for better reconstruction. Experimental results demonstrate that SET achieves significant dehazing performance gains with negligible increases in parameters and computational cost, validating its efficiency and effectiveness. Cheng Tang 0004, Wenqi Lou, Qianyu Cheng, Jiayi Tuo, Tianhao Jiang, Chao Wang 0003, Xuehai Zhou |
ICME | 1 |
| 2024 | SoGraph: A State-Aware Architecture for Out-of-Memory Graph Processing on HBM-Equipped FPGAsabstractEmerging FPGA devices are widely used in graph processing and achieve high performance by exploiting fine-grained parallelism and high-bandwidth memory (HBM) sub-systems with dozens of channels. However, the graph size increases rapidly and often exceeds the capacity of the FPGA device or one of the memory channels, incurring out-of-memory (OoM) faults and performance degradation. Existing approaches focus on improving the performance of on-device full graph processing, ignoring the support for hyper-scale graphs and cross-iteration acceleration. In this work, we propose SoGraph, a graph processing architecture for large graph acceleration on HBM-equipped FPGAs, cooperating with the host. To minimize PCIe traffic during the processing, we adopt three hardware-oriented optimizations: 1) activeness-driven subgraph scheduler, which minimizes on-device graph resident set; 2) lightweight state-aware processing extension in edge-centric paradigm; 3) bit-level update analyzer for to-be-transfer subgraph reduction. The results show that the prototype on Alveo U280 FPGA achieves an average of 3.18x performance speedup over the modified state-of-the-art FPGA design and 1.3x energy efficiency over the GPU solution. Qianyu Cheng, Zhendong Zheng, Tianhao Jiang, Cheng Tang 0004, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
FPL | 4 |
| 2024 | LORA: A Latency-Oriented Recurrent Architecture for GPT Model on Multi-FPGA Platform with Communication OptimizationabstractLarge Language Models (LLMs) have been widely deployed in data centers to provide various services, among which the most representative is the Generative Pre-trained Transformer (GPT). The GPT model has heavy memory and computing overhead, and its inference process has two stages with distinct computing characteristics: Prefill and Decode. Utilizing existing GPUs and FPGA accelerators to construct a platform for deploying GPT in data centers faces the challenges of needing more effective synchronization schemes or structures with higher computational intensity. This paper proposes LORA, a low latency end-to-end GPT acceleration platform utilizing multiple FPGAs. Firstly, we optimize the synchronization timing of the GPT model to reduce the computation and communication overhead. Secondly, we devise some efficient synchronization steps for specific layers of the GPT model that overlap part of the computation and communication delay to improve the latency of our platform. Finally, we deploy recurrent structures on each FPGA to accelerate the different stages of the GPT model. Implemented on the Xilinx Alveo U280 FPGAs, LORA achieves an average $11.1 \times$ speedup over NVIDIA V100 GPUs on the modern GPT-2 model. Compared to the existing multi-FPGA accelerator appliance, LORA shows performance improvements of up to $4 \times$ and $2.7 \times$ in the Prefill and Decode stages. Zhendong Zheng, Qianyu Cheng, Lei Gong 0003, Xianglan Chen, Cheng Tang 0004, Chao Wang 0003, Xuehai Zhou |
FPL | 6 |
| 2024 | FlexWalker: An Efficient Multi-Objective Design Space Exploration Framework for HLS DesignabstractThe HLS toolchain effectively reduces the design complexity of FPGA hardware accelerators. However, in scenarios involving the multi-objective optimization of large-scale HLS designs, determining the knob configurations of Pareto design points remains a challenging task for designers. Our work re-evaluates the key factors affecting the efficiency of multiobjective design space exploration in HLS design and proposes an efficient framework named FlexWalker. It utilizes the upper confidence bound algorithm to organize various heterogeneous regression models for predicting the quality of HLS designs with different knob configurations in the design space and introduces a probability sampling algorithm and an elastic Pareto frontier to counteract the negative impact of regression model errors. Experimental results show that our work can stably eliminate over 90% of non-Pareto frontier design points in the tested HLS design space, effectively enhancing the efficiency of multiobjective design space exploration. Zheyuan Zou, Cheng Tang 0004, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
FPL | 2 |
| 2022 | TCL-Net: A Lightweight and Efficient Dehazing Network with Frequency-Domain Fusion and Multi-Angle Attention
Cheng Tang 0004, Wenqi Lou |
ACCV (4) | 1 |