Yanheng Lu

dblp:345/1999 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Flexible Zero-Shot Approach to Tone Mapping via Structure-Preserving Diffusion Models
abstract
With the prevalence of high dynamic range (HDR) imaging, tone mapping techniques, which convert HDR images to high-quality standard dynamic range (SDR) images for display, have become increasingly important. However, obtaining paired HDR and high-quality SDR images is almost impossible, posing challenges to learning-based tone mapping methods. To address this issue, we propose a zero-shot tone mapping framework without requiring any HDR training samples. Our approach decomposes images into two components: structural information and tonal information. A diffusion-based mapping model taking the structural information as input is first trained in the high-quality SDR domain, then transferred to the HDR domain that has less readily available training data for inference, leveraging the equivalent distribution of the structural information across both domains. To preserve the original image’s structure, we modify the reverse sampling process and explicitly incorporate the original structural information into the intermediate results. To improve the image details, we introduce a dual-control network, enabling different conditional inputs to control different scales of the output. Additionally, we devise a flexible tone adjustment strategy, with a bunch of novel loss functions to modify the trained score function dynamically during reverse sampling, allowing users to customize the style of the generated image according to their preference during testing. Initially designed for tone mapping, our model can be applied to various tasks including image fusion, exposure correction, dehazing, etc., without retraining. Experimental results demonstrate that our approach surpasses previous state-of-the-art methods, indicating that it can serve as an effective, flexible and versatile solution to various tone-mapping tasks. Source code is available at https://github.com/ZSDM-HDR/Zero-Shot-Diffusion-HDR.
Ruoxi Zhu, Shusong Xu, Peiye Liu, Yanheng Lu, Dimin Niu, Hongzhong Zheng, Yen-Kuang Chen, Ming-e Jing, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.5
2026 STEED: Space and Time-Efficient Encrypted Database Using FHE
abstract
In the era of Big Data, enterprises and individuals often upload databases to the cloud for storage and querying, which involves the risk of data leakage. Encrypted databases based on fully homomorphic encryption (FHE) theoretically solve the leakage problem, but the actual deployment of such encrypted databases faces the challenge of high economic costs. Cloud service providers charge for data transfer volume and computation time. Unfortunately, FHE is very expensive in both aspects, with more than five orders of magnitude deterioration compared to directly transmitting and computing plaintext. In this paper, we present STEED, a low-cost encrypted database that tackles both bottlenecks simultaneously. In STEED, we first introduce a FHE framework called BatchPBS, a batch pro grammable bootstrapping framework that improves the recent Liu and Wang (ASIACRYPT 2023) amortised scheme from 6.7 ms to 3 msper ciphertext while adding multi-value bootstrapping (MVB) support. Based on BatchPBS, we propose efficient SQL algorithms in SIMD-style to reduce the computation time and a novel AES transcipher protocol to reduce the data transfer volume. Thus, STEED reduces query time by 13 × and data transfer amount by 165 to 534.9 × compared with SOTA work. Considering end-to-end economic cost of TPC-H query on a database with 1 million rows, STEED reduces the expense of deploying on AWS by $28444.8 per 100 queries. (The code can be found at https://github.com/alibaba-damo-academy/ctl-he)
Fahong Zhang 0002, Cheng Hong 0001, Yanheng Lu, Meng Li 0004, Leibo Liu, Sheng Wang 0011, Feifei Li 0001, Chen Yang 0005, Dimin Niu, Yuan Xie 0001
IEEE Trans. Dependable Secur. Comput.8
2025 Frequency-Biased Synergistic Design for Image Compression and Compensation
abstract
Compression artifacts removal (CAR), an effective post-processing method to reduce compression distortion in edge-side codecs, demonstrates remarkable results by utilizing convolutional neural networks (CNNs) on high computational power cloud side. Traditional image compression reduces redundancy in the frequency domain, and we observed that CNNs also exhibit a bias in frequency domain when handling compression distortions. However, no prior research leverages this frequency bias to design compression methods tailored to CAR CNNs, or vice versa. In this paper, we present a synergistic design that bridges the gap between image compression and learnable compensation for CAR. Our investigation reveals that different compensation networks have varying effects on low and high-frequencies. Building upon these insights, we propose a pioneering redesign of the quantization process, a fundamental component in lossy image compression, to more effectively compress low-frequency information. Additionally, we devise a novel compensation framework that applies different neural networks for reconstructing different frequencies, incorporating a basis attention block to prioritize intentionally dropped low-frequency information, thereby enhancing the overall compensation. We instantiate two compensation networks based on this synergistic design and conduct extensive experiments on three image compression standards, demonstrating that our approach significantly reduces bitrate consumption while delivering high perceptual quality.
Qi Zheng 0004, Zihao Liu 0015, Yilian Zhong, Peiye Liu, Tao Liu 0023, Shusong Xu, Yanheng Lu, Sicheng Li 0001, Dimin Niu, Yibo Fan
CVPR8
2025 ANS-LIC: A High-Throughput Parallel Hardware Implementation of ANS for Learned Imagination Codecs
abstract
Asymmetric Numeral Systems (ANS) play a significant role in learned image codecs (LIC) because of their high coding efficiency. However, it constitutes a substantial portion of inference time, making it the main bottleneck in real-time LIC due to its high computational demands, complex control logic, and serial execution flow. To address these challenges, this paper introduces a hardware-oriented ANS algorithm hANS that reduces complex calculations for state encoding and state-symbol decoding. Furthermore, hANS employs fixed-latency calculation to eliminate control logic, which often causes inconsistent delays. To further enhance throughput, we propose a hardware architecture of ANS for LIC (ANS-LIC), introducing a novel hardware parallelism scheme that incorporates pipeline execution and multi-bin parallelism for encoding, along with multi-stream parallelism for decoding. Additionally, by optimizing the execution order, we achieve a reduction in hardware resource utilization during the decoding process. The proposed ANS-LIC hardware is implemented in RTL and synthesized using TSMC 65nm technology and the Alveo U250 Data Center Accelerator Card. We evaluate ANS-LIC on the Kodak and DIV2K LIC datasets, achieving a 1.17% compression ratio improvement over the SOTA method, Recoil. The implementation results and comparison with other works are presented in Table 1. The synthesis indicates that ANS-LIC requires only 385.5/393.0k gates for encoding and decoding, without SRAM. ANS-LIC achieves throughput improvements of 13.29×/1.47× for encoding and decoding over Recoil. In summary, the proposed ANS-LIC demonstrates substantial advantages.
Shiyan Yi, Guohao Xu, Boyuan Shan, Yanheng Lu, Xiaoyang Zeng, Yibo Fan
DCC6
2025 Unicorn: Unified Neural Image Compression with One Number Reconstruction
abstract
Prevalent lossy image compression schemes can be divided into: 1) explicit image compression (EIC), including traditional standards and neural end-to-end algorithms; 2) implicit image compression (IIC) based on implicit neural representations (INR). The former is encountering impasses of leveling off bitrate reduction at a cost of tremendous complexity while the latter suffers from excessive smoothing quality as well as lengthy decoder models. In this paper, we propose an innovative paradigm, which we dub Unicorn (Unified Neural Image Compression with One Nnumber Reconstruction). By conceptualizing the images as index-image pairs and learning the inherent distribution of pairs in a subtle neural network model, Unicorn can reconstruct a visually pleasing image from a randomly generated noise with only one index number. The neural model serves as the unified decoder of images while the noises and indexes corresponds to explicit representations. As a proof of concept, we propose an effective and efficient prototype of Unicorn based on latent diffusion models with tailored model designs. Quantitive and qualitative experimental results demonstrate that our prototype achieves significant bitrates reduction compared with EIC and IIC algorithms. More impressively, benefitting from the unified decoder, our compression ratio escalates as the quantity of images increases. We envision that more advanced model designs will endow Unicorn with greater potential in image compression. The code will be made publicly available upon publication.
Qi Zheng 0004, Haozhi Wang, Zihao Liu 0015, Zhijian Hao, Bu Chen, Min Li 0033, Rui Wan, Peiye Liu, Yanheng Lu, Dimin Niu, Jinjia Zhou, Minge Jing, Yibo Fan
ACM Multimedia10
2025 SAFE: A Scalable Homomorphic Encryption Accelerator for Vertical Federated Learning
abstract
Privacy preservation has become a critical concern for governments, hospitals, and large corporations. Homomorphic encryption (HE) enables a ciphertext-based computation paradigm with strong security guarantees. In emerging cross-agency data cooperation scenarios like vertical federated learning (VFL), HE protects the data interaction from exposure to counterparts. However, computation on ciphertext has significant performance challenges due to increased data size and substantial overhead. Related work has been proposed to accelerate HE using parallel hardware, such as GPUs, FPGAs, and ASICs. However, many existing hardware accelerators target specific HE operations, such as number theoretic transform (NTT) and key switching, providing limited performance improvement for end-to-end applications. Others support bootstrapping, which requires quite a large ASIC design. To better support existing VFL training applications, we propose SAFE, an HE accelerator for scalable homomorphic matrix-vector products (HMVPs), which is the performance bottleneck. SAFE adopts a coefficient-wise encoded HMVP algorithm, despite a vanilla mode, we further explore the compressed and concatenated modes, which can fully utilize the polynomial encoding slots. The proposed hardware architecture, customized for HMVP dataflow, supports spatial and temporal parallelization of function units. The most costly polynomial function, NTT, is implemented with a low-area constant geometry unit which improves efficiency by$2.43\times $. SAFE is implemented as a CPU-FPGA heterogeneous acceleration system, unleashing the multithread potential. The evaluation demonstrates an up to$36\times $speed-up in end-to-end federated logistic regression training.
Yanheng Lu, Xuanle Ren, Ruiguang Zhong, Jiansong Zhang 0001, Hanghang Wu, Xiaofu Zheng, Tingqiang Chu, Cheng Hong 0001, Changzheng Wei, Dimin Niu, Yuan Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 A Tightly Coupled AI-ISP Vision Processor
abstract
To achieve high-quality and high-resolution image processing, this work presents a novel vision processor that facilitates deep learning-enhanced image processing pipelines. At the system level, by identifying that a divide-and-conquer approach is essential to synergize both classical image processing and image enhancement networks, we develop a tightly coupled system with strip-tile conversion dataflow to enable fine-grained low-latency data interactions between image signal processors (ISPs) and the deep learning accelerator (DLA). At the architecture level, we design a comprehensive set of 21 efficient image processing modules to construct classical ISP pipelines, a tile-based strip layer fusion DLA specifically optimized for networks, and a programmable pixel pool that seamlessly supports the data access patterns of the ISP and the DLA. At the software and hardware co-design level, we propose a comprehensive optimization framework to address the implementation overhead of networks while maintaining the image quality. Finally, evaluations of the AI-ISP vision processor demonstrate 53.95% external memory access reduction and 35.51% latency reduction, delivering superior image quality with minimal on-chip memory overhead. A throughput of up to 168.5 frames per second facilitates efficient processing of ultra-high definition (UHD) resolution images.
Hao Zhang 0126, Sicheng Li 0001, Yupeng Gui, Zhiyong Li 0016, Shusong Xu, Yanheng Lu, Dimin Niu, Hongzhong Zheng, Yen-Kuang Chen, Yuan Xie 0001, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.6
2024 Salus: A Practical Trusted Execution Environment for CPU-FPGA Heterogeneous Cloud Platforms
abstract
CPU-FPGA heterogeneous architectures have become increasingly popular in cloud environments for accelerating compute-intensive tasks. Ensuring the protection of sensitive data processed by these architectures requires the presence of a trusted execution environment (TEE). This work highlights the requirements for designing an FPGA TEE, the challenges faced in deploying existing solutions on commercial-off-the-shelf (COTS) cloud FPGA services, and the limitations of previous works that primarily focus on standalone FPGA TEEs. In response to these challenges, Salus introduces an innovative approach by leveraging an enclave running on the host with a TEE-enabled CPU. This approach aims to protect and attest the bitstream loaded on the FPGA side. By repurposing COTS FPGA bitstream utilities in a novel manner and adopting a proposed security-enhanced FPGA IP, Salus presents a practical design for an FPGA TEE, with minor efforts required.
Sheng Wang 0011, Le Su, Yanheng Lu, Yijin Guan, Dimin Niu, Mingyu Gao 0001, Yuan Xie 0001, Feifei Li 0001
ASPLOS (4)6
2024 Zero-Shot Structure-Preserving Diffusion Model for High Dynamic Range Tone Mapping
abstract
Tone mapping techniques, aiming to convert high dynamic range (HDR) images to high-quality low dynamic range (LDR) images for display, play a more crucial role in real-world vision systems with the increasing application of HDR images. However, obtaining paired HDR and high-quality LDR images is difficult, posing a challenge to deep learning based tone mapping methods. To over-come this challenge, we propose a novel zero-shot tone mapping framework that utilizes shared structure knowl-edge, allowing us to transfer a pre-trained mapping model from the LDR domain to HDR fields without paired training data. Our approach involves decomposing both the LDR and HDR images into two components: structural in-formation and tonal information. To preserve the original image's structure, we modify the reverse sampling process of a diffusion model and explicitly incorporate the struc-ture information into the intermediate results. Additionally, for improved image details, we introduce a dual-control network architecture that enables different types of conditional inputs to control different scales of the output. Experimental results demonstrate the effectiveness of our approach, surpassing previous state-of-the-art methods both qualitatively and quantitatively. Moreover, our model ex-hibits versatility and can be applied to other low-level vi-sion tasks without retraining. The code is available at https://github.com/ZSDM-HDRIZero-Shot-Diffusion-HDR.
Ruoxi Zhu, Shusong Xu, Peiye Liu, Sicheng Li 0001, Yanheng Lu, Dimin Niu, Zihao Liu 0015, Zihao Meng, Zhiyong Li 0016, Xinhua Chen, Yibo Fan
CVPR5
2024 A High-Throughput Private Inference Engine Based on 3D Stacked Memory
abstract
Fully Homomorphic Encryption (FHE) enables unlimited computation depth, allowing privacy-enhanced neural network inference tasks directly on the ciphertext. However, existing FHE architectures suffer from the memory access bottleneck. This work proposes a High-throughput FHE engine for private inference (PI) based on 3D stacked memory (H3). H3 adopts the software-hardware co-design that dynamically adjusts the polynomial decomposition during the PI process to minimize the computation and storage overhead at a fine granularity. With 3D hybrid bonding, H3 integrates a logic die with a multi-layer embedded DRAM, routing data efficiently to the processing unit array through an efficient broadcast mechanism. H3 consumes 192mm2 when implemented using a 28nm logic process. It achieves 1.36 million LeNet-5 or 920 ResNet-20 PI per minute, surpassing existing 7nm accelerators by 52%. This demonstrates that 3D memory is a promising technology to promote the performance of FHE.
Ling Liang 0003, Zhirui Li, Fahong Zhang 0004, Yanheng Lu
DAC6
2023 CHAM: A Customized Homomorphic Encryption Accelerator for Fast Matrix-Vector Product
abstract
Homomorphic encryption (HE) is a promising technique for privacy-preserving computing because it allows computation on encrypted data without decryption. HE, however, suffers from poor performance due to enlarged data size and exploded amount of computation. Related work has been proposed to accelerate HE using GPUs, FPGAs, and ASICs. The existing work, however, aims at specific HE schemes and fails to consider the fast-evolving algorithms. For example, HE algorithms that combine different HE schemes have demonstrated capability of supporting more types of HE operations and ciphertexts. Moreover, some existing hardware accelerators target small HE operations (such as number theoretic transform and key-switch), which however provides limited or even neglected performance improvement for end-to-end applications. To better support existing privacy-preserving applications (e.g., logistic regression and neural network inference), we propose CHAM, an HE accelerator, for high-performance matrix-vector product, which can be easily extended to 2-D and 3-D convolutions. Motivated by the evolution of algorithms, CHAM supports not only traditional HE operations, but also different types of ciphertexts and the conversion between them. We implement CHAM with Xilinx FPGAs. The evaluation demonstrates 1800× speed-up for matrix-vector product, 36× speed-up for logistic regression, and 144× speed-up for Beaver triple generation compared to the existing work.
Xuanle Ren, Yanheng Lu, Ruiguang Zhong, Jiansong Zhang 0001, Hanghang Wu, Xiaofu Zheng, Tingqiang Chu, Cheng Hong 0001, Changzheng Wei, Dimin Niu, Yuan Xie 0001
DAC4