Yanze Wu

dblp:239/3708 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GPU Acceleration of the Sum-Check Protocol Over Towers of Binary Fields for Verifiable Computing
abstract
Emerging zero-knowledge proof protocols such as Binius and Binius-FRI operate over towers of binary fields, allowing for ultra-fast polynomial commitments over a base field. Sum-check, a key protocol in algebraic proof systems, is one of the key implementation bottlenecks for Binius and similar protocols. While sum-check is a massively parallel algorithm, GPU acceleration of sum-check has received little attention due to the lack of native GPU support for binary field multiplication. Hence, in this paper, we explore the key issues in existing GPU-based sum-check accelerators and present SumCATS - an efficient GPU implementation for sum-check acceleration. SumCATS leverages two fundamental improvements over the existing solutions. First, it adapts a CPU-based algorithmic improvement to sum-check proving and applies it to GPUs by recognizing the reduction pattern and shared memory optimizations. Secondly, SumCATS reduces the number of global memory accesses by precomputing products of random challenges and using base field operations to reconstruct extension field elements. When these optimizations are combined, SumCATS achieves a significant speedup (1.81× on NVIDIA RTX 3090 Ti, 1.62× on NVIDIA A100) over the baseline GPU implementation (Binius-GPU) for sum-check over binary tower fields. The code and research artifacts for SumCATS design are available at https://github.com/SPIRE-GMU/sum_cats.
Andrew Fan, Yanze Wu, Harry Han, Md Tanvir Arafin
DATE2
2025 Energy-Efficient Acceleration of Hash-Based Post-Quantum Cryptographic Schemes on Embedded Spatial Architectures
abstract
This work introduces AXIOS, a novel spatial architecture for accelerating hash-based post-quantum cryptography (PQC) primitives. AXIOS demonstrates that structural regularities in hash-based algorithms can be efficiently mapped to spatial accelerators that support FPGA-based programming for granular control and coarse-grained reconfigurable arrays (CGRA) for repeated tasks. AXIOS selects the key generation task of the eXtended Markle Signature Scheme (XMSS), which embodies critical implementation challenges in modern hash-based (PQC) algorithms. The AXIOS implementation on AMD’s VCK190 platform demonstrates an $8.54 \times$ improvement in runtime and a $71.65 \times$ improvement in energy efficiency compared to a benchmark implementation on Intel’s Core $\mathbf{i 9 - 1 4 9 0 0 K}$. AXIOS also breaks the current record of XMSS acceleration in terms of execution time on an embedded SoC or an FPGA platform. To our knowledge, this is the first efficient hardware implementation of compute-intensive hash-based PQC schemes in an embedded spatial architecture. Albeit complex, this FPGA+CGRA-based design is a promising step to support compute-intensive PQC applications at the edge. This work’s code and experimental artifacts are publicly available at https://github.com/SPIRE-GMU/AXIOS.
Yanze Wu, Md Tanvir Arafin
PACT1
2025 Urban Outdoor Propagation Measurements and Channel Models at 6.75 GHz FR1(C) and 16.95 GHz FR3 Upper Mid-Band Spectrum for 5G and 6G
abstract
Global allocations in the upper mid-band spectrum (4-24 GHz) necessitate a comprehensive exploration of the propagation behavior to meet the promise of coverage and capacity. This paper presents an extensive Urban Microcell (UMi) outdoor propagation measurement campaign at 6.75 GHz and 16.95 GHz conducted in Downtown Brooklyn, USA, using a 1 GHz bandwidth sliding correlation channel sounder over 40-880 m propagation distance, encompassing seven Line of Sight (LOS) and 13 Non-Line of Sight (NLOS) locations. Analysis of the path loss (PL) reveals lower directional and omnidirectional PL exponents compared to mmWave and sub-THz frequencies in the UMi environment, using the close-in (CI) free space PL (FSPL) model with a 1 m reference distance. Additionally, a decreasing trend in root mean square (RMS) delay spread (DS) and angular spread (AS) with increasing frequency was observed. The measured NLOS RMS DS and RMS AS mean values (as computed by 3GPP methods) are found to be consistently lower compared to 3GPP model predictions. Point-data tables with corresponding site-specific environmental information for all measured statistics at each TX-RX location are provided to support the models and results. The spatio-temporal statistics presented here offer valuable insights for the design of nextgeneration wireless systems and networks.
Dipankar Shakya, Mingjun Ying, Theodore S. Rappaport, Peijie Ma, Idris Al-Wazani, Yanze Wu, Doru Calin, Hitesh Poddar, Ahmad Bazzi, Marwa Chafii, Yunchou Xing, Amitava Ghosh
ICC6
2025 Upper Mid-Band Channel Measurements and Characterization at 6.75 GHz FR1(C) and 16.95 GHz FR3 in an Indoor Factory Scenario
abstract
This paper presents detailed radio propagation measurements for an indoor factory (InF) environment at$\mathbf{6. 7 5 ~ G H z}$and 16.95 GHz using a 1 GHz bandwidth channel sounder. Conducted at the NYU MakerSpace in the NYU Tandon School of Engineering campus in Brooklyn, NY, USA, our measurement campaign characterizes the radio propagation in a representative small factory with diverse machinery and open workspaces across 12 locations, comprising five line-of-sight (LOS) and seven non-line-of-sight (NLOS) scenarios. Analysis using the close-in (CI) free space path loss (FSPL) model with a 1 m reference distance reveals path loss exponents (PLE) below 2 in LOS at 6.75 GHz and 16.95 GHz, while in NLOS, PLE is similar to free-space propagation (e.g., PLE = 2). The RMS delay spread (DS) decreases at higher frequencies with a clear frequency dependence. Also, measurements show a wider RMS angular spread (AS) in NLOS compared to LOS at both frequency bands, with a decreasing trend as frequency increases. These observations in a densescatterer factory environment demonstrate frequency-dependent behavior that differs from existing industry-standard 3GPP models. Our findings provide crucial insights into complex propagation mechanisms in factory environments, essential for designing robust air interface and industrial wireless networks at the upper mid-band FR3 spectrum.
Mingjun Ying, Dipankar Shakya, Theodore S. Rappaport, Peijie Ma, Idris Al-Wazani, Yanze Wu, Hitesh Poddar
ICC7
2025 Thena: Torus Fully Homomorphic Encryption on Energy-Efficient Heterogeneous Architecture
abstract
Fully Homomorphic Encryption (FHE) enables privacy-preserving computations on encrypted data with strong security guarantees. Torus-based FHE (TFHE) emerges as a promising candidate among FHE variants due to its efficient Boolean logic operation and unlimited computational depth. However, it heavily relies on bootstrapping, a computationally intensive technique. Although there has been significant progress in improving the throughput and latency of the bootstrapping process, there exists a gap in the energy efficiency research of this process without compromising its speed. Also, energy-efficient implementation of TFHE is a key requirement for its application in energy-constrained systems. This work introduces THENA, an energy-efficient bootstrapping accelerator for TFHE built on a heterogeneous Versal adaptive system on chip (ASoC) platform to address this gap. THENA partitions the bootstrapping workload into different parts of ASoC: the serial operations are handled by the processing system (PS), the compute-intensive torus multiplications are mapped to the adaptive intelligent engine (AIE), and the memory and communication operations are allocated on the programming logic (PL). THENA derives a wavefront arraybased energy-efficient multiplier, achieving a higher$(2 \times)$improvement in throughput over a similar implementation (SaberNTT, TCAS '23). THENA uses this multiplier to deliver an end-to-end bootstrapping accelerator on the Versal VCK-190 platform. THENA delivers$7 \times$better energy efficiency for bootstrapping than GPU-based CuFHE (RTX 3090) and outperforms existing complete FPGA designs, such as YKP (HPEC '22) by demonstrating 17.5%, and 35.6% decrease in latency and energy consumption. To the best of our knowledge, this is the first PS+PL+AIE-based heterogeneous TFHE accelerator on Versal ASoCs. THENA's code and experimental artifacts are published at https://github.com/SPIRE-GMU/tfhe-aie/.
Yanze Wu, Md Tanvir Arafin
ICCD1
2025 DreamO: A Unified Framework for Image Customization
abstract
Recently, extensive research on image customization (e.g., identity, subject, style, background, etc.) demonstrates strong customization capabilities in large-scale generative models. However, most approaches are designed for specific tasks, restricting their generalizability to combine different types of condition. Developing a unified framework for image customization remains an open challenge. In this paper, we present DreamO, an image customization framework designed to support a wide range of tasks while facilitating seamless integration of multiple conditions. Specifically, DreamO utilizes a diffusion transformer (DiT) framework to uniformly process input of different types. During training, we introduce a feature routing constraint to facilitate the precise querying of relevant information from reference images. Additionally, we design a placeholder strategy that associates specific placeholders with conditions at particular positions, enabling control over the placement of conditions in the generated results. Moreover, we employ a progressive training strategy to ensure smooth model convergence and correct the generation quality of the final output. Extensive experiments demonstrate that the proposed DreamO can effectively perform various image customization tasks with high quality and flexibly integrate different types of control conditions. Project page: https://mc-e.github.io/project/DreamO
Chong Mou, Yanze Wu, Wenxu Wu, Pengze Zhang, Yufeng Cheng, Xinghui Li, Mengtian Li 0003, Mingcong Liu, Yunsheng Jiang, Shaojin Wu, Songtao Zhao, Jian Zhang 0018
SIGGRAPH Asia2
2024 T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion Models
abstract
The incredible generative ability of large-scale text-to-image (T2I) models has demonstrated strong power of learning complex structures and meaningful semantics. However, relying solely on text prompts cannot fully take advantage of the knowledge learned by the model, especially when flexible and accurate controlling (e.g., structure and color) is needed. In this paper, we aim to ``dig out" the capabilities that T2I models have implicitly learned, and then explicitly use them to control the generation more granularly. Specifically, we propose to learn low-cost T2I-Adapters to align internal knowledge in T2I models with external control signals, while freezing the original large T2I models. In this way, we can train various adapters according to different conditions, achieving rich control and editing effects in the color and structure of the generation results. Further, the proposed T2I-Adapters have attractive properties of practical value, such as composability and generalization ability. Extensive experiments demonstrate that our T2I-Adapter has promising generation quality and a wide range of applications. Our code is available at https://github.com/TencentARC/T2I-Adapter.
Chong Mou, Xintao Wang 0002, Liangbin Xie, Yanze Wu, Jian Zhang 0018, Zhongang Qi, Ying Shan
AAAI4
2024 DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations
abstract
The diffusion-based text-to-image model harbors im-mense potential in transferring reference style. However, current encoder-based approaches significantly impair the text controllability of text-to-image models while transfer-ring styles. In this paper, we introduce DEADiff to address this issue using the following two strategies: 1) a mecha-nism to decouple the style and semantics of reference images. The decoupled feature representations are first extracted by Q-Formers which are instructed by different text descriptions. Then they are injected into mutually exclusive subsets of cross-attention layers for better disentanglement. 2) A non-reconstructive learning method. The Q-Formers are trained using paired images rather than the identical target, in which the reference image and the ground-truth image are with the same style or semantics. We show that DEADiff attains the best visual stylization results and optimal balance between the text controllability inherent in the text-to-image model and style similarity to the reference image, as demonstrated both quantitatively and qualitatively. Our project page is https://tianhao-qi.github.io/DEADiff‘/.
Tianhao Qi, Shancheng Fang, Yanze Wu, Hongtao Xie 0001, Jiawei Liu 0001, Lang Chen, Yongdong Zhang 0001
CVPR3
2024 PuLID: Pure and Lightning ID Customization via Contrastive Alignment
abstract
We propose Pure and Lightning ID customization (PuLID), a novel tuning-free ID customization method for text-to-image generation. By incorporating a Lightning T2I branch with a standard diffusion one, PuLID introduces both contrastive alignment loss and accurate ID loss, minimizing disruption to the original model and ensuring high ID fidelity. Experiments show that PuLID achieves superior performance in both ID fidelity and editability. Another attractive property of PuLID is that the image elements (\eg, background, lighting, composition, and style) before and after the ID insertion are kept as consistent as possible. Codes and models are available at https://github.com/ToTheBeginning/PuLID
Yanze Wu, Zhuowei Chen, Lang Chen
NeurIPS2
2024 Empowering Real-World Image Super-Resolution With Flexible Interactive Modulation
abstract
Interactive image restoration aims to construct an interactive pathway between users and restoration networks, which empowers users to modulate the restoration results according to their own demands. However, existing methods are primarily limited to training their networks with predefined and simplistic synthetic degradations. Consequently, these methods often encounter significant performance degradation when confronted with real-world degradations that deviate from their assumptions. Furthermore, existing interactive image restoration approaches solely support global modulation, wherein a single modulation factor governs the reconstruction process for the entire image. In this paper, we propose a novel method to perform real-world and intricate image super-resolution in an interactive manner. Specifically, we propose a metric-learning-based degradation estimation strategy to estimate not only the overall degradation level of the entire image but also the finer-grained, pixel-wise degradation within real-world scenarios. This enables local control over the restoration results by selectively modulating the corresponding regions based on the densely-estimated degradation map. Additionally, a new metric-argumented loss is proposed to further enhance the performance of real-world image super-resolution. Through extensive experimentation, we demonstrate the efficacy of our method in achieving exceptional modulation and restoration performance in real-world image super-resolution tasks, all while maintaining an appealing model complexity.
Chong Mou, Xintao Wang 0002, Yanze Wu, Ying Shan, Jian Zhang 0018
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Metric Learning Based Interactive Modulation for Real-World Super-Resolution
Chong Mou, Yanze Wu, Xintao Wang 0002, Chao Dong 0005, Jian Zhang 0018, Ying Shan
ECCV (17)2
2022 AnimeSR: Learning Real-World Super-Resolution Models for Animation Videos
abstract
This paper studies the problem of real-world video super-resolution (VSR) for animation videos, and reveals three key improvements for practical animation VSR. First, recent real-world super-resolution approaches typically rely on degradation simulation using basic operators without any learning capability, such as blur, noise, and compression. In this work, we propose to learn such basic operators from real low-quality animation videos, and incorporate the learned ones into the degradation generation pipeline. Such neural-network-based basic operators could help to better capture the distribution of real degradations. Second, a large-scale high-quality animation video dataset, AVC, is built to facilitate comprehensive training and evaluations for animation VSR. Third, we further investigate an efficient multi-scale network structure. It takes advantage of the efficiency of unidirectional recurrent networks and the effectiveness of sliding-window-based methods. Thanks to the above delicate designs, our method, AnimeSR, is capable of restoring real-world low-quality animation videos effectively and efficiently, achieving superior performance to previous state-of-the-art methods.
Yanze Wu, Xintao Wang 0002, Gen Li 0011, Ying Shan
NeurIPS1
2022 Area, Time and Energy Efficient Multicore Hardware Accelerators for Extended Merkle Signature Scheme
abstract
This paper addresses a barrier that prevents the timely adoption of post-quantum signature algorithms, such as the eXtended Merkle Signature Scheme (XMSS), due to its lack of fast, cost-effective and energy-efficient hardware accelerators. Two new architectures that use more than one hash core are proposed for the first time to significantly reduce the latency of two bottleneck XMSS operations, namely key generation and signature generation, for which the speed of existing hardware accelerators is still apparently inadequate. The first proposed multi-core design uses block RAM and a simplified data flow to maximize the use of$p$hash cores concurrently in three major sequential stages of computation, i. e., Winternitz One-time Signature (WOTS), L-tree and Merkle tree. The second proposed multi-core design adds a dedicated hash core for tree hashing in the L-tree and Merkle tree while keeping the$p$hash cores solely for chain hashing in WOTS. The dedicated hash core leapfrogs between the L-tree and Merkle tree and computes concurrently with the$p$hash cores to keep the$p+1$hash cores active most of the time while minimizing the storage requirement and energy consumption. Both designs are implemented on a 28 nm ATRIX-7 FPGA chip. Experimental results show that both proposed accelerators with$p=8$operate at a much faster speed and consume significantly less hardware resources and energy than all existing XMSS accelerators. Specifically, they are$\sim 8\times $and$\sim 6\times $faster than the fastest reported design in key generation and signature generation operations, respectively.
Yuan Cao 0003, Yanze Wu, Lan Qin, Chip-Hong Chang
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 An Efficient Full Hardware Implementation of Extended Merkle Signature Scheme
abstract
This paper presents a full hardware implementation of the eXtended Merkle Signature Scheme (XMSS), a NIST approved and IETF RFC specified post-quantum cryptography (PQC) algorithm. An optimized node traversal is proposed to enable efficient memory utilization without compromising the computational latency of the L-tree and Merkle tree construction, which are two key components used for the compression of the Winternitz One-Time Signature (WOTS) public key in XMSS. The computation of the authentication path during signature generation has also been significantly sped up by our proposed hardware implementation of the Buchmann, Dahmen, and Schneider (BDS) algorithm. Our implementation has completely avoided the use of block random-access memory, which is known to be vulnerable to side-channel attacks. The memory requirement has been highly optimized for implementation with small flip-flop chains and register counters as pointers for fast data access. To the best of our knowledge, this is the first full hardware implementation of all threekey generation,signingandverificationoperations of XMSS. The design has been prototyped and evaluated on a 28 nm FPGA platform to demonstrate its performance improvements over the most efficient software and hardware/software co-design methods reported to date. Specifically, it increases the computational efficiency of the best reported XMSS implementation forkey generationandsignature generationby about 20% and 50%, respectively. It can also run at 10% higher clock speed than the fastest hardware implementation ofsignature verificationin FPGA with 8% lower hardware resource utilization.
Yuan Cao 0003, Yanze Wu, Wen Wang 0007, Jing Ye 0001, Chip-Hong Chang
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Towards Vivid and Diverse Image Colorization with Generative Color Prior
abstract
Colorization has attracted increasing interest in recent years. Classic reference-based methods usually rely on external color images for plausible results. A large image database or online search engine is inevitably required for retrieving such exemplars. Recent deep-learning-based methods could automatically colorize images at a low cost. However, unsatisfactory artifacts and incoherent colors are always accompanied. In this work, we aim at recovering vivid colors by leveraging the rich and diverse color priors encapsulated in a pretrained Generative Adversarial Networks (GAN). Specifically, we first "retrieve" matched features (similar to exemplars) via a GAN encoder and then incorporate these features into the colorization process with feature modulations. Thanks to the powerful generative color prior and delicate designs, our method could produce vivid colors with a single forward pass. Moreover, it is highly convenient to obtain diverse results by modifying GAN latent codes. Our method also inherits the merit of interpretable controls of GANs and could attain controllable and smooth transitions by walking through GAN latent space. Extensive experiments and user studies demonstrate that our method achieves superior performance than previous works.
Yanze Wu, Xintao Wang 0002, Yu Li 0003, Honglun Zhang, Ying Shan
ICCV1