Leyuan Wang

dblp:165/7400 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Security and privacy · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 5 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Amplitude -phase decomposition-based latent diffusion model for underwater image enhancement
Hansen Zhang, Can Pan, Leyuan Wang, Jiaju Tao
Expert Syst. Appl.4
2026 Adaptive spatial-aware non-maximum suppression for dense object detection
Can Pan, Leyuan Wang, Jiaju Tao
Neurocomputing3
2026 Symmetric Image-Text Tuning With Entropy-Guided Fusion for Online Continual Learning in Non-Stationary Visual Streams
abstract
Online continual learning studies how models learn from continuous and non-stationary data streams. In this paper, we observe that CLIP models exhibit an asymmetric image-text interaction under online continual learning. Specifically, text features of previously seen classes may introduce unfavorable supervision when paired with visual features of newly observed data, leading to catastrophic forgetting. To alleviate this issue, we propose a simple yet effective symmetric image-text tuning (SIT) strategy that removes such asymmetric text supervision during online learning. We further introduce an entropy-guided fusion (EGF) mechanism that adaptively combines predictions from the pretrained and finetuned branches based on their relative uncertainty. This design allows the model to recover pretrained knowledge when the finetuned branch becomes unreliable, while still preserving plasticity on recently observed classes when confidence is high. In addition, we present MiD-Blurry, an online continual learning benchmark that combines multiple class distribution patterns to better reflect realistic data streams with blurred temporal boundaries. Extensive experiments on standard continual learning benchmarks and the MiD-Blurry setting evaluate inference-at-any-time performance and generalization to future data. The results show that the proposed approach maintains a practical balance between adapting to new data and preserving previously learned information in realistic online learning scenarios.
Leyuan Wang, Liuyu Xiang, Yiwei Ru, Yunlong Wang 0003, Zhaofeng He 0001
IEEE Trans. Image Process.1
2026 Rethinking Class-Incremental Learning From a Dynamic Imbalanced Learning Perspective
abstract
Deep neural networks suffer from catastrophic forgetting when continually learning new concepts. In this paper, we analyze this problem from a data imbalance point of view. We argue that the imbalance between old task and new task data contributes to forgetting of the old tasks. Moreover, the increasing imbalance ratio during incremental learning further aggravates the problem. To address the dynamic imbalance issue, we propose Uniform Prototype Contrastive Learning (UPCL), where uniform and compact features are learned. Specifically, we generate a set of non-learnable uniform prototypes before each task starts. Then we assign these uniform prototypes to each class and guide the feature learning through prototype contrastive learning. We also dynamically adjust the relative margin between old and new classes so that the feature distribution will be maintained balanced and compact. Finally, we demonstrate through extensive experiments that the proposed method achieves state-of-the-art performance on several benchmark including CIFAR-100, ImageNet-100, TinyImageNet, Food-101, and CUB-200. Experimental results show that our approach not only effectively addresses the issue of imbalanced old data in memory but also tackles the problem of imbalanced new data distributions.
Leyuan Wang, Liuyu Xiang, Yunlong Wang 0003, Huijia Wu, Huafeng Yang, Jingqian Liu, Zhaofeng He 0001
IEEE Trans. Multim.1
2025 LAMAR: LLM-Guided Adaptive Perceptual Modeling for Micro-Action Recognition
abstract
We present LAMAR, a novel framework for Micro-Action Recognition that addresses the challenges of identifying subtle, ephemeral human movements lasting less than one-third of a second. Our approach leverages large language models to estimate semantic complexity of micro-actions and dynamically configure a hierarchical Vision Transformer architecture accordingly. LAMAR introduces: (1) a principled complexity estimation module that quantifies recognition difficulty by analyzing subtlety, noise susceptibility, and intra-class ambiguity; and (2) an adaptive perception pipeline that dynamically adjusts spatiotemporal resolution and attention mechanisms based on estimated complexity. Experiments on the MA-52 benchmark demonstrate that LAMAR outperforms state-of-the-art methods by 7.20% in accuracy while maintaining computational efficiency, establishing a new paradigm for context-aware visual analysis that intelligently allocates resources based on task difficulty.
Yiwei Ru, Leyuan Wang, Ma He, Zhaofeng He 0001, Zhenan Sun
IJCB2
2025 Cross-Optical Property Image Translation for Face Anti-Spoofing: From Visible to Polarization
abstract
Despite the development of spectral sensors and spectral data-driven learning methods which have led to significant advances in face anti-spoofing (FAS), the singular dimensionality of spectral information often results in poor robustness and weak generalization. Polarization, another fundamental property of light, can reveal intrinsic differences between genuine and fake faces with advantaged performance in precision, robustness, and generalizability. In this paper, we propose a facial image translation method from visible light (VIS) to polarization (VPT), capable of generating valuable polarimetric optical characteristics for facial presentation attack detection using VIS spectrum information input only. Specifically, the VPT method adopts a multi-stream network structure, comprising a main network and two branch networks, to translate VIS images into degree of polarization (DoP) images and Stokes polarization parameters${S}_{1}$and${S}_{2}$. To further improve image translation quality, we introduce a frequency-domain consistency loss as a complement to the existing spatial losses to narrow the gap in the frequency domain. The physical mapping relations for the DoP and Stokes parameters are employed, and the Stokes loss is designed to ensure that the generated polarization modalities conform to objective physical laws. Extensive experiments on the CASIA-Polar and CASIA-SURF datasets demonstrate the superiority of VPT over other baseline methods in terms of polarization image quality and its remarkable performance in the FAS task. This work leverages the inherent physical advantages of polarization information in material discrimination tasks while addressing hardware limitations in polarization image collection, proposing a novel solution for face recognition system security control.
Yu Tian 0017, Kunbo Zhang, Yalin Huang, Leyuan Wang, Yue Liu 0005, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.4
2023 ByteTransformer: A High-Performance Transformer Boosted for Variable-Length Inputs
abstract
Transformers have become keystone models in natural language processing over the past decade. They have achieved great popularity in deep learning applications, but the increasing sizes of the parameter spaces required by transformer models generate a commensurate need to accelerate performance. Natural language processing problems are also routinely faced with variable-length sequences, as word counts commonly vary among sentences. Existing deep learning frameworks pad variable-length sequences to a maximal length, which adds significant memory and computational overhead. In this paper, we present ByteTransformer, a high-performance transformer boosted for variable-length inputs. We propose a padding-free algorithm that liberates the entire transformer from redundant computations on zero padded tokens. In addition to algorithmic-level optimization, we provide architecture-aware optimizations for transformer functional modules, especially the performance-critical algorithm Multi-Head Attention (MHA). Experimental results on an NVIDIA A100 GPU with variable-length sequence inputs validate that our fused MHA outperforms PyTorch by 6.13x. The end-to-end performance of ByteTransformer for a forward BERT transformer surpasses state-of-the-art transformer frameworks, such as PyTorch JIT, TensorFlow XLA, Tencent TurboTransformer, Microsoft DeepSpeed-Inference and NVIDIA FasterTransformer, by 87%, 131%, 138%, 74% and 55%, respectively. We also demonstrate the general applicability of our optimization methods to other BERT-like models, including ALBERT, DistilBERT, and DeBERTa.
Chengquan Jiang, Leyuan Wang, Xiaoying Jia 0001, Zizhong Chen, Xin Liu 0086, Yibo Zhu 0001
IPDPS3
2021 UNIT: Unifying Tensorized Instruction Compilation
abstract
Because of the increasing demand for intensive computation in deep neural networks, researchers have developed both hardware and software mechanisms to reduce the compute and memory burden. A widely adopted approach is to use mixed precision data types. However, it is hard to benefit from mixed precision without hardware specialization because of the overhead of data casting. Recently, hardware vendors offer tensorized instructions specialized for mixed-precision tensor operations, such as Intel VNNI, Nvidia Tensor Core, and ARM DOT. These instructions involve a new computing idiom, which reduces multiple low precision elements into one high precision element. The lack of compilation techniques for this emerging idiom makes it hard to utilize these instructions. In practice, one approach is to use vendor-provided libraries for computationally-intensive kernels, but this is inflexible and prevents further optimizations. Another approach is to manually write hardware intrinsics, which is error-prone and difficult for programmers. Some prior works tried to address this problem by creating compilers for each instruction. This requires excessive efforts when it comes to many tensorized instructions. In this work, we develop a compiler framework, UNIT, to unify the compilation for tensorized instructions. The key to this approach is a unified semantics abstraction which makes the integration of new instructions easy, and the reuse of the analysis and transformations possible. Tensorized instructions from different platforms can be compiled via UNIT with moderate effort for favorable performance. Given a tensorized instruction and a tensor operation, UNIT automatically detects the applicability of the instruction, transforms the loop organization of the operation, and rewrites the loop body to take advantage of the tensorized instruction. According to our evaluation, UNIT is able to target various mainstream hardware platforms. The generated end-to-end inference model achieves 1.3 x speedup over Intel oneDNN on an x86 CPU, 1.75x speedup over Nvidia cuDNN on an Nvidia GPU, and 1.13x speedup over a carefully tuned TVM solution for ARM DOT on an ARM CPU.
Jian Weng 0002, Animesh Jain, Jie Wang 0022, Leyuan Wang, Yida Wang 0003, Tony Nowatzki
CGO4
2021 A Large-scale Database for Less Cooperative Iris Recognition
abstract
Since the outbreak of the COVID-19 pandemic, iris recognition has been used increasingly as contactless and unaffected by face masks. Although less user cooperation is an urgent demand for existing systems, corresponding manually annotated databases could hardly be obtained. This paper presents a large-scale database of near-infrared iris images named CASIA-Iris-Degradation Version 1.0 (DV1), which consists of 15 subsets of various degraded images, simulating less cooperative situations such as illumination, off-angle, occlusion, and nonideal eye state. A lot of open-source segmentation and recognition methods are compared comprehensively on the DV1 using multiple evaluations, and the best among them are exploited to conduct ablation studies on each subset. Experimental results show that even the best deep learning frameworks are not robust enough on the database, and further improvements are recommended for challenging factors such as half-open eyes, off-angle, and pupil dilation. Therefore, we publish the DV1 with manual annotations online to promote iris recognition. (http://www.cripacsir.cn/dataset/)
Junxing Hu, Leyuan Wang, Zhengquan Luo, Yunlong Wang 0003, Zhenan Sun
IJCB2
2021 Avoiding Spectacles Reflections on Iris Images Using A Ray-tracing Method
abstract
Spectacles reflection removal is a challenging problem in iris recognition research. The reflection of the spectacles usually contaminates the iris image acquired under infrared illumination. The intense light reflection caused by the active light source makes reflection removal more challenging than normal scenes since important iris texture features are entirely obscured. Eliminating unnecessary reflections can effectively improve iris recognition system performance. This paper proposes a spectacle reflection removal algorithm based on ray coding and ray tracking to remove spectacle reflection in iris images. By decoding the light source’s encoded light beam, the iris imaging device eliminates most of the stray light. Our binocular imaging device tracks the light path to obtain parallax information and realizes reflected light spot removal through image fusion. We designed a prototype system to verify our proposed method in this paper. This method can effectively eliminate reflections without changing iris texture and improve iris recognition in complex scenarios.
Kunbo Zhang, Leyuan Wang
IJCB3
2021 An End-to-End Autofocus Camera for Iris on the Move
abstract
For distant iris recognition, a long focal length lens is generally used to ensure the resolution of iris images, which reduces the depth of field and leads to potential defocus blur. To accommodate users standing statically at different distances, it is necessary to control focus quickly and accurately. And for users in motion, it is also expected to acquire a sufficient amount of accurately focused iris images. In this paper, we introduced a novel rapid auto-focus camera for active refocusing of the iris area of the moving objects with a focus-tunable lens. Our end-to-end computational algorithm can predict the best focus position from one single blurred image and generate the proper lens diopter control signal automatically. This scene-based active manipulation method enables real-time focus tracking of the iris area of a moving object. We built a testing bench to collect real-world focal stacks for evaluation of the autofocus methods. Our camera has reached an autofocus speed of over 50 fps. The results demonstrate the advantages of our proposed camera for biometric perception in static and dynamic scenes. The code is available at https://github.com/Debatrix/AquulaCam.
Leyuan Wang, Kunbo Zhang, Yunlong Wang 0003, Zhenan Sun
IJCB1
2021 HAWQ-V3: Dyadic Neural Network Quantization
abstract
Current low-precision quantization algorithms often have the hidden cost of conversion back and forth from floating point to quantized integer values. This hidden cost limits the latency improvement realized by quantizing Neural Networks. To address this, we present HAWQ-V3, a novel mixed-precision integer-only quantization framework. The contributions of HAWQ-V3 are the following: (i) An integer-only inference where the entire computational graph is performed only with integer multiplication, addition, and bit shifting, without any floating point operations or even integer division; (ii) A novel hardware-aware mixed-precision quantization method where the bit-precision is calculated by solving an integer linear programming problem that balances the trade-off between model perturbation and other constraints, e.g., memory footprint and latency; (iii) Direct hardware deployment and open source contribution for 4-bit uniform/mixed-precision quantization in TVM, achieving an average speed up of 1.45x for uniform 4-bit, as compared to uniform 8-bit for ResNet50 on T4 GPUs; and (iv) extensive evaluation of the proposed methods on ResNet18/50 and InceptionV3, for various model compression levels with/without mixed precision. For ResNet50, our INT8 quantization achieves an accuracy of 77.58%, which is 2.68% higher than prior integer-only work, and our mixed-precision INT4/8 quantization can reduce INT8 latency by 23% and still achieve 76.73% accuracy. Our framework and the TVM implementation have been open sourced (HAWQ, 2020).
Zhewei Yao, Zhen Dong 0003, Zhangcheng Zheng, Amir Gholami, Eric Tan, Leyuan Wang, Qijing Huang 0001, Michael W. Mahoney, Kurt Keutzer
ICML7
2020 Recognition Oriented Iris Image Quality Assessment in the Feature Space
abstract
A large portion of iris images captured in real world scenarios are poor quality due to the uncontrolled environment and the non-cooperative subject. To ensure that the recognition algorithm is not affected by low-quality images, traditional hand-crafted factors based methods discard most images, which will cause system timeout and disrupt user experience. In this paper, we propose a recognition-oriented quality metric and assessment method for iris image to deal with the problem. The method regards the iris image em-beddings Distance in Feature Space (DFS) as the quality metric and the prediction is based on deep neural networks with the attention mechanism. The quality metric proposed in this paper can significantly improve the performance of the recognition algorithm while reducing the number of images discarded for recognition, which is advantageous over hand-crafted factors based iris quality assessment methods. The relationship between Image Rejection Rate (IRR) and Equal Error Rate (EER) is proposed to evaluate the performance of the quality assessment algorithm under the same image quality distribution and the same recognition algorithm. Compared with hand-crafted factors based methods, the proposed method is a trial to bridge the gap between the image quality assessment and biometric recognition.
Leyuan Wang, Kunbo Zhang, Yunlong Wang 0003, Zhenan Sun
IJCB1
2019 A Unified Optimization Approach for CNN Model Inference on Integrated GPUs
abstract
Modern deep learning applications urge to push the model inference taking place at the edge devices for multiple reasons such as achieving shorter latency, relieving the burden of the network connecting to the cloud, and protecting user privacy. The Convolutional Neural Network (CNN) is one of the most widely used model family in the applications. Given the high computational complexity of the CNN models, it is favorable to execute them on the integrated GPUs at the edge devices, which are ubiquitous and have more power and better energy efficiency than the accompanying CPUs. However, programming on integrated GPUs efficiently is challenging due to the variety of their architectures and programming interfaces. This paper proposes an end-to-end solution to execute CNN model inference on the integrated GPUs at the edge, which uses a unified IR to represent and optimize vision-specific operators on integrated GPUs from multiple vendors, as well as leverages machine learning-based scheduling search schemes to optimize computationally-intensive operators like convolution. Our solution even provides a fallback mechanism for operators not suitable or convenient to run on GPUs. The evaluation results suggest that compared to state-of-the-art solutions backed up by the vendor-provided high-performance libraries on Intel Graphics, ARM Mali GPU, and Nvidia integrated Maxwell GPU, our solution achieves similar, or even better (up to 1.62×), performance on a number of popular image classification and object detection models. In addition, our solution has a wider model coverage and is more flexible to embrace new models. Our solution has been adopted in production services in AWS and is open-sourced.
Leyuan Wang, Zhi Chen 0030, Lianmin Zheng, Mu Li 0003, Yida Wang 0003
ICPP1
2018 TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
Tianqi Chen 0001, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy
OSDI8
2016 Fast parallel skew and prefix-doubling suffix array construction on the GPU
abstract
Summary Suffix arrays are fundamental full‐text index data structures of importance to a broad spectrum of applications in such fields as bioinformatics, Burrows–Wheeler transform‐based lossless data compression, and information retrieval. In this work, we propose and implement two massively parallel approaches on the graphics processing unit (GPU) based on two classes of suffix array construction algorithms. The first, parallelskew, makes algorithmic improvements to the previous work of Deo and Keely to achieve a speedup of 1.45x over their work. The second, a hybridskewandprefix‐doublingimplementation, is the first of its kind on the GPU and achieves a speedup of 2.3–4.4x over Osipov's prefix‐doubling and 2.4–7.9x over our skew implementation on large datasets. Our implementations rely on two efficient parallel primitives, a merge and a segmented sort. We theoretically analyze the two formulations of suffix array construction algorithms and show performance comparisons on a large variety of practical inputs. We conclude that, with the novel use of our efficient segmented sort,prefix‐doublingis more competitive thanskewon the GPU. We also demonstrate the effectiveness of our methods in our implementations of the Burrows‐Wheeler transform and in a parallel full‐text, minute‐space‐index for pattern searching. Copyright © 2016 John Wiley & Sons, Ltd.
Leyuan Wang, Sean Baxter, John D. Owens
Concurr. Comput. Pract. Exp.1
2015 Fast Parallel Suffix Array on the GPU
Leyuan Wang, Sean Baxter, John D. Owens
Euro-Par1