Wenze Shao

dblp:49/4514 · also Wen-Ze Shao · DBLP profile ↗
← Back
22ranked-venue papers
12as first author
9since 2021 · last 2025
0000-0001-6869-7789ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 LNLFace: Enhanced Blind Face Restoration With Local and Non-local Lookups
abstract
Existing reference-based blind face restoration (BFR) methods tend to either focus on the detailed texture of facial components or only conduct code prediction in the latent space, often neglecting their complementary relationship. To deal with it, the integration of local and non-local lookups that interact to ensure both fine-grained details and global geometric consistency is an eminently practical yet challenging solution. Specifically, Facial Component Dictionaries and High-Quality Feature Codebooks, which are pre-constructed from a large corpus of high-quality face images, perform their own functions of low- and high-level features. Thus, we first introduce an xLSTM-based Degradation-Aware Module (DAM) to mitigate code prediction biases suffering from degradation, then develop several two-stage Global Semantic Attention Modules (GSAM) to refine the multi-scale local details by the non-local insights. Finally, our proposed method, termed Local and Non-local Lookups for BFR (LNLFace), has demonstrated comparable or even superior performance than state-of-the-art methods, on both synthetic and real-world datasets. The codes are available at: https://github.com/yanwd628/LNLFace.
Weidan Yan, Wenze Shao, Dengyin Zhang
ICASSP2
2025 Explanation-Inspired Transferable Adversarial Attacks with Layer-Wise Increment Decomposition
Shizhe Xue, Yiyuan Chen, Wenze Shao
ICIC (4)3
2025 An Event-tailored State-Space Based Model for Pedestrian Detection
abstract
Event cameras, as emerging bio-inspired sensors, endow us with a unique scene perception capability with sub-millisecond latency in challenging environments, such as high-dynamic range and motion blur, as to which a plausible yet efficient exploration on spatiotemporal characteristics of the sparse, asynchronous event data remains an open problem. Event-based pedestrian detection, considered as a promising alternate for road safety in autonomous driving, is chosen as the testbed in this paper for pursuing a specific event-tailored spatiotemporal model. Note that, heterogeneous architectures are generally used in literature, such as building on a CNN/Transformer-style model for capturing the spatial features and a RNN/LSTM model for mining the temporal coherence, respectively. However, existing methods still face significant limitations, particularly as deployed in multi-rate dynamic environments, characterized by pronounced sparsity patterns in slow-motion or other scenarios. As such, a homogeneous neural network for robust pedestrian detection is proposed, with Event-tailored Recurrent Spatiotemporal State-Space Module (ERS3M) as the core innovation, for a joint meticulous modeling of spatiotemporal sparsity and dynamics over event data. On one hand, inspired by Vision Mamba, ERS3M is equipped with adaptive spatiotemporal state propagation as well as multi-directional compensatory scanning, enabling elaborate detection even as to observation intervals with extremely limited events triggered. On the other hand, ERS3 M is augmented with an additional block termed Temporal-Entropy Synergy, offering a collaborative spatiotemporal event purification mechanism, so as to enhance the probability credibility of event streams in visual semantics considering their complicated dynamics. Finally, ERS3M ends with an aliasing-alleviated S5 block to transit information between consecutive time steps, facilitating the temporal consistent pedestrian detection. Evaluations on the PEDRo dataset demonstrate that, the proposed detection method with ERS3 M as backbone has achieved a comparable or even superior performance to state-of-the-art approaches in terms of both accuracy and efficiency.
Liuyi Li, Jian Wang 0145, Jinjing Zhu, Wenze Shao
ACM Multimedia5
2025 FaceGCN: Structured Priors Inspired Graph Convolutional Networks for Face Restoration With Unknown Degradations
abstract
Facial image restoration has gained a tremendous progress since the increasing boom of the deep learning methods. Owing to its nature of strong ill-posedness, different categories of a-priori constraints have been harnessed or embedded in the existing deep architectures. While, as it turns to blind face restoration with more complicated degradations, the challenge becomes greater. In this paper, a further insightful step is taken by exploring the potentials of the graph convolutional networks (GCN) in conjunction with the structured priors for the blind problem. Specifically, a lightweight yet physically more intuitive model termed FaceGCN is proposed. On the one hand, a dynamic generator of facial adjacency matrices is constructed assisted by two self-supervised losses, allowing a sparse, accurate, and adaptive construction of case-specific face graphs with facial feature components as nodes. On the other hand, to model well the joint local-nonlocal correlations among various facial feature components, a kind of novel strip-attention GCN modules is correspondingly developed by splitting facial feature maps into intra- and inter-strips in both horizontal and vertical orientations, respectively. Extensive experimental results show that FaceGCN has achieved comparable or even superior performance to state-of-the-art methods, yet at a considerably less computational cost.
Weidan Yan, Wenze Shao, Dengyin Zhang, Liang Xiao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Boosting Adversarial Transferability via Relative Feature Importance-Aware Attacks
abstract
Modern deep neural networks are known highly vulnerable to adversarial examples. As a pioneering work, the fast gradient sign method (FGSM) is proved more transferable in black-box attacks than its multi-small-step extension, i.e., iterative-FGSM, particularly being restricted by a limited number of iterations. This paper revisits their early, representative successor MI-FGSM as a baseline, i.e., iterative-FGSM with momentum, and introduces an innovative boosting idea different from either FGSM-inspired algorithms or other mainstream methods. For one thing, during gradient backpropogation of MI-FGSM, the proposed approach merely requires amending the chain rule with respect to adversarial images using the counterpart original images. For another, a credible analysis has revealed that such a naively boosted MI-FGSM essentially performs a special kind of intermediate-layer attacks. In specific, the notable finding in the paper is a new principle of adversarial transferability guided by the relative feature importance, emphasizing the significance of semantically non-critical information for the first time in the literature, although originally thought to be weak in large. Experimental results on various leading victim models, both undefended and defended, demonstrate that the new approach incorporating robust gradients has indeed attained stronger adversarial transferability than state-of-the-art works. The code is available at:https://github.com/ljwooo/RFIA-main.
Wenze Shao, Yubao Sun, Li-Qian Wang, Qi Ge, Liang Xiao 0001
IEEE Trans. Inf. Forensics Secur.2
2024 M2Mamba: Multi-Scale in Multi-Scale Mamba for Blind Face Restoration
abstract
With the evolution of telemedicine, clean images, especially facial images, are crucial in areas such as symptom evaluation and cosmetic medicine. However, dealing with indiscernible facial images, a condition known as blind degradation, poses a formidable challenge for existing blind face restoration (BFR) techniques due to their inherently ill-posed nature. Thus, a delicate interaction between preserving local details and maintaining global geometries is desired. Despite advances in convolutional neural networks (CNNs), transformers, and denoising diffusion, where no matter serried convolutions, elaborately-designed self-attentions, or stochastic noises, tend to isolate either local or non-local features. To this end, a new candidate model termed Multi-Scale in Multi-Scale Mamba (M2Mamba) is proposed, which builds on a pioneering structured state space model-based network and includes three new components: multi-scale learned fusion module (MSLFM), multi-scale attention fusion module (MSAFM), and multi-scale inspired Mamba (MS-Mamba). Firstly, the MSLFM is adopted by extracting image-level global guidance from inputs of different scales, preserving intuitive perception yet enriching semantic understanding. Secondly, the MSAFM dynamically integrates features from various encoder stages. Thirdly, the MS-Mamba employs separate branches for both small- and large-scale receptive fields, benefiting modeling of long-range dependencies. In the final, M2Mamba is demonstrated on both synthetic and realistic benchmarks, showing comparable or better performance than state-of-the-art methods. The codes are available at: https://github.com/yanwd628/M2Mamba.
Weidan Yan, Wenze Shao, Dengyin Zhang
BIBM2
2024 A novel regularization method for decorrelation learning of non-parallel hyperplanes
Wenze Shao, Yuan-Hai Shao 0001, Chun-Na Li 0001
Inf. Sci.1
2024 Revisiting reweighted graph total variation blind deconvolution and beyond
Wenze Shao, Haisong Deng, Wei-Wei Luo, Jinye Li, Mei-Lin Liu
Vis. Comput.1
2023 Revisiting the Regularizers in Blind Image Deblurring With a New One
abstract
Image deblurring and its counterpart blind problem are undoubtedly two fundamental tasks in computational imaging and computer vision. Interestingly, deterministic edge-preserving regularization for maximum-a-posteriori (MAP) based non-blind image deblurring has been largely made clear 25 years ago. As for the blind task, the state-of-the-art MAP-based approaches seem to also reach a consensus on the characteristic of deterministic image regularization, i.e., formulated in anL0composite style or termed asL0+Xstyle, whereXis often a discriminative term such as dark channels-based sparsity regularization. However, with a modeling perspective as such, non-blind and blind deblurring are entirely disconnected from each other. Additionally, becauseL0andXare motivated very differently in general, it is not easy in practice to derive an efficient numerical scheme. In fact, since the prosperity of modern blind deblurring 15 years ago, a physically intuitive yet practically effective and efficient regularization has been always desired. In this paper, representative deterministic image regularization terms in MAP-based blind deblurring are firstly revisited, with an emphasis on their differences from edge-preserving regularization for non-blind deblurring. Inspired by existing robust losses in the statistical and deep learning literature, an insightful conjecture is then made. That is, deterministic image regularization for blind deblurring can be naively formulated using a type of redescending potential functions (RDP), and interestingly, a RDP-induced blind deblurring regularization term is actually the 1rst-order derivative of a nonconvex edge-preserving regularization for non-blind image deblurring. An intimate relationship in regularization is therefore established between the two problems, differing much from the mainstream modeling perspective on blind deblurring. Via above principle analysis, the conjecture is demonstrated on benchmark deblurring problems in the final, accompanied with comparisons against several top-performingL0+Xstyle methods. We note that, the rationality and practicality of the RDP-induced regularization is particularly highlighted here, aiming to open up an alternative line of possibility for modeling blind deblurring.
Wenze Shao
IEEE Trans. Image Process.1
2020 Gradient-based discriminative modeling for blind image deblurring
Wenze Shao, Yunzhi Lin, Li-Qian Wang, Qi Ge, Bing-Kun Bao, Haibo Li 0001
Neurocomputing1
2020 DeblurGAN+: Revisiting blind motion deblurring using conditional adversarial networks
Wenze Shao, Lu-Yue Ye, Li-Qian Wang, Qi Ge, Bing-Kun Bao, Haibo Li 0001
Signal Process.1
2019 Adversarial Representation Learning for Dynamic Scene Deblurring: A Simple, Fast and Robust Approach
abstract
In this paper, we investigate a novel learning-based method for dynamic scene deblurring. Since the inference model is formulated as an encoder-decoder, the core task has turned to learning blur-invariant hidden features from the blurred images to a great degree. To achieve state-of-the-art results in terms of both deblurring accuracy and efficiency, a simple, robust and computationally efficient deep auto-encoder is developed tailored specifically for blind deblurring, which is learned in an adversarial fashion based on use of the recent Wasserstein generative adversarial networks. Thanks to the designed framework, the new model has shown comparable or superior performance both qualitatively and quantitatively to existing state-of-the-art methods. What is more important, instead of exploiting the multi-scale strategy as previous methods, our model is just single-scale capable of achieving 6 times more efficiency than the closest competitor by Nah et al. [1], which is a more complex multi-scale deep method.
Lu-Yue Ye, Wenze Shao, Qi Ge, Li-Qian Wang, Bing-Kun Bao, Haibo Li 0001
ICIP3
2019 On potentials of regularized Wasserstein generative adversarial networks for realistic hallucination of tiny faces
Wenze Shao, Qi Ge, Li-Qian Wang, Bing-Kun Bao, Haibo Li 0001
Neurocomputing1
2019 Enhancing Blurred Low-Resolution Images via Exploring the Potentials of Learning-Based Super-Resolution
abstract
This paper aims to propose a candidate solution to the challenging task of single-image blind super-resolution (SR), via extensively exploring the potentials of learning-based SR schemes in the literature. The task is formulated into an energy functional to be minimized with respect to both an intermediate super-resolved image and a nonparametric blur-kernel. The functional includes a so-called convolutional consistency term which incorporates a nonblind learning-based SR result to better guide the kernel estimation process, and a bi-[Formula: see text]-[Formula: see text]-norm regularization imposed on both the super-resolved sharp image and the nonparametric blur-kernel. A numerical algorithm is deduced via coupling the splitting augmented Lagrangian (SAL) and the conjugate gradient (CG) method. With the estimated blur-kernel, the final SR image is reconstructed using a simple TV-based nonblind SR method. The proposed blind SR approach is demonstrated to achieve better performance than [T. Michaeli and M. Irani, Nonparametric Blind Super-resolution, in Proc. IEEE Conf. Comput. Vision (IEEE Press, Washington, 2013), pp. 945–952.] in terms of both blur-kernel estimation accuracy and image ehancement quality. In the meanwhile, the experimental results demonstrate surprisingly that the local linear regression-based SR method, anchored neighbor regression (ANR) serves the proposed functional more appropriately than those harnessing the deep convolutional neural networks.
Wenze Shao, Bing-Kun Bao, Haibo Li 0001
Int. J. Pattern Recognit. Artif. Intell.1
2018 Blind Deblurring Using Discriminative Image Smoothing
Wenze Shao, Yunzhi Lin, Bing-Kun Bao, Liqian Wang, Qi Ge, Haibo Li 0001
PRCV (1)1
2017 Color demosaicking via nonlocal tensor representation
abstract
A single sensor camera can capture scenes by means of color filter array. Each pixel samples only one of the three primary colors. Color demosaicking (CDM) is a process of reconstruction a full color image from this sensor data. In this paper, we propose a novel CDM scheme based on learned simultaneous sparse coding over nonlocal tensor representation. First, similar 2D patches are grouped to form a three-order tensor, that is, 3D array. Then, three sub-dictionaries, which characterize the coherent structures that appear in each dimension of the grouped tensor, are learned jointly by using Tucker decomposition. The consequent coefficient tensor is imposed by the grouped-block-sparsity constraint, which forces the similar patches to share the same atoms of the dictionaries in their sparse decomposition. Experimental results demonstrate the effectiveness both in the average CPSNR and visual quality.
Wenze Shao, Hongyi Liu 0001, Zhihui Wei, Liang Xiao 0001
ICASSP3
2017 Structure-Based Low-Rank Model With Graph Nuclear Norm Regularization for Noise Removal
abstract
Nonlocal image representation methods, including group-based sparse coding and block-matching 3-D filtering, have shown their great performance in application to low-level tasks. The nonlocal prior is extracted from each group consisting of patches with similar intensities. Grouping patches based on intensity similarity, however, gives rise to disturbance and inaccuracy in estimation of the true images. To address this problem, we propose a structure-based low-rank model with graph nuclear norm regularization. We exploit the local manifold structure inside a patch and group the patches by the distance metric of manifold structure. With the manifold structure information, a graph nuclear norm regularization is established and incorporated into a low-rank approximation model. We then prove that the graph-based regularization is equivalent to a weighted nuclear norm and the proposed model can be solved by a weighted singular-value thresholding algorithm. Extensive experiments on additive white Gaussian noise removal and mixed noise removal demonstrate that the proposed method achieves a better performance than several state-of-the-art algorithms.
Qi Ge, Xiaoyuan Jing, Fei Wu 0004, Zhihui Wei, Liang Xiao 0001, Wenze Shao, Dong Yue 0001, Haibo Li 0001
IEEE Trans. Image Process.6
2016 Regularized motion blur-kernel estimation with adaptive sparse image prior learning
Wenze Shao, Haisong Deng, Qi Ge, Haibo Li 0001, Zhihui Wei
Pattern Recognit.1
2015 Simple, Accurate, and Robust Nonparametric Blind Super-Resolution
Wenze Shao, Michael Elad
ICIG (3)1
2015 Bi-l0-l2-norm regularization for blind motion deblurring
Wenze Shao, Haibo Li 0001, Michael Elad
J. Vis. Commun. Image Represent.1
2015 A hybrid active contour model with structured feature for image segmentation
Qi Ge, Chuansong Li, Wenze Shao, Haibo Li 0001
Signal Process.3
2008 Edge-and-corner preserving regularization for image interpolation and reconstruction
Wenze Shao, Zhihui Wei
Image Vis. Comput.1