Jian Lu 0002

dblp:225/6330-2 · DBLP profile ↗
← Back
43ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0003-4599-7281ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 17 · 12 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Weighted learning via logarithmic sparsity and hyper-Laplacian regularization for multi-view subspace clustering
Min Li 0024, Jian Lu 0002, Mingqing Xiao 0001
Neurocomputing3
2026 Synchrosqueezed windowed linear canonical transform: A method for mode retrieval from multicomponent signals with crossing instantaneous frequencies
Shuixin Li, Jiecheng Chen, Qingtang Jiang, Jian Lu 0002
Signal Process.4
2026 Dual-feature attention for robust point cloud registration: integrating transformation-variant and invariant features
Saifullahi Aminu Bello, Ali Aljofey, Jian Lu 0002, Chen Xu 0004, Yuru Zou
Vis. Comput.3
2025 Geometric edge convolution for rigid transformation invariant features in 3D point clouds
Saifullahi Aminu Bello, Saghir Ahmed Saghir Alfasly, Jian Lu 0002, Lin Li 0050, Chen Xu 0004, Yuru Zou
Neurocomputing4
2025 Comprehensive phishing detection: A multi-channel approach with variants TCN fusion leveraging URL and HTML features
Ali Aljofey, Saifullahi Aminu Bello, Jian Lu 0002, Chen Xu 0004
J. Netw. Comput. Appl.3
2025 EMBANet: A flexible efficient multi-branch attention network
abstract
Recent advances in the design of convolutional neural networks have shown that performance can be enhanced by improving the ability to represent multi-scale features. However, most existing methods either focus on designing more sophisticated attention modules, which leads to higher computational costs, or fail to effectively establish long-range channel dependencies, or neglect the extraction and utilization of structural information. This work introduces a novel module, the Multi-Branch Concatenation (MBC), designed to process input tensors and extract multi-scale feature maps. The MBC module introduces new degrees of freedom (DoF) in the design of attention networks by allowing for flexible adjustments to the types of transformation operators and the number of branches. This study considers two key transformation operators: multiplexing and splitting, both of which facilitate a more granular representation of multi-scale features and enhance the receptive field range. By integrating the MBC with an attention module, a Multi-Branch Attention (MBA) module is developed to capture channel-wise interactions within feature maps, thereby establishing long-range channel dependencies. Replacing the 3x3 convolutions in the bottleneck blocks of ResNet with the proposed MBA yields a new block, the Efficient Multi-Branch Attention (EMBA), which can be seamlessly integrated into state-of-the-art backbone CNN models. Furthermore, a new backbone network, named EMBANet, is constructed by stacking EMBA blocks. The proposed EMBANet has been thoroughly evaluated across various computer vision tasks, including classification, detection, and segmentation, consistently demonstrating superior performance compared to popular backbones.
Keke Zu, Lei Zhang 0006, Jian Lu 0002, Chen Xu 0004, Hongyang Chen 0001, Yu Zheng 0004
Neural Networks4
2025 Orthogonal Constrained Minimization with Tensor \(\ell_{2,{p}}\) Regularization for HSI Denoising and Destriping
abstract
Abstract. Hyperspectral images (HSIs) are often contaminated by a mixture of noise such as Gaussian noise, dead lines, stripes, and so on. In this paper, we propose a multiscale low-rank tensor regularized [Formula: see text] (MLTL2p) approach for HSI denoising and destriping, which consists of an orthogonal constrained minimization model and an iterative algorithm with convergence guarantees. The model of the proposed MLTL2p approach is built based on a new sparsity-enhanced Multiscale Low-rank Tensor regularization and a tensor [Formula: see text] norm with [Formula: see text]. The multiscale low-rank regularization for HSI denoising utilizes the global and local spectral correlation as well as the spatial nonlocal self-similarity priors of HSIs. The corresponding low-rank constraints are formulated based on independent higher-order singular value decomposition with sparsity enhancement on its core tensor to prompt more low-rankness. The tensor [Formula: see text] norm for HSI destriping is extended from the matrix [Formula: see text] norm. A proximal block coordinate descent algorithm is proposed in the MLTL2p approach to solve the resulting nonconvex nonsmooth minimization with orthogonal constraints. We show any accumulation point of the sequence generated by the proposed algorithm converges to a first-order stationary point, which is defined using three equalities of substationarity, symmetry, and feasibility for orthogonal constraints. In the numerical experiments, we compare the proposed method with state-of-the-art methods, including a deep learning based method, and test the methods on both simulated and real HSI datasets. Our proposed MLTL2p method demonstrates outperformance in terms of metrics such as mean peak signal-to-noise ratio as well as visual quality.
Shijie Yu, Jian Lu 0002, Xiaojun Chen 0001
SIAM J. Imaging Sci.3
2025 BERT-PhishFinder: A Robust Model for Accurate Phishing URL Detection With Optimized DistilBERT
abstract
Phishing URL detection has become a critical challenge in cybersecurity, with existing methods often struggling to maintain high accuracy while generalizing across diverse datasets. In this article, we introduce BERT-PhishFinder, a novel and efficient transformer-based model designed to tackle this problem. While most traditional approaches rely heavily on lexical features or complex convolutional architectures, BERT-PhishFinder leverages the power of DistilBERT, a lightweight yet highly effective transformer, to capture rich contextual representations of URL sequences. To enhance the model’s robustness and reduce overfitting, we strategically incorporate SpatialDropout1D in the embedding layers, along with global average pooling and global max pooling techniques to extract both comprehensive and key discriminative features. The pooled representations are thoughtfully concatenated to form a comprehensive feature representation. Through this carefully crafted design, our model adopts ensemble learning, as it undergoes multiple parallel dense layers, each with distinct parameters and dropout regularization. This facilitates learning diverse patterns and features from the input URL sequence, culminating in exceptional phishing URL detection performance. Extensive evaluations against conventional deep learning algorithms, transformer models (XLNet, RoBERTa, ALBERT), and other existing methods on five benchmark datasets show that BERT-PhishFinder not only achieves the state of-the-art real phishing URL detection but also accomplishes this with reduced label dependency.
Ali Aljofey, Saifullahi Aminu Bello, Jian Lu 0002, Chen Xu 0004
IEEE Trans. Dependable Secur. Comput.3
2024 Auxiliary audio-textual modalities for better action recognition on vision-specific annotated videos
Saghir Ahmed Saghir Alfasly, Jian Lu 0002, Chen Xu 0004, Yuru Zou
Pattern Recognit.2
2024 Stochastic Variance Reduced Gradient for Affine Rank Minimization Problem
abstract
Abstract. In this paper, we develop an efficient stochastic variance reduced gradient descent algorithm to solve the affine rank minimization problem consisting of finding a matrix of minimum rank from linear measurements. The proposed algorithm as a stochastic gradient descent strategy enjoys a more favorable complexity than that using full gradients. It also reduces the variance of the stochastic gradient at each iteration and accelerates the rate of convergence. We prove that the proposed algorithm converges linearly in expectation to the solution under a restricted isometry condition. Numerical experimental results demonstrate that the proposed algorithm has a clear advantageous balance of efficiency, adaptivity, and accuracy compared with other state-of-the-art algorithms.
Ningning Han, Juan Nie, Jian Lu 0002, Michael Kwok-Po Ng
SIAM J. Imaging Sci.3
2024 Tensor recovery using the tensor nuclear norm based on nonconvex and nonlinear transformations
Zhihui Tu, Kaitao Yang, Jian Lu 0002, Qingtang Jiang
Signal Process.3
2024 Cyclic tensor singular value decomposition with applications in low-rank high-order tensor recovery
Yigong Zhang, Zhihui Tu, Jian Lu 0002, Chen Xu 0004, Michael Kwok-Po Ng
Signal Process.3
2024 OSRE: Object-to-Spot Rotation Estimation for Bike Parking Assessment
abstract
Current deep models excel in object detection for classification and localization. However, precise object rotation estimation within the visual context of an input image remains underexplored due to the lack of object datasets with rotation annotations. This paper addresses these challenges by tackling rotation estimation for parked bikes with respect to their parking area. Firstly, 3D graphics were leveraged to build a camera-agnostic well-annotated Synthetic Bike Rotation Dataset (SynthBRSet). Subsequently, an object-to-spot rotation estimator (OSRE) is introduced by extending object detection to regress bike rotations in two axes. As the proposed model trained purely on synthetic data, image smoothing techniques adopted during deployment on real-world images. The proposed OSRE has undergone evaluation on both synthetic and real-world data, showing promising results. Our data and code are available at https://saghiralfasly.github.io/OSRE-Project/.
Saghir Ahmed Saghir Alfasly, Zaid Al-Huda, Saifullahi Aminu Bello, Ahmed El-Azab, Jian Lu 0002, Chen Xu 0004
IEEE Trans. Intell. Transp. Syst.5
2024 An Effective Video Transformer With Synchronized Spatiotemporal and Spatial Self-Attention for Action Recognition
abstract
Convolutional neural networks (CNNs) have come to dominate vision-based deep neural network structures in both image and video models over the past decade. However, convolution-free vision Transformers (ViTs) have recently outperformed CNN-based models in image recognition. Despite this progress, building and designing video Transformers have not yet obtained the same attention in research as image-based Transformers. While there have been attempts to build video Transformers by adapting image-based Transformers for video understanding, these Transformers still lack efficiency due to the large gap between CNN-based models and Transformers regarding the number of parameters and the training settings. In this work, we propose three techniques to improve video understanding with video Transformers. First, to derive better spatiotemporal feature representation, we propose a new spatiotemporal attention scheme, termed synchronized spatiotemporal and spatial attention (SSTSA), which derives the spatiotemporal features with temporal and spatial multiheaded self-attention (MSA) modules. It also preserves the best spatial attention by another spatial self-attention module in parallel, thereby resulting in an effective Transformer encoder. Second, a motion spotlighting module is proposed to embed the short-term motion of the consecutive input frames to the regular RGB input, which is then processed with a single-stream video Transformer. Third, a simple intraclass frame interlacing method of the input clips is proposed that serves as an effective video augmentation method. Finally, our proposed techniques have been evaluated and validated with a set of extensive experiments in this study. Our video Transformer outperforms its previous counterparts on two well-known datasets, Kinetics400 and Something-Something-v2.
Saghir Ahmed Saghir Alfasly, Charles K. Chui, Qingtang Jiang, Jian Lu 0002, Chen Xu 0004
IEEE Trans. Neural Networks Learn. Syst.4
2023 FastPicker: Adaptive independent two-stage video-to-video summarization for efficient action recognition
Saghir Ahmed Saghir Alfasly, Jian Lu 0002, Chen Xu 0004, Zaid Al-Huda, Qingtang Jiang, Zhaosong Lu, Charles K. Chui
Neurocomputing2
2023 A meta-framework for multi-label active learning based on deep reinforcement learning
Shuyue Chen, Ran Wang 0001, Jian Lu 0002
Neural Networks3
2023 Nonlocal ultrasound image despeckling via improved statistics and rank constraint
Hanmei Yang, Jian Lu 0002, Ye Luo 0004, Heng Zhang 0014
Pattern Anal. Appl.2
2023 Generalized Tag-Based Physical-Layer Authentication Under Frequency Selective Fading Channels
abstract
Physical Layer Authentication (PLA) in wireless systems has drawn much attention because it can provide information-theoretic security. However, the performance of the prior PLA schemes significantly declines under a frequency selective fading channel and the theoretical analysis of a PLA scheme becomes extremely challenging under a frequency selective fading channel as compared to that under a frequency flat fading channel. This paper addresses the problem of authenticating transmitters at the physical layer under a frequency selective fading channel. We propose two tag-based PLA schemes under a frequency selective fading channel. The first scheme is the Generalized Tag-based PLA (GT-PLA) scheme, where the negative effect caused by a frequency selective fading channel is compensated for. If the channel reciprocity holds, we will propose the second scheme to further improve the performance of the GT-PLA scheme, named the Adaptive Generalized Tag-based PLA (AGT-PLA) scheme. We provide the theoretical analysis of both proposed schemes over wireless fading channels in terms of robustness, security, and compatibility, and derive their closed-form expressions. Moreover, we implement the proposed schemes and conduct extensive performance comparisons. Our experimental results demonstrate that the GT-PLA scheme achieves higher robustness over the prior PLA schemes while the AGT-PLA scheme can further improve the performance of the GT-PLA scheme in terms of robustness, security, and compatibility.
Haijun Tan, Ning Xie 0007, Jian Lu 0002, Dusit Niyato
IEEE Trans. Commun.3
2023 Physical Layer Authentication in Spatial Modulation
abstract
Spatial Modulation (SM) is a promising low-complexity modulation scheme for Multiple-Input Multiple-Output (MIMO) systems. In this paper, we address the problem of authenticating the transmitter device in the SM. We propose an authentication approach for an SM system by using Physical-Layer Authentication (PLA) mechanisms because the PLA has the following advantages: high security and low complexity. Based on the features of an SM system, we propose two PLA schemes:PLA with Superimposed Authentication Tag(PLA-SAT) andPLA with Superimposed Imaginary authentication Tag(PLA-SIT). We provide performance analyses of our schemes over fading channels in terms of robustness, compatibility, and security. Moreover, we derive their closed-form expressions under both perfect and imperfect channel estimates, including the Probability of Detection (PD), Probability of False Alarm (PFA), and Average Error Probability (AEP). Although the two proposed schemes have the same robustness and security, the PLA-SIT scheme has better compatibility than the PLA-SAT scheme. Our schemes were implemented and extensive performance comparisons through simulations were conducted. We observe that the simulation results of the two proposed schemes perfectly match their corresponding theoretical analyses. The authentication accuracy of the two proposed schemes is close to one when the received SNR is greater than 20 dB and the security performances of the two proposed schemes improve as the variance of estimation errors increases.
Jiaheng Zhang, Qihong Zhang, Peichang Zhang, Lei Huang 0001, Ning Xie 0007, Jian Lu 0002
IEEE Trans. Commun.7
2023 Detection of Jamming Attacks for the Physical-Layer Authentication
abstract
This paper concerns the problem of defending against the jamming attack for the Physical-Layer Authentication (PLA). The problem is important due to the fact that jamming attacks can make the legitimate receiver reject a legitimate signal or accept an impersonating signal. In this paper, we propose two jamming detection schemes to defend against jamming attacks for the PLA. The first scheme is the Jamming-Attack Detection (JAD) scheme, which exploits the difference of noise variances between the handshaking and communication stages. The second scheme is the Composite Jamming-Attack Detection (CJAD) scheme, which exploits the difference of noise variances between Alice and Bob. Moreover, we provide the theoretical analysis of the proposed schemes over wireless fading channels and derive their closed-form expressions. We implement our schemes and conduct extensive performance comparisons. Our theoretical analysis and simulation results show that if security is the priority, the CJAD scheme is the best option. If we further consider the overhead, the JAD scheme is the best option under the strategy of jamming attacks and the CJAD scheme may be a better option under the strategy of enhanced jamming attacks, especially for the larger power of a jamming signal.
Haijun Tan, Ning Xie 0007, Jian Lu 0002, Dusit Niyato
IEEE Trans. Wirel. Commun.4
2022 EPSANet: An Efficient Pyramid Squeeze Attention Block on Convolutional Neural Network
Keke Zu, Jian Lu 0002, Yuru Zou, Deyu Meng
ACCV (3)3
2022 Learnable Irrelevant Modality Dropout for Multimodal Action Recognition on Modality-Specific Annotated Videos
abstract
With the assumption that a video dataset is multimodality annotated in which auditory and visual modalities both are labeled or class-relevant, current multimodal methods apply modality fusion or cross-modality attention. However, effectively leveraging the audio modality in vision-specific annotated videos for action recognition is of particular challenge. To tackle this challenge, we propose a novel audio-visual framework that effectively leverages the audio modality in any solely vision-specific annotated dataset. We adopt the language models (e.g., BERT) to build a semantic audio-video label dictionary (SAVLD) that maps each video label to its most K-relevant audio labels in which SAVLD serves as a bridge between audio and video datasets. Then, SAVLD along with a pretrained audio multi-label model are used to estimate the audio-visual modality relevance during the training phase. Accordingly, a novel learnable irrelevant modality dropout (IMD) is proposed to completely drop out the irrelevant audio modality and fuse only the relevant modalities. Moreover, we present a new two-stream video Transformer for efficiently modeling the visual modalities. Results on several vision-specific annotated datasets including Kinetics400 and UCF-101 validated our framework as it outperforms most relevant action recognition methods.
Saghir Ahmed Saghir Alfasly, Jian Lu 0002, Chen Xu 0004, Yuru Zou
CVPR2
2022 Stable matching-based two-way selection in multi-label active learning with imbalanced data
Shuyue Chen, Ran Wang 0001, Jian Lu 0002, Xizhao Wang
Inf. Sci.3
2022 H∞ stabilization problem for memristive neural networks with time-varying delays
Imran Ghous, Jian Lu 0002, Zhaoxia Duan
Inf. Sci.2
2022 Deep image prior based defense against adversarial examples
Tao Dai 0001, Bin Chen 0011, Jian Lu 0002, Shutao Xia
Pattern Recognit.4
2021 HOCA: Higher-Order Channel Attention for Single Image Super-Resolution
abstract
Convolutional neural networks (CNNs) have obtained great success in single image super-resolution (SR). More recent works (e.g., RCAN and SAN) have obtained remarkable performance with channel attention based on first- or second-order statistics of features. However, these methods neglect the rich feature statistics higher than second-order, thus hindering the representation ability of CNNs. To address this issue, we propose a higher-order channel attention (HOCA) module to enhance the representation ability of CNNs. In our HOCA module, to capture different types of semantic information, we first compute k-order of feature statistics, followed by channel attention to learn the feature interdependencies. Considering the diversity of input contents, we design a gate mechanism to adaptively select a specific k-order channel attention. Besides, our HOCA module serves as a plug-and-play module and can be easily plugged into existing state-of-art CNN-based SR methods. Extensive experiments on public benchmarks show that our HOCA module effectively improves the performance of various CNN-based SR methods.
Yalei Lv, Tao Dai 0001, Bin Chen 0011, Jian Lu 0002, Shutao Xia, Jingchao Cao
ICASSP4
2021 Correlation-based structural dropout for convolutional neural networks
Yuyuan Zeng, Tao Dai 0001, Bin Chen 0011, Shutao Xia, Jian Lu 0002
Pattern Recognit.5
2021 Orthogonal Subspace Based Fast Iterative Thresholding Algorithms for Joint Sparsity Recovery
abstract
Sparse signal recoveries from multiple measurement vectors (MMV) with joint sparsity property have many applications in signal, image, and video processing. The problem becomes much more involved when snapshots of the signal matrix are temporally correlated. With signal's temporal correlation in mind, we provide a framework of iterative MMV algorithms based on thresholding, functional feedback and null space tuning. Convergence analysis for exact recovery is established. Unlike most of iterative greedy algorithms that select indices in a measurement/solution space, we determine indices based on an orthogonal subspace spanned by the iterative sequence. In addition, a functional feedback that controls the amount of energy relocation from the “tails” is implemented and analyzed. It is seen that the principle of functional feedback is capable to lower the number of iteration and speed up the convergence of the algorithm. Numerical experiments demonstrate that the proposed algorithm has a clearly advantageous balance of efficiency, adaptivity and accuracy compared with other state-of-the-art algorithms.
Ningning Han, Shidong Li, Jian Lu 0002
IEEE Signal Process. Lett.3
2020 Enhanced Image Restoration Via Supervised Target Feature Transfer
abstract
Deep learning has obtained remarkable success for image restoration. However, most existing deep image restoration models are trained by minimizing the pixel-level reconstruction error between restored images and target images (ground truth), while neglecting the rich information from the intermediate feature layers, thus hindering the representational power of networks. To address this problem, we propose a Supervised Target Feature Transfer (STFT) framework to enhance the power of feature expression of the deep image restoration models. Specifically, we introduce a self-supervised antoencoder-based target feature extractor to extract compact feature representation of target images, which serves as supervision signals to train the deep backbone models at the same time. With such feature-level supervised information, deep backbone model can be enhanced by transfer learning of such target features. Moreover, we theoretically analyze our STFT training strategies and demonstrate that it imposes learnable prior information on the backbone restoration model. Extensive experiments demonstrate the effectiveness of our proposed framework compared with the state-of-the-art image restoration models.
Yuzhao Chen, Tao Dai 0001, Xi Xiao 0001, Jian Lu 0002, Shutao Xia
ICIP4
2020 Fakd: Feature-Affinity Based Knowledge Distillation for Efficient Image Super-Resolution
abstract
Convolutional neural networks (CNNs) have been widely used in image super-resolution (SR). Most existing CNN-based methods focus on achieving better performance by designing deeper/wider networks, while suffering from heavy computational cost problem, thus hindering the deployment of such models in mobile devices with limited resources. To relieve such problem, we propose a novel and efficient SR model, named Feature Affinity-based Knowledge Distillation (FAKD), by transferring the structural knowledge of a heavy teacher model to a lightweight student model. To transfer the structural knowledge effectively, FAKD aims to distill the second-order statistical information from feature maps and trains a lightweight student network with low computational and memory cost. Experimental results demonstrate the efficacy of our method and the effectiveness over other knowledge distillation based methods in terms of both quantitative and visual metrics.
Zibin He, Tao Dai 0001, Jian Lu 0002, Yong Jiang 0001, Shutao Xia
ICIP3
2020 Sample-aware Data Augmentor for Scene Text Recognition
abstract
Deep neural networks (DNNs) have been widely used in scene text recognition, and achieved remarkable performance. Such DNN-based scene text recognizers usually require plenty of training data for training, but data collection and annotation is usually cost-expensive in practice. To alleviate this issue, data augmentation is often applied to train the scene text recognizers. However, existing data augmentation methods including affine transformation and elastic transformation methods suffer from the problems of under- and over-diversity, due to the complexity of text contents and shapes. In this paper, we propose a sample-aware data augmentor to transform samples adaptively based on the contents of samples. Specifically, our data augmentor consists of three parts: gated module, affine transformation module, and elastic transformation module. In our data augmentor, affine transformation module focuses on keeping the affinity of samples, while elastic transformation module aims to improve the diversity of samples. With the gated module, our data augmentor determines transformation type adaptively based on the properties of training samples and the recognizer capability during the training process. Besides, our framework introduces an adversarial learning strategy to optimize the augmentor and the recognizer jointly. Extensive experiments on scene text recognition benchmarks show that our sample-aware data augmentor significantly improves the performance of state-of-the-art scene text recognizer.
Guanghao Meng, Tao Dai 0001, Shudeng Wu, Bin Chen 0011, Jian Lu 0002, Yong Jiang 0001, Shutao Xia
ICPR5
2020 Transferable Adversarial Attacks for Deep Scene Text Detection
abstract
Scene text detection (STD) aims to locate text in images and plays an important role in many computer vision tasks including automatic driving and text recognition systems. Recently, deep neural networks (DNNs) have been widely and successfully used in scene text detection, leading to plenty of DNN-based STD methods including regression-based and segmentation-based STD methods. However, recent studies have also shown that DNN is vulnerable to adversarial attacks, which can significantly degrade the performance of DNN models. In this paper, we investigate the robustness of DNN-based STD methods against adversarial attacks. To this end, we propose a generic and efficient attack method to generate adversarial examples, which are produced by adding small but imperceptible adversarial perturbation to the input images. Experiments on attacking four various models and a real-world STD engine of Google optical character recognition (OCR) show that the state-of-the-art DNN-based STD methods including regression-based and segmentation-based methods are vulnerable to adversarial attacks.
Shudeng Wu, Tao Dai 0001, Guanghao Meng, Bin Chen 0011, Jian Lu 0002, Shutao Xia
ICPR5
2020 Ultrasound Image Restoration Using Weighted Nuclear Norm Minimization
abstract
Ultrasound images are often contaminated by speckle noise during the acquisition process, which influences the performance of subsequent applications. The paper introduces a nonconvex low-rank matrix approximation model for ultrasound images restoration, which integrates the weighted nuclear norm minimization (WNNM) and data fidelity term. WNNM can adaptively assign weights on different singular values to preserve more details in restored images. The fidelity term about ultrasound images do not be utilized in existing low-rank ultrasound denoising methods. This optimization question can effectively solved by alternating direction method of multipliers (ADMM). The experimental results on simulated images and real medical ultrasound images demonstrate the excellent performance of the proposed method compared with other four state-of-the-art methods.
Hanmei Yang, Heng Zhang 0014, Ye Luo 0004, Jian Lu 0002
ICPR5
2020 DIPDefend: Deep Image Prior Driven Defense against Adversarial Examples
abstract
Deep neural networks (DNNs) have shown serious vulnerability to adversarial examples with imperceptible perturbation to clean images. Most existing input-transformation based defense methods (e.g., ComDefend) rely heavily on the learned external priors from an external large training dataset, while neglecting the rich image internal priors of the input itself, thus limiting the generalization of the defense models against the adversarial examples with biased image statistics from the external training dataset. Motivated by deep image prior that can capture rich image statistics from a single image, we propose an effective Deep Image Prior Driven Defense (DIPDefend) method against adversarial examples. With a DIP generator to fit the target/adversarial input, we find that our image reconstruction exhibits quite interesting learning preference from a feature learning perspectives, i.e., the early stage primarily learns the robust features resistant to adversarial perturbation, followed by learning non-robust features that are sensitive to adversarial perturbation. Besides, we develop an adaptive stopping strategy that adapts our method to diverse images. In this way, the proposed model obtains a unique defender for each individual adversarial input, thus being robust to various attackers. Experimental results demonstrate the superiority of our method over the state-of-the-art defense methods against white-box and black-box adversarial attacks.
Tao Dai 0001, Dongxian Wu, Bin Chen 0011, Jian Lu 0002, Yong Jiang 0001, Shutao Xia
ACM Multimedia5
2020 Multiplicative Noise Removal: Nonlocal Low-Rank Model and Its Proximal Alternating Reweighted Minimization Algorithm
abstract
The goal of this paper is to develop a novel numerical method for efficient multiplicative noise removal. The nonlocal self-similarity of natural images implies that the matrices formed by their nonlocal similar patches are low-rank. By exploiting this low-rank prior with application to multiplicative noise removal, we propose a nonlocal low-rank model for this task and develop a proximal alternating reweighted minimization (PARM) algorithm to solve the optimization problem resulting from the model. Specifically, we utilize a generalized nonconvex surrogate of the rank function to regularize the patch matrices and develop a new nonlocal low-rank model, which is a nonconvex nonsmooth optimization problem having a patchwise data fidelity and a generalized nonlocal low-rank regularization term. To solve this optimization problem, we propose the PARM algorithm, which has a proximal alternating scheme with a reweighted approximation of its subproblem. A theoretical analysis of the proposed PARM algorithm is conducted to guarantee its global convergence to a critical point. Numerical experiments demonstrate that the proposed method for multiplicative noise removal significantly outperforms existing methods, such as the benchmark SAR-BM3D method, in terms of the visual quality of the denoised images, and of the peak-signal-to-noise ratio (PSNR) and the structural similarity index measure (SSIM) values.
Jian Lu 0002, Lixin Shen, Chen Xu 0004, Yuesheng Xu
SIAM J. Imaging Sci.2
2018 Adaptive multiple-elites-guided composite differential evolution algorithm with a shift mechanism
Laizhong Cui, Genghui Li, Zexuan Zhu 0001, Qiuzhen Lin, Ka-Chun Wong, Jianyong Chen, Jian Lu 0002
Inf. Sci.8
2018 A novel differential evolution algorithm with a self-adaptation parameter control method by differential evolution
Laizhong Cui, Genghui Li, Zexuan Zhu 0001, Zhenkun Wen, Jian Lu 0002
Soft Comput.6
2018 Modified Gbest-guided artificial bee colony algorithm with new probability model
Laizhong Cui, Kai Zhang 0049, Genghui Li, Xianghua Fu, Zhenkun Wen, Jian Lu 0002
Soft Comput.7
2017 A ranking-based adaptive artificial bee colony algorithm for global numerical optimization
Laizhong Cui, Genghui Li, Xizhao Wang, Qiuzhen Lin, Jianyong Chen, Jian Lu 0002
Inf. Sci.7
2013 Huber Fractal Image Coding Based on a Fitting Plane
abstract
Recently, there has been significant interest in robust fractal image coding for the purpose of robustness against outliers. However, the known robust fractal coding methods (HFIC and LAD-FIC, etc.) are not optimal, since, besides the high computational cost, they use the corrupted domain block as the independent variable in the robust regression model, which may adversely affect the robust estimator to calculate the fractal parameters (depending on the noise level). This paper presents a Huber fitting plane-based fractal image coding (HFPFIC) method. This method builds Huber fitting planes (HFPs) for the domain and range blocks, respectively, ensuring the use of an uncorrupted independent variable in the robust model. On this basis, a new matching error function is introduced to robustly evaluate the best scaling factor. Meanwhile, a median absolute deviation (MAD) about the median decomposition criterion is proposed to achieve fast adaptive quadtree partitioning for the image corrupted by salt & pepper noise. In order to reduce computational cost, the no-search method is applied to speedup the encoding process. Experimental results show that the proposed HFPFIC can yield superior performance over conventional robust fractal image coding methods in encoding speed and the quality of the restored image. Furthermore, the no-search method can significantly reduce encoding time and achieve less than 2.0 s for the HFPFIC with acceptable image quality degradation. In addition, we show that, combined with the MAD decomposition scheme, the HFP technique used as a robust method can further reduce the encoding time while maintaining image quality.
Jian Lu 0002, Zhongxing Ye, Yuru Zou
IEEE Trans. Image Process.1
2007 Automatic generation of colorful patterns with wallpaper symmetries from dynamics
Jian Lu 0002, Zhongxing Ye, Yuru Zou
Vis. Comput.1
2006 Orbit trap rendering method for generating artistic images with cyclic or dihedral symmetry
Yuru Zou, Wenxia Li, Jian Lu 0002, Ruisong Ye
Comput. Graph.3
2005 Orbit trap rendering methods for generating artistic images with crystallographic symmetries
Jian Lu 0002, Zhongxing Ye, Yuru Zou, Ruisong Ye
Comput. Graph.1