EDBT 2026 Demo / reviewers in the wild / expert
Yao Lu 0008
dblp:26/5662-8
· DBLP profile ↗
48ranked-venue papers
14as first author
35since 2021 · last 2025
0000-0002-3147-2081ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 11 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ALRMR-GEC: Adjusting Learning Rate Based on Memory Rate to Optimize the Edit Scorer for Grammatical Error CorrectionabstractEdit-based approaches for Grammatical Error Correction (GEC) have attracted volume attention due to their outstanding explanations of the correction process and rapid inference. Through exploring the characteristics of the generalized and specific knowledge learning for GEC, we discover that efficiently training GEC systems with satisfactory generalization capacity prefers more generalized knowledge rather than specific knowledge. Current gradient-based methods for training GEC systems, however, usually prioritize minimizing training loss over generalization loss. This paper proposes the strategy of Adjusting Learning Rate Based on Mermory Rate to optimize the edit-based GEC scorer (ALRMR-GEC). Specifically, we introduce the memory rate, a novel metric, to provide an explicit indicator for the model’s state of learning generalized and specific knowledge, which can effectively guide the GEC system to adjust the learning rate timely. Extensive experiments, conducted by optimizing the published edit scorer on the BEA2019 dataset, have shown our ALRMR-GEC significantly enhances the model generalization ability with stable and satisfactory performance nearly irrespective of the initial learning rate selection. Also, our method can accelerate the training over tenfold faster in certain cases. Finally, the experiments indicate the memory rate introduced in our ALRMR-GEC guides the GEC editscorer to learn more generalized knowledge. Zhixiao Wu, Yao Lu 0008, Jie Wen 0001, Guangming Lu 0002 |
AAAI | 2 |
| 2025 | Among General Spine Segmentation with Multi-scale and DiscriminateFeature Fusion
Tingwei Wen, Yao Lu 0008, Xiaosheng Chen, Xinhai Lu, Guangming Lu 0002 |
CVM (1) | 2 |
| 2025 | SSHR: More Secure Generative Steganography with High-Quality Revealed Secret ImagesabstractImage steganography ensures secure information transmission and storage by concealing secret messages within images. Recently, the diffusion model has been incorporated into the generative image steganography task, with text prompts being employed to guide the entire process. However, existing methods are plagued by three problems: (1) the restricted control exerted by text prompts causes generated stego images resemble the secret images and seem unnatural, raising the severe detection risk; (2) inconsistent intermediate states between Denoising Diffusion Implicit Models and its inversion, coupled with limited control of text prompts degrade the revealed secret images; (3) the descriptive text of images(i.e. text prompts) are also deployed as the keys, but this incurs significant security risks for both the keys and the secret images.To tackle these drawbacks, we systematically propose the SSHR, which joints the Reference Images with the adaptive keys to govern the entire process, enhancing the naturalness and imperceptibility of stego images. Additionally, we methodically construct an Exact Reveal Process to improve the quality of the revealed secret images. Furthermore, adaptive Reference-Secret Image Related Symmetric Keys are generated to enhance the security of both the keys and the concealed secret images. Various experiments indicate that our model outperforms existing methods in terms of recovery quality and secret image security. Jiannian Wang, Yao Lu 0008, Guangming Lu 0002 |
ICML | 2 |
| 2025 | Efficient and Separate Authentication Image Steganography NetworkabstractImage steganography hides multiple images for multiple recipients into a single cover image. All secret images are usually revealed without authentication, which reduces security among multiple recipients. It is elegant to design an authentication mechanism for isolated reception. We explore such mechanism through sufficient experiments, and uncover that additional authentication information will affect the distribution of hidden information and occupy more hiding space of the cover image. This severely decreases effectiveness and efficiency in large-capacity hiding. To overcome such a challenge, we first prove the authentication feasibility within image steganography. Then, this paper proposes an image steganography network collaborating with separate authentication and efficient scheme. Specifically, multiple pairs of lock-key are generated during hiding and revealing. Unlike traditional methods, our method has two stages to make appropriate distribution adaptation between locks and secret images, simultaneously extracting more reasonable primary information from secret images, which can release hiding space of the cover image to some extent. Furthermore, due to separate authentication, fused information can be hidden in parallel with a single network rather than traditional serial hiding with multiple networks, which can largely decrease the model size. Extensive experiments demonstrate that the proposed method achieves more secure, effective, and efficient image steganography. Code is available at https://github.com/Revive624/Authentication-Image-Steganography. Junchao Zhou, Yao Lu 0008, Jie Wen 0001, Guangming Lu 0002 |
ICML | 2 |
| 2025 | A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and TriggersabstractPoison-only Clean-label Backdoor Attacks (PCBAs) aim to covertly inject attacker-desired behavior into DNNs by merely poisoning the dataset without changing the labels. To effectively implant a backdoor, multiple triggers are proposed for various attack requirements of Attack Success Rate (ASR) and stealthiness. Additionally, sample selection enhances clean-label backdoor attacks' ASR by meticulously selecting "hard'' samples instead of random samples to poison. Current methods, however, 1) usually handle the sample selection and triggers in isolation, leading to severely limited improvements on both ASR and stealthiness. Consequently, attacks exhibit unsatisfactory performance on evaluation metrics when converted to PCBAs via a mere stacking of methods. Therefore, we seek to explore the bi-directional collaborative relations between the sample selection and triggers to address the above dilemma. 2) Since the strong specificity within triggers, the simple combination of sample selection and triggers fails to substantially enhance both evaluation metrics, with generalization preserved among various attacks. Therefore, we seek to propose a set of components to significantly improve both stealthiness and ASR based on the commonalities of attacks. Specifically, Component A ascertains two critical selection factors, and then makes them an appropriate combination based on the trigger scale to select more reasonable "hard'' samples for improving ASR. Component B is proposed to select samples with similarities to relevant trigger implanted samples to promote stealthiness. Component C reassigns trigger poisoning intensity on RGB colors through distinct sensitivity of the human visual system to RGB for higher ASR, with stealthiness ensured by sample selection including Component B. Furthermore, all components can be strategically integrated into diverse PCBAs, enabling tailored solutions that balance ASR and stealthiness enhancement for specific attack requirements. Extensive experiments demonstrate the superiority of our components in stealthiness, ASR, and generalization. Our code will be released as soon as possible. Zhixiao Wu, Yao Lu 0008, Jie Wen 0001, Guangming Lu 0002 |
NeurIPS | 2 |
| 2025 | Efficient U-shape invertible neural network for large-capacity image steganography
Le Zhang 0016, Yao Lu 0008, Yuanrong Xu, Guangming Lu 0002 |
J. Inf. Secur. Appl. | 3 |
| 2025 | Individualized image steganography method with Dynamic Separable Key and Adaptive Redundancy Anchor
Junchao Zhou, Yao Lu 0008, Guangming Lu 0002 |
Knowl. Based Syst. | 2 |
| 2025 | ACTN: Adaptive Coupling Transformer Network for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) and Transformer networks have shown impressive performance in hyperspectral image (HSI) classification. However, these models usually concentrate on examining either local or global representations of HSI data, frequently falling short of capturing multidimensional representations. Furthermore, these methods fail to fully leverage the strengths of CNNs and Transformers. This article presents the adaptive coupling Transformer network (ACTN), a parallel-hybrid network aiming to improve representation learning for HSI classification. ACTN can capture different types of representation and facilitate mutual learning. Specifically, we introduce a parallel-hybrid module called the adaptive coupling module (ACM), which is designed to capture multifaceted representations from the HSI cube. The ACM consists of two branches: a CNN branch that extracts local contextual representations and a Transformer branch that captures global dependency representations. Our proposal is an adaptive response fusion module (ARFM) that interacts with the hybrid module to merge local and global representations at different resolutions in an adaptive way. In addition, we utilize a cosine similarity function to restrict the loss function in mutual learning, guaranteeing the preservation of both local and global representations to the maximum extent. Extensive experiments conducted on three public HSI datasets demonstrate that ACTN outperforms state-of-the-art methods based on Transformers and CNNs. Xiaofei Yang 0002, Weijia Cao, Yicong Zhou, Yao Lu 0008 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Efficient U-Shape Invertible Neural Network for Image SteganographyabstractCurrently, it is challenging to recover high-quality secret images from highly secure stego images while maintaining a limited computational cost for image steganography. This paper proposes an Efficient U-shape Invertible Neural Network (EUIN-Net) for image steganography. Due to the gradual fusion and separation properties of the U-shape invertible mechanism, our EUIN-Net comprehensively couples and decouples the secret-cover information on different scales and depths. Besides, using the skip connections between each pair of U-shape invertible blocks, the long-range dependency can be retrieved. Such above two factors can drive our EUIN-Net to promote the quality of both stego and revealed secret images. Furthermore, the shared and multi-scale characteristics of the U-shaped invertible blocks during the hiding and revealing stages contribute to significant reductions of our EUIN-Net in the model size and Flops. Extensive experiments demonstrate that the proposed EUIN-Net is efficient and can achieve state-of-the-art performances for image steganography. Le Zhang 0016, Yao Lu 0008, Mi-Xiao Hou, Guangming Lu 0002 |
ICME | 3 |
| 2024 | QFormer: An Efficient Quaternion Transformer for Image Denoising
Bo Jiang 0017, Yao Lu 0008, Guangming Lu 0002, Bob Zhang 0001 |
IJCAI | 2 |
| 2024 | Implicit Prompt Learning for Image Denoising
Yao Lu 0008, Bo Jiang 0017, Guangming Lu 0002, Bob Zhang 0001 |
IJCAI | 1 |
| 2024 | Frequency Adapter and Spatial Prompt Network for All-in-One Blind Image Restoration
Shuoming Chen, Wenjie Pei, Yao Lu 0008, Guangming Lu 0002 |
PRCV (8) | 3 |
| 2024 | AGP-Net: Adaptive Graph Prior Network for Image DenoisingabstractImage denoising is a critical problem in industrial information applications since noisy images can have adverse effects on the performance of many industrial tasks. Currently, Transformer structures and graph convolutional networks (GCNs) have been widely employed in image denoising to capture long-range dependencies for the performance promotion. These methods, however, severely suffer from three major problems. Initially, the long-range dependencies captured by Transformers and GCNs are only focused on the pixel level and patch level, respectively. This leads to the coarse retrieved feature, hindering further performance promotion. In addition, due to the limited training data, especially for the noisy images with highly diverse and complex noise, the denoising process may lack sufficient feature for reconstructing denoised images. Eventually, the limited training data may also results in over-fitting, leading to poor generalization in the denoising process. This article first proposes adaptive graph prior network (AGP-Net) using a novel graph construction method to capture the long-range dependencies on both the pixel and patch levels. Then, we propose graph supplementary prior and graph noise prior in AGP-Net to adaptively generate supplementary feature and regularization noise for improving the performance and generalization of image denoising. Extensive ablation and benchmark tests show our AGP-Net achieve the most advanced image denoising performance. Bo Jiang 0017, Yao Lu 0008, Bob Zhang 0001, Guangming Lu 0002 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Joint Adaptive Robust Steganography NetworkabstractTransmission distortions within steganography systems easily cause dramatic degradations of revealing and invisibility performances. Previous works lacked sufficient adaptation for different distortions, which hinders the performance improvement of robust image steganography. This article proposes joint adaptive robust steganography network (JARS-Net). Specifically, the hierarchical attentive invertible (HAI) mechanism is first proposed to achieve adaptive feature tuning by gradually adjusting and fusing the cover-secret information from different depths and scales. Moreover, adaptive key learning (AKL) is proposed as an adaptive steganography strategy to generate adaptive keys for secret recovery under different distortions. Furthermore, benefiting from the joint of reversible HAI and the soft AKL, revealed secret images can be progressively decoupled from the received stego images along the backward HAI flow. Extensive experiments demonstrate that the proposed JARS-Net can significantly promote the invisibility and revealing performances of covert communication under different distortions. Le Zhang 0016, Yao Lu 0008, Guangming Lu 0002 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Efficient Harmonic Neural Networks With Compound Discrete Cosine Transform Filters and Shared Reconstruction FiltersabstractThe harmonic neural network (HNN) learns a combination of discrete cosine transform (DCT) filters to obtain an integrated feature from all spectra in the frequency domain. HNN, however, faces two challenges in learning and inference processes. First, the spectrum feature learned by HNN is insufficient and limited because the number of DCT filters is much smaller than that of feature maps. In addition, the number of parameters and the computation costs of HNN are significantly high because the intermediate spectrum layers are expanded multiple times. These two challenges will severely harm the performance and efficiency of HNN. To solve these problems, we first propose the compound DCT (C-DCT) filters integrating the nearest DCT filters to retrieve rich spectrum features to improve the performance. To significantly reduce the model size and computation complexity for improving the efficiency, the shared reconstruction filter is then proposed to share and dynamically drop the meta-filters in every frequency branch. Integrating the C-DCT filters with the shared reconstruction filters, the efficient harmonic network (EH-Net) is introduced. Extensive experiments on different datasets demonstrate that the proposed EH-Nets can effectively reduce the model size and computation complexity while maintaining the model performance. The code has been released at https://github.com/zhangle408/EH-Nets. Yao Lu 0008, Le Zhang 0016, Xiaofei Yang 0002, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Lightweight image denoising network with four-channel interaction transform
Yao Lu 0008, Guangming Lu 0002 |
Image Vis. Comput. | 2 |
| 2023 | Deep adaptive hiding network for image hiding using attentive frequency extraction and gradual depth extraction
Le Zhang 0016, Yao Lu 0008, Jinxing Li 0003, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001 |
Neural Comput. Appl. | 2 |
| 2023 | Joint adjustment image steganography networks
Le Zhang 0016, Yao Lu 0008, Guangming Lu 0002 |
Signal Process. Image Commun. | 2 |
| 2023 | Few-Shot Learning for Image DenoisingabstractDeep Neural Networks (DNNs) have achieved impressive results on the task of image denoising, but there are two serious problems. First, the denoising ability of DNNs-based image denoising models using traditional training strategies heavily relies on extensive training on clean-noise image pairs. Second, image denoising models based on DNNs usually have large parameters and high computational complexity. To address these issues, this paper proposes a two-stage Few-Shot Learning for Image Denoising (FSLID). Our FSLID is a two-stage denoising strategy integrating Basic Feature Learner (BFL), Denoising Feature Inducer (DFI), and Shared Image Reconstructor (SIR). BFL and SIR are first jointly unsupervised to train on the base image dataset$\mathcal {D}_{base}$consisting of easily collected high-quality clean images. Following this, the trained BFL extracts the guided features and constraint features for the noisy and corresponding clean images in the novel image dataset$\mathcal {D}_{novel}$, respectively. Furthermore, DFI encodes the noisy features of the noisy images in$\mathcal {D}_{novel}$. Then, inducing both the guided features and noisy features, DFI can generate the denoising prior features for the SIR with frozen weights to adaptively denoise the noisy images. Furthermore, we propose refined, low-channel-count, recursive multi-branch Multi-Scale Feature Recursive (MSFR) to modularly formulate an efficient DFI to capture more diverse contextual features information under a limited number of feature channels. Thus, compared with the baseline models, the FSLID composed of the proposed MSFR can significantly reduce the number of model parameters and computational complexity. Extensive experimental results demonstrate our FSLID significantly outperforms well-established baselines on multiple datasets and settings. We hope that our work will encourage further research to explore the field of few-shot image denoising. Bo Jiang 0017, Yao Lu 0008, Bob Zhang 0001, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | QTN: Quaternion Transformer Network for Hyperspectral Image ClassificationabstractNumerous state-of-the-art transformer-based techniques with self-attention mechanisms have recently been demonstrated to be quite effective in the classification of hyperspectral images (HSIs). However, traditional transformer-based methods severely suffer from the following problems when processing HSIs with three dimensions: (1) processing the HSIs using 1D sequences misses the 3D structure information; (2) too expensive numerous parameters for hyperspectral image classification tasks; (3) only capturing spatial information while lacking the spectral information. To solve these problems, we propose a novel Quaternion Transformer Network (QTN) for recovering self-adaptive and long-range correlations in HSIs. Specially, we first develop a band adaptive selection module (BASM) for producing Quaternion data from HSIs. And then, we propose a new and novel quaternion self-attention (QSA) mechanism to capture the local and global representations. Finally, we propose a new and novel transformer method, i.e., QTN by stacking a series of QSA for hyperspectral classification. The proposed QTN could exploit computation using Quaternion algebra in hypercomplex spaces. Extensive experiments on three public datasets demonstrate that the QTN outperforms the state-of-the-art vision transformers and convolution neural networks. Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Graph Attention in Attention Network for Image DenoisingabstractImage denoising aims to remove the noise from noisy images. With the increasing complexity of the noise within the noisy images, current denoising methods cannot satisfactorily address this issue. This article proposes a graph attention in attention network (GAiA-Net) for image denoising. First, we introduce a novel approach to graph construction for the GAiA-Net. In the process of such graph construction, the noisy images are divided into patches to formulate the nodes in a graph. The edges are initialized using$k $-nearest neighbors. Hence, through iterative transformation and learning, both the pixel-level and structure-level features can be captured by different information exchanges and aggregation within (pixel-level) and outside (structure-level) of the nodes, respectively. Second, we propose the graph attention in attention (GAiA) in the GAiA-Net. The proposed GAiA produces the pixel-level attention within nodes to be further induced to the nodes with various distances to generate the final attention. Therefore, our GAiA-Net can capture the long dependencies on both the pixel-level and structure-level features, which can effectively reduce the complex noise in the denoising process. Comprehensive experiments demonstrate that the proposed GAiA-Net produces state-of-the-art performances on both synthetic noise image and real noise image datasets. Especially, when experimenting on complex noisy Nam datasets, our GAiA-Net achieves a PSNR of 40.40 dB and SSIM of 0.989. These results prove the satisfactory potential and effectiveness of our GAiA-Net. Bo Jiang 0017, Yao Lu 0008, Xiaosheng Chen, Xinhai Lu, Guangming Lu 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | AIA: Attention in Attention Within Collaborate Domains
Le Zhang 0016, Yao Lu 0008, Guangming Lu 0002 |
PRCV (1) | 3 |
| 2022 | Real noise image adjustment networks for saliency-aware stylistic color retouch
Bo Jiang 0017, Yao Lu 0008, Guangming Lu 0002, David Zhang 0001 |
Knowl. Based Syst. | 2 |
| 2022 | Recursive Feature Diversity Network for audio super-resolution
Bo Jiang 0017, Mi-Xiao Hou, Yao Lu 0008, David Zhang 0001, Guangming Lu 0002 |
Speech Commun. | 4 |
| 2022 | Deep Image Denoising With Adaptive PriorsabstractImage denoising methods using deep neural networks have achieved a great progress in the image restoration. However, the recovered images restored by these deep denoising methods usually suffer from severe over-smoothness, artifacts, and detail loss. To improve the quality of restored images, we first propose Supplemental Priors (SP) method to adaptively predict depth-directed and sample-directed prior information for the reconstruction (decoder) networks. Furthermore, the over-parameterized deep neural networks and too precise supplemental prior information may cause an over-fitting, restricting the performance promotion. To improve the generalization of denoising networks, we further propose Regularization Priors (RP) method to flexibly learn depth-directed and dataset-directed regularization noise for the retrieving (encoder) networks. By respectively integrating the encoder and decoder with these plug-and-play RP block and SP block, we propose the final Adaptive Prior Denoising Networks, called APD-Nets. APD-Nets is the first attempt to simultaneously regularize and supplement denoising networks from the adaptive priors’ view with drawing learning-based mechanism into producing adaptive regularization noise and supplemental information. Extensive experiment results demonstrate our method significantly improves the generalization of denoising networks and the quality of restored images with greatly outperforming the traditional deep denoising methods both quantitatively and visually.The code will be released athttps://github.com/JiangBoCS/APD-Nets. Bo Jiang 0017, Yao Lu 0008, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Learning Informative and Discriminative Features for Facial Expression Recognition in the WildabstractThe informativeness and discriminativeness of features collaboratively ensure high-accuracy Facial Expression Recognition (FER) in the wild. Most of existing methods use the single-path deep convolutional neural network with softmax loss for basic FER, while they cannot deal with the challenging situations of the compound FER in the wild, because they fail to learn informative and discriminative features in a targeted manner. To this end, we present an Informative and Discriminative Feature Learning (IDFL) framework that consists of two key components: the Multi-Path Attention Convolutional Neural Network (MPACNN) and Balanced Separate loss (BS loss), for both basic and compound high-accuracy FER in the wild. Specifically, MPACNN leverages different paths to learn diverse features. These features are then adaptively fused into informative ones via an attention module, such that the model can adequately capture detailed information for both basic and compound FER. The BS loss maximizes the inter-class distance of features and minimizes the intra-class one. In this way, the features are discriminative enough for high-accuracy FER in the wild. Particularly, the BS loss is invoked as the objective function of MPACNN, so the model can learn informative and discriminative features at the same time, yielding better performance. Seven databases are utilized to evaluate the proposed method, and the results demonstrate that our method achieves state-of-the-art performance on both basic and compound expressions with good generalization ability. Moreover, our model contains fewer parameters and can be trained faster than other related models. Yingjian Li 0001, Yao Lu 0008, Bingzhi Chen, Zheng Zhang 0006, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Multiscale Conditional Regularization for Convolutional Neural NetworksabstractWith the increased model size of convolutional neural networks (CNNs), overfitting has become the main bottleneck to further improve the performance of networks. Currently, the weighting regularization methods have been proposed to address the overfitting problem and they perform satisfactorily. Since these regularization methods cannot be used in all the networks and they are usually not flexible enough in different phases of the training and test processes, this article proposes a multiscale conditional (MSC) regularization method. MSC divides the intermediate features into different scales and then generates new data for each scale features, respectively. In addition, the new data are generated by employing the information from two conditions: 1) each sample feature and 2) each layer pattern. Finally, a self-identity structure is proposed to supplement the features with the generated data. Therefore, MSC can adaptively and efficiently generate much finer and individualized data to make the entire regularization more flexible. Furthermore, MSC is more general and can be applied to all kinds of networks through the proposed self-identity structure. The experimental results on all the benchmark datasets showed that the proposed MSC regularization method achieves the best performances in all the networks. Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Yuanrong Xu, Zheng Zhang 0006, David Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Addi-Reg: A Better Generalization-Optimization Tradeoff Regularization Method for Convolutional Neural NetworksabstractIn convolutional neural networks (CNNs), generating noise for the intermediate feature is a hot research topic in improving generalization. The existing methods usually regularize the CNNs by producing multiplicative noise (regularization weights), called multiplicative regularization (Multi-Reg). However, Multi-Reg methods usually focus on improving generalization but fail to jointly consider optimization, leading to unstable learning with slow convergence. Moreover, Multi-Reg methods are not flexible enough since the regularization weights are generated from a definite manual-design distribution. Besides, most popular methods are not universal enough, because these methods are only designed for the residual networks. In this article, we, for the first time, experimentally and theoretically explore the nature of generating noise in the intermediate features for popular CNNs. We demonstrate that injecting noise in the feature space can be transformed to generating noise in the input space, and these methods regularize the networks in a Mini-batch in Mini-batch (MiM) sampling manner. Based on these observations, this article further discovers that generating multiplicative noise can easily degenerate the optimization due to its high dependence on the intermediate feature. Based on these studies, we propose a novel additional regularization (Addi-Reg) method, which can adaptively produce additional noise with low dependence on intermediate feature in CNNs by employing a series of mechanisms. Particularly, these well-designed mechanisms can stabilize the learning process in training, and our Addi-Reg method can pertinently learn the noise distributions for every layer in CNNs. Extensive experiments demonstrate that the proposed Addi-Reg method is more flexible and universal, and meanwhile achieves better generalization performance with faster convergence against the state-of-the-art Multi-Reg methods. Yao Lu 0008, Zheng Zhang 0006, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Hyperspectral Image Transformer Classification NetworksabstractHyperspectral image (HSI) classification is an important task in earth observation missions. Convolution neural networks (CNNs) with the powerful ability of feature extraction have shown prominence in HSI classification tasks. However, existing CNN-based approaches cannot sufficiently mine the sequence attributes of spectral features, hindering the further performance promotion of HSI classification. This article presents a hyperspectral image transformer (HiT) classification network by embedding convolution operations into the transformer structure to capture the subtle spectral discrepancies and convey the local spatial context information. HiT consists of two key modules, i.e., spectral-adaptive 3-D convolution projection module and convolution permutator (ConV-Permutator) to retrieve the subtle spatial–spectral discrepancies. The spectral-adaptive 3-D convolution projection module produces the local spatial–spectral information from HSIs using two spectral-adaptive 3-D convolution layers instead of the linear projection layer. In addition, the Conv-Permutator module utilizes the depthwise convolution operations to separately encode the spatial–spectral representations along the height, width, and spectral dimensions, respectively. Extensive experiments on four benchmark HSI datasets, including Indian Pines, Pavia University, Houston2013, and Xiongan (XA) datasets, show the superiority of the proposed HiT over existing transformers and the state-of-the-art CNN-based methods. Our codes of this work are available athttps://github.com/xiachangxue/DeepHyperXfor the sake of reproducibility. Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Self-Supervised Learning With Prediction of Image Scale and Spectral Order for Hyperspectral Image ClassificationabstractIn recent years, Convolutional Neural Networks (CNNs) have achieved great success in hyperspectral image classification attributed to their unparalleled capacity to extract the local information. However, to successfully learn the high-level semantic image features, they always require massive amounts of manually labeled data during the training process, which is expensive, scarce, and impractical, and severely hinders the improvement of supervised deep learning methods. To alleviate these burdens, we present Self-Supervised Learning methods for hyperspectral image classification by a pre-training model using extensive unlabeled data and fine-tuning the hyperspectral image target classification. In this paper, we propose a new method for learning image characteristics by training a CNN to recognize the image scale that is applied to the hyperspectral images (HSIs). In addition, we propose a multi-pretext task method to learn stable and good feature representations combing two different pretext task methods and contrastive loss function. We evaluate the proposed methods in Self-Supervised Learning benchmarks on four benchmark HSIs datasets. The experiment results demonstrate that the proposed methods outperform the traditional supervised deep learning methods when large amounts of unlabeled HSIs data are used. Moreover, it demonstrates that the Self-Supervised Learning method is promising to alleviate dependence on manually labeled data of hyperspectral image classification. Finally, our research contributes to the creation and refinement of Self-Supervised Learning methods for pretextual tasks within the HSIs community. Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | High Resolution Fingerprint Retrieval Based on Pore Indexing and Graph ComparisonabstractFingerprint retrieval aims to identify a query fingerprint image in a large database using indexing algorithms. Because of the abundant level 3 pore features within high-resolution fingerprint images, pore-based fingerprint retrieval algorithms have been rapidly developed. These retrieval algorithms, however, suffer from severe calculation-consuming problems with the pores increasing. This paper proposes a pore-based fingerprint retrieval method for high-resolution fingerprint images. The proposed method consists of two main steps. 1) In the pore indexing step, an indexing space is constructed using the binary codes of pores in enrolled images. Then, a designed graph-based searching algorithm searches the nearest neighbors of pores from the query image to construct one-to-many correspondences. 2) In the refinement step, the one-to-many correspondences are refined by a random walker-based graph comparison algorithm to remove the false correspondences. The remained nearest neighbors are used to calculate the similarities between the query image and the enrolled images. The proposed method is evaluated on two databases, showing that our method achieves better retrieval accuracies with a higher speed than the existing pore-based retrieval algorithms. Yuanrong Xu, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Semantic-Interactive Graph Convolutional Network for Multilabel Image RecognitionabstractMultilabel image recognition, a critically practical task in computer vision, aims to predict multiple objects present in each image. The existing studies mainly focus on conceptual visual cues but fail to reconcile the visual information with their semantic guidance. Intuitively, humans can not only associate extra topological concepts but also imagine other approximate scenes based on a semantic description. Inspired by such semantic-interactive capability, two different types of semantic priors, i.e., the concept correlations of the same scene and semantic similarities among different scenes, should be further explored for the recognition decisions. To efficiently interact with these semantic relationships, in this article, we propose a novel semantic-interactive graph convolutional network (SI-GCN), which can leverage the topological information learned from knowledge graphs to boost the performance of multilabel recognition. Specifically, the proposed SI-GCN framework consists of two different GCN-based branches in parallel, i.e., concept correlations learning (CCL) branch and semantic similarity learning (SSL) branch. Inputting the semantic-embedding vectors of all the concepts, the CCL branch maps the label co-occurrence graph into a set of interdependent concept classifiers. Recalibrating the image feature embedding with the standardized supervision of the semantic similarity graph, the SSL branch learns the semantically consistent in-batch visual representations. Finally, a well-established interactive learning scheme is formulated to concurrently optimize the obtained concept classifiers and the visual representation learning in an end-to-end manner. Extensive experiments on the MS-COCO and Pascal VOC 2007 & 2012 benchmarks demonstrate the superiorities of the proposed SI-GCN method compared to the state-of-the-art baselines. Bingzhi Chen, Zheng Zhang 0006, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Highly shared Convolutional Neural Networks
Yao Lu 0008, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001 |
Expert Syst. Appl. | 1 |
| 2021 | Fully shared convolutional neural networks
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Zheng Zhang 0006, Yuanrong Xu |
Neural Comput. Appl. | 1 |
| 2021 | Fast Pore Comparison for High Resolution Fingerprint Images Based on Multiple Co-Occurrence Descriptors and Local Topology SimilaritiesabstractPore-based fingerprint recognition has been researched for decades. Many algorithms have been proposed to improve the recognition accuracy of the system. However, the accuracies are always improved at the cost of speed. This article proposes a novel method to compare the pores in high-resolution fingerprint images using the popular coarse-to-fine strategy. A multiple spatial pairwise local co-occurrence descriptor is proposed to improve the calculation of the similarities between pores. It calculates multiple local co-occurrence statistics for each pore using its neighbors. The proposed method can establish correspondences between pores more accurately. The refinement of the correspondences is then achieved by using a local topology-preserving matching algorithm. The algorithm uses rotational invariant local structures and pore pair local topology similarities to calculate the cost of each correspondence. It can remove the mismatches more accurately and efficiently. The experimental results on two high-resolution fingerprint image databases show that the proposed algorithm perform well in both accuracy and speed comparing to the existing algorithms. Yuanrong Xu, Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | High-parameter-efficiency convolutional neural networks
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001 |
Neural Comput. Appl. | 1 |
| 2020 | SRGC-Nets: Sparse Repeated Group Convolutional Neural NetworksabstractGroup convolution is widely used in many mobile networks to remove the filter's redundancy from the channel extent. In order to further reduce the redundancy of group convolution, this article proposes a novel repeated group convolutional (RGC) kernel, which has M primary groups, and each primary group includes N tiny groups. In every primary group, the same convolutional kernel is repeated in all the tiny groups. The RGC filter is the first kernel to remove the redundancy from group extent. Based on RGC, a sparse RGC (SRGC) kernel is also introduced in this article, and its corresponding network is called SRGC neural networks (SRGC-Net). The SRGC kernel is the summation of RGC kernel and pointwise group convolutional (PGC) kernel. The number of PGC's groups is M . Accordingly, in each primary group, besides the center locations in all channels, the values of parameters located in other N-1 tiny groups are all zero. Therefore, SRGC can significantly reduce the parameters. Moreover, it can also effectively retrieve spatial and channel-difference features by utilizing RGC and PGC to preserve the richness of produced features. Comparative experiments were performed on the benchmark classification data sets. Compared with the traditional popular networks, SRGC-Nets can perform better with timely reducing the model size and computational complexity. Furthermore, it can also achieve better performances than other latest state-of-the-art mobile networks on most of the databases and effectively decrease the test and training runtime. Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Super Sparse Convolutional Neural NetworksabstractTo construct small mobile networks without performance loss and address the over-fitting issues caused by the less abundant training datasets, this paper proposes a novel super sparse convolutional (SSC) kernel, and its corresponding network is called SSC-Net. In a SSC kernel, every spatial kernel has only one non-zero parameter and these non-zero spatial positions are all different. The SSC kernel can effectively select the pixels from the feature maps according to its non-zero positions and perform on them. Therefore, SSC can preserve the general characteristics of the geometric and the channels’ differences, resulting in preserving the quality of the retrieved features and meeting the general accuracy requirements. Furthermore, SSC can be entirely implemented by the “shift” and “group point-wise” convolutional operations without any spatial kernels (e.g., “3×3”). Therefore, SSC is the first method to remove the parameters’ redundancy from the both spatial extent and the channel extent, leading to largely decreasing the parameters and Flops as well as further reducing the img2col and col2img operations implemented by the low leveled libraries. Meanwhile, SSC-Net can improve the sparsity and overcome the over-fitting more effectively than the other mobile networks. Comparative experiments were performed on the less abundant CIFAR and low resolution ImageNet datasets. The results showed that the SSC-Nets can significantly decrease the parameters and the computational Flops without any performance losses. Additionally, it can also improve the ability of addressing the over-fitting problem on the more challenging less abundant datasets. Yao Lu 0008, Guangming Lu 0002, Bob Zhang 0001, Yuanrong Xu, Jinxing Li 0003 |
AAAI | 1 |
| 2019 | Separate Loss for Basic and Compound Facial Expression Recognition in the WildabstractIn the past few years, facial expression recognition has made great progress because of the development of convolutional neural networks. However, the features learned only using the softmax loss are not discriminative enough for highly accurate facial expression recognition in the wild, especially for the compound facial expression recognition. To enhance the discriminative power of the learned features, we propose the separate loss for both basic and compound facial expression recognition in the wild in this paper. Such loss maximizes intra-class similarity while minimizing the similarity between different classes. The qualitative and quantitative analysis shows that the features learned using such loss function are characterized by intra-class compactness and inter-class separation. Experiments are performed on two databases in the wild and the proposed method achieves state-of-the-art results on both basic and compound expressions. Furthermore, another two databases are used to perform cross database experiments to show the generalization ability of our method. Yingjian Li 0001, Yao Lu 0008, Jinxing Li 0003, Guangming Lu 0002 |
ACML | 2 |
| 2019 | Multi-label Chest X-Ray Image Classification via Label Co-occurrence Learning
Bingzhi Chen, Yao Lu 0008, Guangming Lu 0002 |
PRCV (2) | 2 |
| 2019 | Weighted Channel-Wise Decomposed Convolutional Neural Networks
Yao Lu 0008, Guangming Lu 0002, Yuanrong Xu |
Neural Process. Lett. | 1 |
| 2019 | High resolution fingerprint recognition using pore and edge descriptors
Yuanrong Xu, Guangming Lu 0002, Yao Lu 0008, David Zhang 0001 |
Pattern Recognit. Lett. | 3 |
| 2019 | Fingerprint Pore Comparison Using Local Features and Spatial RelationsabstractHigh-resolution fingerprint recognition has been a hot topic for many years. Compared with a traditional fingerprint image, a high-resolution fingerprint image can provide more features, such as pores and ridge contours. Introducing these features into fingerprint comparison and recognition can improve the recognition accuracy and reduce the risk of identification errors. This paper proposes a novel method for comparing pores on high-resolution fingerprint images. The method can be divided into two steps. In the first step, fingerprints are aligned using the pixel-category-distance-based data-driven descending algorithm. Traditionally, fingerprints are aligned based on feature points, such as minutiae and singular points. Such alignment methods are not suitable when dealing with partial fingerprints because small overlapping areas often do not contain enough features to guarantee a correct alignment. In this research, the ridges and valleys on fingerprints are used in combination with the orientation field for alignment. The proposed algorithm performs well when aligning both partial and full fingerprints. The common areas between the two images can be estimated based on the alignment result. In the second step, pores lying in the common areas are selected for comparison. To improve the comparison accuracy, pores are compared using local features and spatial relations. A graph comparison algorithm is designed in this step. The experimental results show that the proposed method is more accurate than other state-of-the-art pore comparison algorithms. Yuanrong Xu, Guangming Lu 0002, Yao Lu 0008, Feng Liu 0013, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | AAR-CNNs: Auto Adaptive Regularized Convolutional Neural NetworksabstractIn order to address the overfitting problem caused by the small or simple training datasets and the large model’s size in Convolutional Neural Networks (CNNs), a novel Auto Adaptive Regularization (AAR) method is proposed in this paper. The relevant networks can be called AAR-CNNs. AAR is the first method using the “abstraction extent” (predicted by AE net) and a tiny learnable module (SE net) to auto adaptively predict more accurate and individualized regularization information. The AAR module can be directly inserted into every stage of any popular networks and trained end to end to improve the networks’ flexibility. This method can not only regularize the network at both the forward and the backward processes in the training phase, but also regularize the network on a more refined level (channel or pixel level) depending on the abstraction extent’s form. Comparative experiments are performed on low resolution ImageNet, CIFAR and SVHN datasets. Experimental results show that the AAR-CNNs can achieve state-of-the-art performances on these datasets. Yao Lu 0008, Guangming Lu 0002, Yuanrong Xu, Bob Zhang 0001 |
IJCAI | 1 |
| 2018 | Pyramidal Combination of Separable Branches for Deep Short Connected Neural Networks
Yao Lu 0008, Guangming Lu 0002 |
PRCV (2) | 1 |
| 2017 | A convex multi-view low-rank sparse regression for feature selection and clusteringabstractMany real-world problems involve multi-view high-dimension-small-sample-size data analysis, such as multi-omics data. The combination of multi-view databases is supposed to provide a better biological significance. However, the multi-view data always contain noise and outlying entries that result in inaccurate and unreliable. It has become an urgent need how to effectively analyze these data. We proposed a novel convex multi-view low-rank sparse regression (CMLSR) algorithm to do cluster and feature selection. The model was constructed by imposing L2,1-norm and trace norm constraints on the regularization functions. It can diminish the impact of noises and outliers and produce more precise results. Clustering quality was determined by both sparse constraint and low-rank constraint. Finally, we selected characteristic genes based on the projection matrix. The method was used in TCGA multi-view genes expression data sets, annotated according to Gene Ontology (GO). In this paper, we demonstrated the effectiveness of the proposed algorithm through comparing it with the existing methods. Yao Lu 0008, Jin-Xing Liu 0001, Junliang Shang |
BIBM | 1 |
| 2016 | A p-norm singular value decomposition method for robust tumor clusteringabstractTumor clustering based on biomolecular data plays a very important role for cancer classifications discovery. To further improve the robustness, stability and accuracy of tumor clustering, we develop a novel dimension reduction method named p-norm singular value decomposition (PSVD) to seek a low-rank approximation matrix to the bimolecular data. To enhance the robustness to outliers, the Lp-norm is taken as the error function and the Schatten p-norm is used as the regularization function in our optimization model. To evaluate the performance of PSVD, Kmeans clustering method is then employed for tumor clustering based on the low-rank approximation matrix. The extensive experiments are performed on gene expression dataset and cancer genome dataset respectively. All experimental results demonstrate that the PSVD-based method outperforms many existing methods. Especially it is experimentally proved that the proposed method is efficient for processing higher dimensional data with good robustness and superior time performance. Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Mi-Xiao Hou, Yao Lu 0008 |
BIBM | 5 |
| 2016 | Characteristic gene selection via L2, 1-norm Sparse Principal Component AnalysisabstractSparse Principal Component Analysis (SPCA) is a method that can get the sparse loadings of the principal components (PCs), and it may formulate PCA as a regression-type optimization problem by using the elastic net. But the selected features are different with each PC and generally independent. A new method named SPCA has been proposed for removing these detect, which replaces the elastic net with L2,1-norm penalty. The results of the method on gene expression data are still unknown. Therefore, we will take a test to prove this point in this paper. Firstly, this method is applied to the simulated data for obtaining an optimal parameter. Secondly, the L2,1SPCA method is applied to the gene expression data, that is the head and neck squamous carcinoma data (HNSC). Thirdly, the characteristic genes are selected according the PCs. The results consist of very lower P-value and very higher hit count, which shows the method of L2,1SPCA can obtain higher recognition accuracy and higher relevancy to the genes. Finally, the experimental results demonstrate that the L2,1SPCA works well and has good performances in the gene expression data. Yao Lu 0008, Ying-Lian Gao, Jin-Xing Liu 0001, Chang-Gang Wen, Yaxuan Wang, Jiguo Yu |
BIBM | 1 |