EDBT 2026 Demo / reviewers in the wild / expert
Baodi Liu
dblp:295/5553 · also Bao-Di Liu
· DBLP profile ↗
158ranked-venue papers
16as first author
130since 2021 · last 2026
0000-0002-1408-5514ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 6 first-author · 45 since 2021Artificial intelligence and machine learning · 56 · 7 first-author · 48 since 2021Applied, interdisciplinary, general and emerging computing · 44 · 4 first-author · 40 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncertainty-aware adaptive feature completion networks for incomplete multi-view learning
Sichao Fu, Jun Wang 0085, Baodi Liu, Chaofeng Tang, Weihua Ou |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Multiplex graph prompt collaboration for open-set social event detection
Xiuqin Liang, Jiazhen Chen, Sichao Fu, Wuli Wang, Mingbin Feng, Tony S. Wirjanto, Qinmu Peng, Baodi Liu, Weihua Ou |
Expert Syst. Appl. | 8 |
| 2026 | Physics-guided transformer for build-up rate prediction
Yanfeng Geng, Weiliang Wang, Yucai Shi, Baodi Liu, Yisen Yang, Linyuan Shang |
Expert Syst. Appl. | 5 |
| 2026 | Universal Attack based Focal Enhancement for Bearing Fault Diagnosis
Puhua Jia, Xinghao Yang, Yongwei Tang, Baodi Liu, Wei Li 0032, Weifeng Liu 0001 |
Knowl. Based Syst. | 4 |
| 2026 | Context-aware feature complementary screening network for mass segmentation in whole mammograms
Qingkun Guo, Luhao Sun, Chao Li 0075, Wenzong Jiang, Weifeng Liu 0001, Baodi Liu |
Multim. Syst. | 9 |
| 2026 | Dynamic frequency-band filtering domain generalization for mammogram classification
Shenxiao Li, Yunqi Huang, Wenzong Jiang, Chao Li 0075, Weifeng Liu 0001, Xiongbin Wang, Baodi Liu |
Multim. Syst. | 8 |
| 2026 | Dual graph network for few-shot 3D point cloud classification
Xiaoming Gao, Haizhou Tan, Hengxin Feng, Baodi Liu |
Multim. Syst. | 7 |
| 2026 | Cross-difference-driven dual-stream contrast multi-view network for mammogram classification
Ruijia Tian, Chenteng Zhang, Wenzong Jiang, Chao Li 0075, Weifeng Liu 0001, Xiongbin Wang, Baodi Liu |
Multim. Syst. | 8 |
| 2026 | Diagnosis-driven hard sample generation: low-frequency attenuation supervised contrastive learning for mammogram classification
Changchao Wang, Wenzong Jiang, Chao Li 0075, Weifeng Liu 0001, Xiongbin Wang, Baodi Liu |
Multim. Syst. | 7 |
| 2026 | A pressure-conditioned generative adversarial network for efficient temperature field visualization in combustion simulations
Haoran Yu 0005, Baodi Liu, Weifeng Liu 0001 |
Multim. Syst. | 3 |
| 2026 | Channel-guided dual-pooling multi-scale spatial attention network for mass segmentation in whole mammograms
Wenzong Jiang, Weifeng Liu 0001, Xiongbin Wang, Baodi Liu |
Multim. Syst. | 6 |
| 2026 | Evolving classifiers with background suppression transformer for open-set long-tailed class-incremental remote sensing scene classification
Sichao Fu, Hongquan Xin, Wuli Wang, Peng Ren 0001, Baodi Liu, Weihua Ou, Dapeng Tao |
Neural Networks | 6 |
| 2026 | Causal-guided strength differential independence sample weighting for out-of-distribution generalization
Haoran Yu 0005, Weifeng Liu 0001, Yingjie Wang 0007, Baodi Liu, Dapeng Tao, Honglong Chen |
Pattern Recognit. | 4 |
| 2026 | Joint subgraph independence for graph out-of-distribution generalization
Weifeng Liu 0001, Baodi Liu, Dapeng Tao, Honglong Chen |
Pattern Recognit. | 4 |
| 2026 | CO3+: Improved Collaborative Consortium of Foundation Models for Open-World Few-Shot LearningabstractOpen-World Few-Shot Learning (OFSL) is a critical research domain focused on accurately identifying target samples under conditions where data is scarce and labels are unreliable. This field is highly relevant to real-world scenarios, holding significant practical implications. Currently, the field has only a few solutions, primarily relying on conventional methods such as metric learning and feature aggregation. However, these methods often struggle in more complex scenarios. Recent breakthroughs in foundation models such as CLIP and DINO have demonstrated their strong representational capabilities, even in resource-limited environments. These advancements have led to a shift from “training model from scratch” towards “exploiting the extensive capabilities and expertise of these pre-trained foundation models for OFSL”. Inspired by this shift, we introduce the Improved Collaborative Consortium of Foundation Models (CO+3), an extension of CO3, first presented in AAAI 2024. CO+3significantly improves the accuracy of OFSL by integrating the strengths of four foundational models. It includes three decoupled blocks: (1) The Label Correction Block (LC-Block) rectifies unreliable labels, (2) the Data Augmentation Block (DA-Block) enriches the available data, and (3) the Text-guided Fusion Adapter (TeFu-Adapter) merges various features and reduces the impact of noisy labels through semantic constraints. We evaluate CO+3across eleven benchmark datasets, comparing it against recent state-of-the-art methods. Our thorough evaluations demonstrate that the proposed CO+3consistently surpasses existing methods by a substantial margin, particularly in high-noise scenarios. Shuai Shao 0006, Rui Xu 0012, Bingfeng Zhang, Baodi Liu, Weifeng Liu 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Lesion Asymmetry Screening Assisted Global Awareness Multi-View Network for Mammogram ClassificationabstractMammography is a primary method for early screening, and developing deep learning-based computer-aided systems is of great significance. However, current deep learning models typically treat each image as an independent entity for diagnosis, rather than integrating images from multiple views to diagnose the patient. These methods do not fully consider and address the complex interactions between different views, resulting in poor diagnostic performance and interpretability. To address this issue, this paper proposes a novel end-to-end framework for breast cancer diagnosis: lesion asymmetry screening assisted global awareness multi-view network (LAS-GAM). More than just the most common image-level diagnostic model, LAS-GAM operates at the patient level, simulating the workflow of radiologists analyzing mammographic images. The framework processes the four views of a patient and revolves around two key modules: a global module and a lesion screening module. The global module simulates the comprehensive assessment by radiologists, integrating complementary information from the craniocaudal (CC) and mediolateral oblique (MLO) views of both breasts to generate global features that represent the patient's overall condition. The lesion screening module mimics the process of locating lesions by comparing symmetric regions in contralateral views, identifying potential lesion areas and extracting lesion-specific features using a lightweight model. By combining the global features and lesion-specific features, LAS-GAM simulates the diagnostic process, making patient-level predictions. Moreover, it is trained using only patient-level labels, significantly reducing data annotation costs. Experiments on the Digital Database for Screening Mammography (DDSM) and In-house datasets validate LAS-GAM, achieving AUCs of 0.817 and 0.894, respectively. Xinchuan Liu, Luhao Sun, Chao Li 0075, Bowen Han 0001, Wenzong Jiang, Tianhao Yuan, Weifeng Liu 0001, Zhaoyun Liu, Baodi Liu |
IEEE Trans. Medical Imaging | 10 |
| 2026 | Unbiased Semantic Decoding With Vision Foundation Models for Few-Shot SegmentationabstractFew-shot segmentation (FSS) has garnered significant attention. Many recent approaches attempt to introduce the segment anything model (SAM) to handle this task. With the strong generalization ability and rich object-specific extraction ability of the SAM model, such a solution shows great potential in FSS. However, the decoding process of SAM highly relies on accurate and explicit prompts, making previous approaches mainly focus on extracting prompts from the support set, which is insufficient to activate the generalization ability of SAM, and this design is easy to result in a biased decoding process when adapting to the unknown classes. In this work, we propose an unbiased semantic decoding (USD) strategy integrated with SAM, which extracts target information from both the support and query set simultaneously to perform consistent predictions guided by the semantics of the contrastive language-image pretraining (CLIP) model. Specifically, to enhance the unbiased semantic discrimination of SAM, we design two feature enhancement strategies that leverage the semantic alignment capability of CLIP to enrich the original SAM features, mainly including a global supplement at the image level to provide a generalize category indicate with support image and a local guidance at the pixel level to provide a useful target location with query image. Besides, to generate target-focused prompt embeddings, a learnable visual-text target prompt generator (VTPG) is proposed by interacting target text embeddings and clip visual features. Without requiring retraining of the vision foundation models, the features with semantic discrimination draw attention to the target region through the guidance of prompt with rich target information. Experiments on both the PASCAL- $5^{i}$ and COCO- $20^{i}$ show that our proposed method outperforms the existing approaches by a clear margin and achieves new state-of-the-art performances. Bingfeng Zhang, Jian Pang, Weifeng Liu 0001, Baodi Liu, Honglong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Excluding the Impossible for Open Vocabulary Semantic SegmentationabstractOpen vocabulary semantic segmentation is a hot topic in research, focusing on segmenting and recognizing a diverse array of categories in varied environments, including those previously unknown, thereby holding significant practical value. Mainstream studies utilize the CLIP model for direct semantic segmentation (denoted as “forward methods”), which often struggles to represent underrepresented categories effectively. To address this issue, this paper introduces a novel approach Excluding the ImpossibLe Semantic Segmentation Network (ELSE-Net) based on reverse thinking. By excluding improbable categories, ELSE-Net narrows the selection range for forward methods, significantly reducing the risk of misclassification. In implementation, we initially draw on leading research to design the General Processing Block (GP-Block), which generates inclusion probabilities (the likelihood of belonging to a category) by using the CLIP model cooperated with a Mask Proposal Network (MPN). We then present the EXcluding the ImPossible Block (EXP-Block), which computes exclusion probabilities (the likelihood of not belonging to a category) through the CLIPN model and a custom-designed Reverse Retrieval Adapter (R2-Adapter). These exclusion probabilities are subsequently used to refine the inclusion probabilities, which are ultimately employed to annotate class-agnostic masks. Moreover, the core component of our EXP-Block is model-agnostic, enabling it to enhance the capabilities of existing frameworks. Experimental results from four benchmark datasets validate the effectiveness of ELSE-Net and underscore the seamless model-agnostic functionality of the EXP-Block. Shiyuan Zhao, Baodi Liu, Weifeng Liu 0001, Shuai Shao 0006 |
AAAI | 2 |
| 2025 | Domain Generalization for Mammogram Classification by Suppressing Domain-Specific Features
Jiqun Chen, Luhao Sun, Wenzong Jiang, Weifeng Liu 0001, Chao Li 0075, Baodi Liu |
MICCAI (7) | 7 |
| 2025 | Noise-Robust Few-Shot Classification via Variational Adversarial Data AugmentationabstractFew-shot classification models trained with clean samples poorly classify samples from the real world with various scales of noise. To enhance the model for recognizing noisy samples, researchers usually utilize data augmentation or use noisy samples generated by adversarial training for model training. However, existing methods still have problems: (i) The effects of data augmentation on the robustness of the model are limited. (ii) The noise generated by adversarial training usually causes overfitting and reduces the generalization ability of the model, which is very significant for few-shot classification. (iii) Most existing methods cannot adaptively generate appropriate noise. Given the above three points, this paper proposes a noise-robust few-shot classification algorithm, VADA—Variational Adversarial Data Augmentation. Unlike existing methods, VADA utilizes a variational noise generator to generate an adaptive noise distribution according to different samples based on adversarial learning, and optimizes the generator by minimizing the expectation of the empirical risk. Applying VADA during training can make few-shot classification more robust against noisy data, while retaining generalization ability. In this paper, we utilize FEAT and ProtoNet as baseline models, and accuracy is verified on several common few-shot classification datasets, including MiniImageNet, TieredImageNet, and CUB. After training with VADA, the classification accuracy of the models increases for samples with various scales of noise. Baodi Liu, Kai Zhang 0029, Honglong Chen, Dapeng Tao, Weifeng Liu 0001 |
Comput. Vis. Media | 2 |
| 2025 | WMANet:Weighted multiple adaptive feature attention for self-supervised single remote-sensing image denoising
Weifeng Liu 0001, Dapeng Tao, Baodi Liu, Yanjiang Wang 0001 |
Knowl. Based Syst. | 6 |
| 2025 | PPBU: Progressive Pixel Bank Updating Strategy for Single Remote Sensing Image Denoising
Baodi Liu, Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | GSF-GAN: A Global-Aware Selective Fusion Generative Adversarial Network for Multitemporal Cloud RemovalabstractRemote sensing images are essential for surface observation, yet cloud cover often leads to a lack of spectral information, significantly disrupting the continuity of these images. Therefore, effective de-clouding techniques are crucial. Among these, multi-temporal de-clouding approaches offer notable advantages, as they utilize complementary images from different time periods and avoid the complexity of multimodal data fusion. However, maintaining image continuity and preserving fine details remains a major challenge due to the complex surface features, high detail recovery demands, and subtle differences between multi-temporal images. To address these challenges, we propose a novel de-clouding framework, the Global-aware Selective Fusion Generative Adversarial Network (GSF-GAN). GSF-GAN tackles the problem by introducing a Triplet Weight Selection Module (TWSM), which efficiently filters high-quality features to preserve more surface details, while the RescaleNorm-ReLU Swin Transformer (ReSwin) captures temporal variation patterns through global context modeling, further enhancing the de-clouding effect. Through experimental validation on STGAN and Sen2_MTC datasets, our method shows significant advantages in the de-cloud effect, and the PSNR and SSIM indexes are better than the control model, which verifies its effectiveness and superiority. Aozhe Dou, Weifeng Liu 0001, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Multi-scale region selection network in deep features for full-field mammogram classificationabstractEarly diagnosis and treatment of breast cancer can effectively reduce mortality. Since mammogram is one of the most commonly used methods in the early diagnosis of breast cancer, the classification of mammogram images is an important work of computer-aided diagnosis (CAD) systems. With the development of deep learning in CAD, deep convolutional neural networks have been shown to have the ability to complete the classification of breast cancer tumor patches with high quality, which makes most previous CNN-based full-field mammography classification methods rely on region of interest (ROI) or segmentation annotation to enable the model to locate and focus on small tumor regions. However, the dependence on ROI greatly limits the development of CAD, because obtaining a large number of reliable ROI annotations is expensive and difficult. Some full-field mammography image classification algorithms use multi-stage training or multi-feature extractors to get rid of the dependence on ROI, which increases the computational amount of the model and feature redundancy. In order to reduce the cost of model training and make full use of the feature extraction capability of CNN, we propose a deep multi-scale region selection network (MRSN) in deep features for end-to-end training to classify full-field mammography without ROI or segmentation annotation. Inspired by the idea of multi-example learning and the patch classifier, MRSN filters the feature information and saves only the feature information of the tumor region to make the performance of the full-field image classifier closer to the patch classifier. MRSN first scores different regions under different dimensions to obtain the location information of tumor regions. Then, a few high-scoring regions are selected by location information as feature representations of the entire image, allowing the model to focus on the tumor region. Experiments on two public datasets and one private dataset prove that the proposed MRSN achieves the most advanced performance. Luhao Sun, Bowen Han 0001, Wenzong Jiang, Weifeng Liu 0001, Baodi Liu, Dapeng Tao, Chao Li 0075 |
Medical Image Anal. | 5 |
| 2025 | Target data guided few-shot remote sensing scene classification in reproducing Hilbert kernel space
Chunyu Du, Baodi Liu, Yanjiang Wang 0001 |
Multim. Syst. | 2 |
| 2025 | Parentheses insertion based sentence-level text adversarial attack
Xinghao Yang, Baodi Liu, Honglong Chen, Dapeng Tao, Weifeng Liu 0001 |
Multim. Syst. | 3 |
| 2025 | Correction: Parentheses insertion based sentence-level text adversarial attack
Xinghao Yang, Baodi Liu, Honglong Chen, Dapeng Tao, Weifeng Liu 0001 |
Multim. Syst. | 3 |
| 2025 | Fourier aids CNN and transformer for semantic segmentation of remote sensing images
Jun Wang 0085, Youzhou Wu, Baodi Liu, Keding Wang |
Multim. Syst. | 3 |
| 2025 | FSDMB: few-shot object detection via double matching branch
Baodi Liu, Qingtao Xie |
Multim. Tools Appl. | 1 |
| 2025 | A spatial-spectral fusion convolutional transformer network with contextual multi-head self-attention for hyperspectral image classification
Wuli Wang, Peng Ren 0001, Jianbu Wang, Guangbo Ren, Baodi Liu |
Neural Networks | 7 |
| 2025 | Feature aggregation and connectivity for object re-identification
Dongchen Han, Baodi Liu, Shuai Shao 0006, Weifeng Liu 0001, Yicong Zhou |
Pattern Recognit. | 2 |
| 2025 | IW-ViT: Independence-Driven Weighting Vision Transformer for out-of-distribution generalization
Weifeng Liu 0001, Haoran Yu 0005, Yingjie Wang 0007, Baodi Liu, Dapeng Tao, Honglong Chen |
Pattern Recognit. | 4 |
| 2025 | MSC-GAN: A Multistream Complementary Generative Adversarial Network With Grouping Learning for Multitemporal Cloud RemovalabstractOptical remote sensing images have extensive application value, but cloud contamination greatly limits their potential use in the field of geographic information. Cloud removal aims to restore clear, unobstructed images from cloud-covered ones for subsequent in-depth analysis. Due to severe cloud cover problems such as thick clouds in some areas of remote sensing images, cloud removal tasks have become challenging. Recently, many methods have attempted to incrementally fill in obscured regions by fusing cloud-free information from multitemporal data. However, most of these methods fail to effectively utilize the interaction among different temporal data, and some information of data is easily lost in the process of deep transmission, this causes problems such as inadequate cloud removal and blurred recovery of ground under the clouds. Therefore, we propose a multistream complementary generative adversarial network (MSC-GAN) for cloud removal using multitemporal data. First, it employs a multistream complementary (MSC) architecture in the down-sampling feature encoding stage to effectively promote the interaction of feature information across multitemporal data, alleviating information loss as network depth increases. Second, to reduce the feature blur, we design a group feature reweighting (GFR) module as a complementary connection of long-distance information, in which the grouping learning and multidimensional parallel architecture can cost-effectively enhance semantic fusion between low-level and high-level features. Moreover, a channel enhancement method is introduced to assist in processing the underlying transition information, minimizing the interference of invalid information. Experimental results on multiple benchmark datasets under a series of image quality assessment metrics demonstrate the effectiveness of the proposed method. Yanjiang Wang 0001, Weifeng Liu 0001, Dapeng Tao, Baodi Liu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | See Degraded Objects: A Physics-Guided Approach for Object Detection in Adverse EnvironmentsabstractIn adverse environments, the detector often fails to detect degraded objects because they are almost invisible and their features are weakened by the environment. Common approaches involve image enhancement to support detection, but they inevitably introduce human-invisible noise that negatively impacts the detector. In this work, we propose a physics-guided approach for object detection in adverse environments, which gives a straightforward solution that injects the physical priors into the detector, enabling it to detect poorly visible objects. The physical priors, derived from the imaging mechanism and image property, include environment prior and frequency prior. The environment prior is generated from the physical model, e.g., the atmospheric model, which reflects the density of environmental noise. The frequency prior is explored based on an observation that the amplitude spectrum could highlight object regions from the background. The proposed two priors are complementary in principle. Furthermore, we present a physics-guided loss that incorporates a novel weight item, which is estimated by applying the membership function on physical priors and could capture the extent of degradation. By backpropagating the physics-guided loss, physics knowledge is injected into the detector to aid in locating degraded objects. We conduct experiments in synthetic foggy environment, real foggy environment, and real underwater scenario. The results demonstrate that our method is effective and achieves state-of-the-art performance. The code is available at https://github.com/PangJian123/See-Degraded-Objects. Weifeng Liu 0001, Jian Pang, Bingfeng Zhang, Baodi Liu, Dapeng Tao |
IEEE Trans. Image Process. | 5 |
| 2024 | Collaborative Consortium of Foundation Models for Open-World Few-Shot LearningabstractOpen-World Few-Shot Learning (OFSL) is a crucial research field dedicated to accurately identifying target samples in scenarios where data is limited and labels are unreliable. This research holds significant practical implications and is highly relevant to real-world applications. Recently, the advancements in foundation models like CLIP and DINO have showcased their robust representation capabilities even in resource-constrained settings with scarce data. This realization has brought about a transformative shift in focus, moving away from “building models from scratch” towards “effectively harnessing the potential of foundation models to extract pertinent prior knowledge suitable for OFSL and utilizing it sensibly”. Motivated by this perspective, we introduce the Collaborative Consortium of Foundation Models (CO3), which leverages CLIP, DINO, GPT-3, and DALL-E to collectively address the OFSL problem. CO3 comprises four key blocks: (1) the Label Correction Block (LC-Block) corrects unreliable labels, (2) the Data Augmentation Block (DA-Block) enhances available data, (3) the Feature Extraction Block (FE-Block) extracts multi-modal features, and (4) the Text-guided Fusion Adapter (TeFu-Adapter) integrates multiple features while mitigating the impact of noisy labels through semantic constraints. Only the adapter's parameters are adjustable, while the others remain frozen. Through collaboration among these foundation models, CO3 effectively unlocks their potential and unifies their capabilities to achieve state-of-the-art performance on multiple benchmark datasets. https://github.com/The-Shuai/CO3. Shuai Shao 0006, Yan Wang 0076, Baodi Liu, Bin Liu 0021 |
AAAI | 4 |
| 2024 | DeIL: Direct-and-Inverse CLIP for Open-World Few-Shot LearningabstractOpen-World Few-Shot Learning (OFSL) is a critical field of research, concentrating on the precise identification of target samples in environments with scarce data and unre-liable labels, thus possessing substantial practical signif-icance. Recently, the evolution of foundation models like CLIP has revealed their strong capacity for representation, even in settings with restricted resources and data. This development has led to a significant shift in focus, tran-sitioning from the traditional method of “building models from scratch” to a strategy centered on “efficiently utilizing the capabilities of foundation models to extract rele-vant prior knowledge tailored for OFSL and apply it judi-ciously”. Amidst this backdrop, we unveil the Direct-and-Inverse CLIP (DeIL), an innovative method leveraging our proposed “Direct-and-Inverse” concept to activate CLIP-based methods for addressing OFSL. This concept transforms conventional single-step classification into a nuanced two-stage process: initially filtering out less probable cate-gories, followed by accurately determining the specific cat-egory of samples. DeIL comprises two key components: a pretrainer (frozen) for data denoising, and an adapter (tun-able) for achieving precise final classification. In experiments, DeIL achieves SOTA performance on 11 datasets. https://github.com/The-Shuai/DeIL. Shuai Shao 0006, Yan Wang 0076, Baodi Liu, Yicong Zhou |
CVPR | 4 |
| 2024 | Adaptive Immune-based Sound-Shape Code Substitution for Adversarial Chinese Text AttacksabstractAdversarial textual examples reveal the vulnerability of natural language processing (NLP) models.Most existing text attack methods are designed for English text, while the robust implementation of the second popular language, i.e., Chinese with 1 billion users, is greatly underestimated.Although several Chinese attack methods have been presented, they either directly transfer from English attacks or adopt simple greedy search to optimize the attack priority, usually leading to unnatural sentences.To address these issues, we propose an adaptive Immune-based Sound-Shape Code (ISSC) algorithm for adversarial Chinese text attacks.Firstly, we leverage the Sound-Shape code to generate natural substitutions, which comprehensively integrate multiple Chinese features.Secondly, we employ adaptive immune algorithm (IA) to determine the replacement order, which can reduce the duplication of population to improve the search ability.Extensive experimental results validate the superiority of our ISSC in producing high-quality Chinese adversarial texts. Xinghao Yang, Baodi Liu, Weifeng Liu 0001 |
EMNLP | 4 |
| 2024 | Spectral Channel-Weighting CAT for Hyperspectral Image Classification
Yujuan Qi, Baodi Liu, Yanjiang Wang 0001 |
PRCV (13) | 3 |
| 2024 | Discriminative Representation-Based Classifier for Few-Shot Remote Sensing Classification
Tianhao Yuan, Weifeng Liu 0001, Yingjie Wang 0007, Baodi Liu |
PRCV (13) | 4 |
| 2024 | Multi-View Self-Supervised Auxiliary Task for Few-Shot Remote Sensing ClassificationabstractABSTRACT In the past few years, the swift advancement of remote sensing technology has greatly promoted its widespread application in the agricultural field. For example, remote sensing technology is used to monitor the planting area and growth status of crops, classify crops, and detect agricultural disasters. In these applications, the accuracy of image classification is of great significance in improving the efficiency and sustainability of agricultural production. However, many of the existing studies primarily rely on contrastive self‐supervised learning methods, which come with certain limitations such as complex data construction and a bias towards invariant features. To address these issues, additional techniques like knowledge distillation are often employed to optimize the learned features. In this article, we propose a novel approach to enhance feature acquisition specific to remote sensing images by introducing a classification‐based self‐supervised auxiliary task. This auxiliary task involves performing image transformation self‐supervised learning tasks directly on the remote sensing images, thereby improving the overall capacity for feature representation. In this work, we design a texture fading reinforcement auxiliary task to reinforce texture features and color features that are useful for distinguishing similar classes of remote sensing. Different auxiliary tasks are fused to form a multi‐view self‐supervised auxiliary task and integrated with the main task to optimize the model training in an end‐to‐end manner. The experimental results on several popular few‐shot remote sensing image datasets validate the effectiveness of the proposed method. The performance better than many advanced algorithms is achieved with a more concise structure. Baodi Liu, Xujian Qiao |
Comput. Intell. | 1 |
| 2024 | Ensembling Multi-View Discriminative Semantic Feature for Few-Shot Classification
Rui Xu 0012, Shuai Shao 0006, Lei Xing 0005, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Feedback-Irrelevant Mapping: An evaluation method for decoupled few-shot classification
Rui Xu 0012, Shuai Shao 0006, Lei Xing 0005, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Adaptive Gradient-based Word Saliency for adversarial text attacks
Yupeng Qi, Xinghao Yang, Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001 |
Neurocomputing | 3 |
| 2024 | HSIR-ME-CPMG: A High-Resolved Pulse Sequence for the T₁ - T₂ Measurement of Unconventional Reservoir RocksabstractWe designed an improved pulse sequence combined with the hybrid saturation recovery and inversion recovery and the multiple echo spaced Carr-Purcell-Meiboom-Gill (CPMG) pulse sequence (HSIR-ME-CPMG) to measure the longitudinal relaxation time (T1) and the transverse relaxation time (T2) simultaneously, to overcome the shortcomings of conventional IR-CPMG and SR-CPMG pulse train acquisition time and low image resolution. The longitudinal relaxation time (T1) is encoded using the combination of the saturation recovery and the inversion recovery pulse sequence to improve the resolution between different relaxation components. The transverse relaxation time (T2) are encoded by the CPMG pulse sequence. To reduce the energy consumption and the data storage space, the echo spacing is varied in different windows. Numerical simulations show that the proposed pulse sequence can capture the contrast between different components with similar relaxation times, even in low signal-to-noise ratio (SNR). The proposed pulse sequence can be popularized for better characterizing relaxation components of unconventional reservoirs such as the shale oil. Xinmin Ge, Quansheng Miao, Lei Xing 0006, Baodi Liu, Fu Zuo |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | MWLN: Multilevel Wavelet Learning Network for Continuous-Scale Remote-Sensing Image Super-ResolutionabstractRemote-sensing image super-resolution (SR) reconstructs high resolution (HR) with texture from the input low resolution (LR). It has been widely used and applied in image-processing tasks. However, most algorithms focus on designing more complex structures to enhance performance, ignoring learning frequency information. Moreover, existing methods are designed for SR tasks with specific scales, such as scales of 2 and 4. It limits the network performance in applications. To alleviate the above issues, this letter designs a multilevel wavelet learning network (MWLN) for continuous-scale remote-sensing image SR. MWLN achieves continuous magnification remote-sensing image SR tasks without training at different scales multiple times through multilevel wavelet feature aggregation (MWFA) and self-learning implicit representation (SLIR). MWFA extracts hierarchical features and applies discrete wavelet transforms (DWTs), capturing high-frequency information while avoiding information loss. Moreover, this letter cascades a multidimensional attention mechanism model channel and spatial features and enhances features’ interaction. SLIR maps the image coordinates and red, green, and blue (RGB) value through self-learning, realizing the continuous-scale reconstruction. Extensive experimental results demonstrate that MWLN outperforms the compared methods in quantitative and qualitative results on specific and continuous-scale remote-sensing image SR tasks. Baodi Liu, Lifei Zhao, Weifeng Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | Dual-Branch Feature Fusion Network Based Cross-Modal Enhanced CNN and Transformer for Hyperspectral and LiDAR ClassificationabstractThe joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) data has attracted considerable attention in the field of remote sensing. Integrating the advantages of the two data sources can provide precise data support and analytical decision-making for remote-sensing applications. However, due to the inherent differences in properties and semantic information from heterogeneous data, most existing deep-learning methods suboptimally extract the characteristic features of both data sources while utilizing their interactive information. In this letter, we propose a dual-branch feature fusion network-based cross-modal enhanced CNN and Transformer (DF2NCECT) to make full use of the respective features and interactive information of multisource data. DF2NCECT consists of two main stages. One is the basic feature extraction stage, which builds a hybrid convolution module based on 3DCNN and inception structure to fully extract the joint features of HSI from multiple spatial perspectives. The other is the deep feature fusion stage, where the CNN and Transformer are designed in parallel to fully explore and fuse deep features between HSI and LiDAR. More importantly, to achieve efficacious interactive information between HSI and LiDAR, a cross-modal enhanced CNN and Transformer module (CECT) is designed to deeply enhance the fused interactive features from global/local perspectives. Experiments show that the proposed method is superior and outperforms the comparison methods by an average of 3.06% in OA on Houston2013 and 1.79% on Summer, respectively. Wuli Wang, Chong Li 0006, Peng Ren 0001, Xinchao Lu, Jianbu Wang, Guangbo Ren, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2024 | SELM: Self-Motivated Ensemble Learning Model for Cross-Domain Few-Shot Classification in Hyperspectral ImagesabstractHyperspectral image (HSI) classification is a common task in remote sensing that often faces challenges due to limited samples and cross-domain discrepancies between training and test data. This particular problem is termed as HSI Cross-domain Few-Shot Classification (HSI-CFSC). To solve this problem, we propose a Self-motivated Ensemble Learning Model (SELM). Building upon source pre-trained representations, our end-to-end approach comprises a self-training paradigm to iteratively refine target representations independent of direct source supervision. Moreover, an ensemble classifier suite leveraging diverse decision boundaries is optimized to excavate comprehensive classification cues from limited labeled target data. The OA, AA and Kappa of SELM in UP, PC and Salinas data sets are respectively 86.55%, 82.27%, 82.10%, 98.07%, 94.20%, 97.30% and 91.33%, 94.96%, 90.37%, which achieve the state-of-art performance compared with other classical methods. Shiyuan Zhao, Shuai Shao 0006, Weifeng Liu 0001, Xinmin Ge, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2024 | Remote sensing image cloud removal based on multi-scale spatial information perception
Aozhe Dou, Weifeng Liu 0001, Zhenzhong Wang, Baodi Liu |
Multim. Syst. | 6 |
| 2024 | Design of integrated interactive system for pre-diagnosis of breast cancer pathological images based on CNN and PyQt5
Yunkai Yang, Qijia Yang, Weifeng Liu 0001, Baodi Liu |
Multim. Syst. | 4 |
| 2024 | Central Attention with Multi-Graphs for Image AnnotationabstractAbstract In recent decades, the development of multimedia and computer vision has sparked significant interest among researchers in the field of automatic image annotation. However, much of the research has primarily focused on using a single graph for annotating images in semi-supervised learning. Conversely, numerous approaches have explored the integration of multi-view or image segmentation techniques to create multiple graph structures. Yet, relying solely on a single graph proves to be challenging, as it struggles to capture the complete manifold of structural information. Furthermore, the computational complexity of building multiple graph structures based on multi-view or image segmentation is substantial and time-consuming. To address these issues, we propose a novel method called "Central Attention with Multi-graphs for Image Annotation." Our approach emphasizes the critical role of the central image region in the annotation process. Remarkably, we demonstrate that impressive performance can be achieved by leveraging just two graph structures, composed of central and overall features, in semi-supervised learning. To validate the effectiveness of our proposed method, we conducted a series of experiments on benchmark datasets, including Corel5K, ESPGame, and IAPRTC12. These experiments provide empirical evidence of our method’s capabilities. Baodi Liu, Qianqian Shao, Weifeng Liu 0001 |
Neural Process. Lett. | 1 |
| 2024 | Simplified Multi-head Mechanism for Few-Shot Remote Sensing Image ClassificationabstractAbstract The study of few-shot remote sensing image classification has received significant attention. Although meta-learning-based algorithms have been the primary focus of recent examination, feature fusion methods stress feature extraction and representation. Nonetheless, current feature fusion methods, like the multi-head mechanism, are restricted by their complicated network structure and challenging training process. This manuscript presents a simplified multi-head mechanism for obtaining multiple feature representations from a single sample. Furthermore, we perform specific fundamental transformations on remote-sensing images to obtain more suitable features for information representation. Specifically, we reduce multiple feature extractors of the multi-head mechanism to a single one and add an image transformation module before the feature extractor. After transforming the image, the features are extracted resulting in multiple features for each sample. The feature fusion stage is integrated with the classification prediction stage, and multiple linear classifiers are combined for multi-decision fusion to complete feature fusion and classification. By combining image transformation with feature decision fusion, we compare our results with other methods through validation tests and demonstrate that our algorithm simplifies the multi-head mechanism while maintaining or improving classification performance. Xujian Qiao, Lei Xing 0005, Anxun Han, Weifeng Liu 0001, Baodi Liu |
Neural Process. Lett. | 5 |
| 2024 | Few-shot image classification via hybrid representation
Baodi Liu, Shuai Shao 0006, Lei Xing 0005, Weifeng Liu 0001, Weijia Cao, Yicong Zhou |
Pattern Recognit. | 1 |
| 2024 | MCNet: Magnitude consistency network for domain adaptive object detection under inclement environments
Jian Pang, Weifeng Liu 0001, Bingfeng Zhang, Xinghao Yang, Baodi Liu, Dapeng Tao |
Pattern Recognit. | 5 |
| 2024 | Weight Saliency search with Semantic Constraint for Neural Machine Translation attacks
Wen Han, Xinghao Yang, Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001 |
Pattern Recognit. Lett. | 3 |
| 2024 | FADS: Fourier-Augmentation Based Data-Shunting for Few-Shot ClassificationabstractCollecting a substantial number of labeled samples is infeasible in many real-world scenarios, thereby bringing out challenges for supervised classification. The research on Few-Shot Classification (FSC) aims to address this issue. Current FSC methods mainly leverage ideas such as meta-learning, self-supervised learning, and data augmentation. Among them, data augmentation appears to be an extremely efficient approach to alleviate the aforementioned data-deficiency problem. Here, we propose a novel data augmentation based FSC method termed Fourier-Augmentation based Data-Shunting (FADS). FADS mainly contains two operations, namely Fourier-based data augmentation (FDA) and data shunting. (i) Fourier transform has a desirable property for classification tasks: the image’s phase and amplitude components in the frequency domain correspond to its high-level structure (i.e., semantic) and low-level style (i.e., statistic) information, which do not interfere with each other. Inspired by this observation, we design the FDA operation, which changes the amplitude spectrum of the to-be-augmented images to obtain new images of the same category. (ii) Then we design the data shunting operation to cooperate with the FDA to accomplish FSC. Specifically, it splits the augmented data into different groups to get independent, weak decisions and then fuses them to obtain a unified, strong decision. We conduct experiments on four benchmark datasets. Results show that utilizing our method brings a performance gain of 0.3%-2% in terms of classification accuracy, compared with the classical methods. Shuai Shao 0006, Yan Wang 0076, Bin Liu 0021, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | EME: Energy-Based Multiexpert Model for Long-Tailed Remote Sensing Image ClassificationabstractThe distribution of remote sensing scene images often follows a long-tailed pattern, where there is an abundance of samples in a few dominant classes and a scarcity of samples in most other classes. This presents two major challenges when it comes to identifying this type of data: Head-Dominance: Models trained on such data tend to prioritize the dominant classes, overlooking the tail classes and resulting in poor performance when it comes to recognizing them. Tail-Interference: The presence of tail classes disrupts the learned representations for the head classes, acting as noise that negatively impacts the recognition accuracy of the head data. To address these challenges, we propose an innovative solution called the energy-based multiexpert (EME) model. The core concept behind this approach is to utilize energy-based discriminators (EDors) to separate the data into head and tail categories. Subsequently, we design multiple experts to classify the head and tail data separately, ensuring that the significant differences in data volume between these categories do not interfere with each other. Experimental results obtained by applying the EME model to three remote sensing datasets demonstrate its efficiency, outperforming current state-of-the-art (SOTA) methods. These findings underscore the effectiveness of our proposed approach in addressing the challenges posed by the long-tailed distribution in remote sensing scene images. Shuai Shao 0006, Shiyuan Zhao, Weifeng Liu 0001, Dapeng Tao, Baodi Liu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | An Ultralightweight Hybrid CNN Based on Redundancy Removal for Hyperspectral Image ClassificationabstractConvolutional neural network (CNN)-based hyperspectral image (HSI) classification models often exhibit high volume and complexity. This not only poses challenges in deploying them on mobile and embedded devices due to storage and power constraints but also introduces a dilemma between the growing demand for labeled samples and the high cost associated with manual labeling. To address these challenges, we propose an ultra-lightweight hybrid CNN based on redundancy removal (ULite-R2HCN), specifically designed for HSI classification in scenarios with limited samples. To reduce computational costs and enhance feature extraction effectiveness, we focus on optimizing the widely used depthwise convolution (DW-Conv) and pointwise convolution (PW-Conv) in the lightweight HSI classification model. For DW-Conv, we design a spatial convolution with redundancy removal (R2Spatial-Conv). This involves the design of multi-scale 3D convolution kernels with specific structures instead of 2D convolution kernels, aiming to reduce redundant convolution kernels and extract multi-scale spatial features. Simultaneously, for PW-Conv, we design a spectral convolution with redundancy removal (R2Spectral-Conv). This utilizes a “copy-splicing-grouping” structure to extract spectral features within arbitrary range intervals, effectively reducing redundant spectral extractions and capturing long-range spectral relationships. Numerous experiments have shown that the proposed ULite-R2HCN achieves higher classification accuracy with an ultra-light volume for a few training samples. In addition, sufficient ablation experiments also verified the advanced performance of the designed R2Spatial-Conv and R2Spectral-Conv. Xiaohu Ma, Wuli Wang, Wei Li 0032, Jianbu Wang, Guangbo Ren, Peng Ren 0001, Baodi Liu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Gradient Guided Multiscale Feature Collaboration Networks for Few-Shot Class-Incremental Remote Sensing Scene ClassificationabstractFew-shot class-incremental learning has recently received significant research focus in remote sensing scene classification (FSCIL-RSSC). The success of FSCIL-RSSC relies on the robustness of the feature backbone and classifiers. Existing works focus on improving classifier adaptation, but little attention is paid to the importance of backbone robustness on the recognition ability of new class samples’ embeddings. Due to the large distribution shift between old and new classes, FSCIL-RSSC using high-layer (single-scale) features may not adapt flawlessly to new categories. To solve the issue, we put forward a gradient guided multiscale feature collaboration network (G-MFCN) for FSCIL-RSSC. Specifically, we introduce a parallel hierarchy strategy to simultaneously capture the multifeature discriminative information of the same sample. Then, a gradient guide block is designed to automatically pick out the optimal values of different convolution blocks for multifeature fusion. Finally, the classical feature pyramid network is introduced for multiscale fusion to obtain more obvious discriminative features of RSSC. More importantly, our proposed G-MFCN is a simple and adaptable module, which can combine any existing FSCIL frameworks to further improve the optimized classifiers’ effectiveness for the FSCIL-RSSC scenario. Extensive experiments on four benchmarks demonstrate that the proposed G-MFCN achieves significant improvements in comparison to existing FSCIL-RSSC methods. Wuli Wang, Sichao Fu, Peng Ren 0001, Guangbo Ren, Qinmu Peng, Baodi Liu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Toward Cross-Domain Class-Incremental Remote Sensing Scene ClassificationabstractClass-incremental (CI) learning has recently received extensive research interest in remote sensing scene classification (CI-RSSC). The existing CI-RSSC methods’ superior performance seriously relies on old (base classes) and new classes (incremental classes) sampled independently from an identical distribution (dataset). In real-world RSSC scenarios, there exist significant distribution shifts between old and new classes, leading to the existing CI-RSSC methods being unable to adjust flawlessly to these new classes. In this article, we propose a novel cross-domain (CD) CI-RSSC framework to solve the above-mentioned problems, termed CDCI-RSSC. Specifically, a modular sharing-based dynamic extension module is first designed, which only updates specialized modules to extract new class feature embeddings for reducing memory footprint. Then, an effective dynamic alignment guided domain adaptive module (DAM) is further proposed to calculate the dynamic weights of each sample in various fields, which can minimize distribution shifts between source and target domains. Finally, a foreground enhancement module (FEM) is introduced to alleviate the issue of complex background interference in RSSC by increasing the weight of critical regions. Compared with the existing CI-RSSC and CD-RSSC, our proposed CDCI-RSSC framework surmounts the challenge of handling the distribution shifts between source (base session) and target domains (incremental session) while alleviating the limitations of continuous learning of new classes. Extensive experiments on three CDCI scenarios show that the CDCI-RSSC model achieves significant performance improvements in comparison to existing CI-RSSC and CD-RSSC methods. Sichao Fu, Wuli Wang, Peng Ren 0001, Qinmu Peng, Guangbo Ren, Baodi Liu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Deep Location Soft-Embedding-Based Network With Regional Scoring for Mammogram ClassificationabstractEarly detection and treatment of breast cancer can significantly reduce patient mortality, and mammogram is an effective method for early screening. Computer-aided diagnosis (CAD) of mammography based on deep learning can assist radiologists in making more objective and accurate judgments. However, existing methods often depend on datasets with manual segmentation annotations. In addition, due to the large image sizes and small lesion proportions, many methods that do not use region of interest (ROI) mostly rely on multi-scale and multi-feature fusion models. These shortcomings increase the labor, money, and computational overhead of applying the model. Therefore, a deep location soft-embedding-based network with regional scoring (DLSEN-RS) is proposed. DLSEN-RS is an end-to-end mammography image classification method containing only one feature extractor and relies on positional embedding (PE) and aggregation pooling (AP) modules to locate lesion areas without bounding boxes, transfer learning, or multi-stage training. In particular, the introduced PE and AP modules exhibit versatility across various CNN models and improve the model's tumor localization and diagnostic accuracy for mammography images. Experiments are conducted on published INbreast and CBIS-DDSM datasets, and compared to previous state-of-the-art mammographic image classification methods, DLSEN-RS performed satisfactorily. Bowen Han 0001, Luhao Sun, Chao Li 0075, Wenzong Jiang, Weifeng Liu 0001, Dapeng Tao, Baodi Liu |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Attribute Space Analysis for Image Editing
Shuqi Yang, Baodi Liu, Weifeng Liu 0001 |
ICIG (2) | 3 |
| 2023 | Self-Compensating Learning for Few-Shot SegmentationabstractFew-shot segmentation (FSS) has witnessed rapid development. Most existing approaches extract prototypes from support images to segment query images. However, the integrity and validity of these support prototypes cannot be guaranteed. To solve the above drawbacks, we propose a self-compensating strategy, aiming to provide query-aware support information, to build more effective matching between support information and query images. Specifically, we design a prototype compensating module to mine useful information from the query prediction, to update original support prototypes as new query-aware support prototypes. Then the updated prototypes are utilized to perform the second matching with query features. In addition, we also compensate the information of original prior masks on the second matching phase, to improve the quality of prior masks. With improved prototype representations and prior knowledge, our approach can directly improve the performance of different approaches with new state-of-the-art performances. Bingfeng Zhang, Weifeng Liu 0001, Baodi Liu, Siyue Yu |
ICIP | 4 |
| 2023 | Annealing Genetic-based Preposition Substitution for Text Rubbish Example GenerationabstractModern Natural Language Processing (NLP) models expose under-sensitivity towards text rubbish examples. The text rubbish example is the heavily modified input text which is nonsensical to humans but does not change the model’s prediction. Prior work crafts rubbish examples by iteratively deleting words and determining the deletion order with beam search. However, the produced rubbish examples usually cause a reduction in model confidence and sometimes deliver human-readable text. To address these problems, we propose an Annealing Genetic based Preposition Substitution (AGPS) algorithm for text rubbish sample generation with two major merits. Firstly, the AGPS crafts rubbish text examples by substituting input words with meaningless prepositions instead of directly removing them, which brings less degradation to the model’s confidence. Secondly, we design an Annealing Genetic algorithm to optimize the word replacement priority, which allows the Genetic Algorithm (GA) to jump out the local optima with probabilities. This is significant in achieving better objectives, i.e., a high word modification rate and a high model confidence. Experimental results on five popular datasets manifest the superiority of AGPS compared with the baseline and expose the fact: the NLP models can not really understand the semantics of sentences, as they give the same prediction with even higher confidence for the nonsensical preposition sequences. Xinghao Yang, Baodi Liu, Weifeng Liu 0001, Honglong Chen |
IJCAI | 3 |
| 2023 | A Stable Vision Transformer for Out-of-Distribution Generalization
Haoran Yu 0005, Baodi Liu, Yingjie Wang 0007, Kai Zhang 0029, Dapeng Tao, Weifeng Liu 0001 |
PRCV (8) | 2 |
| 2023 | Deep Positional-Representation-Based Local Information Retention Networks for Mammography ClassificationabstractEarly diagnosis of breast cancer is challenging because in the most common mammogram images, the tumor usually occupies only a very small part of the entire image, which often makes deep learning models lose attention to the tumor area. In previous work, most models solved this problem by using ROI labeling to train models, which was expensive and difficult to widely apply. Some recent ROI-free methods use multi-scale features or multi-stage training, which gets rid of the model's dependence on ROI but greatly increases the computational complexity and deployment difficulty, limiting the potential of deep neural networks. Therefore, a deep positional-representation-based local information retention networks (PR-LIR) was proposed. PR-LIR is a lightweight, end-to-end mammogram classification model, which uses positional representation (PR) and multi-scale regional pooling (MRP) modules to locate tumor regions and retain regional semantic information of small target tumors at different scales, without ROI labeling and multi-stage training, and almost no increasement in parameters and computational complexity. In particular, the proposed PR and MRP modules have good generalization performance, which can be applied to most CNN models and improve the classification accuracy of mammography images. Experimental results on two publicly available datasets show that PR-LIR achieves the best AUC and satisfactory accuracy compared to the previous state-of-the-art mammogram classification method. Bowen Han 0001, Luhao Sun, Chao Li 0075, Wenzong Jiang, Weifeng Liu 0001, Dapeng Tao, Baodi Liu |
SMC | 8 |
| 2023 | Object tracking based on siamese network with 3D attention and multiple graph attention
Shilei Yan, Yujuan Qi, Yanjiang Wang 0001, Baodi Liu |
Comput. Vis. Image Underst. | 5 |
| 2023 | CSN: Component supervised network for few-shot classification
Rui Xu 0012, Shuai Shao 0006, Lei Xing 0005, Yujun Wei, Weifeng Liu 0001, Baodi Liu, Yanjiang Wang 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | Generation-based parallel particle swarm optimization for adversarial text attacks
Xinghao Yang, Yupeng Qi, Honglong Chen, Baodi Liu, Weifeng Liu 0001 |
Inf. Sci. | 4 |
| 2023 | Selecting Information Fusion Generative Adversarial Network for Remote-Sensing Image Cloud RemovalabstractThe multi-temporal remote sensing cloud removal method has improved performance, but it lacks a screening mechanism during feature fusion, simply summing and fusing features from different temporal states. This results in the inclusion of unwanted clouds and redundant feature information, hindering the restoration of the landscape under the clouds. To address this, we propose a selective information fusion generative adversarial network (SIF-GAN) for remote sensing image cloud removal. SIF-GAN incorporates channel attention during feature extraction to capture important information in different channels and uses the selective information fusion network to assign weights to the feature information from other temporal states, selecting the crucial features for fusion. The feature of cloud-free regions in different temporal states is utilized maximally by the selection process to recover the image features under clouds. The results of the experiments show that SIF-GAN achieves superior cloud removal performance compared to other methods. Wenzong Jiang, Weifeng Liu 0001, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Deformable Convolutional Network Constrained by Contrastive Learning for Underwater Image EnhancementabstractAutonomous underwater vehicles (AUVs) based on remote sensing technology have been widely applied in various underwater tasks. However, the complex underwater environment leads to challenges such as color distortion, blurred details, and fog effects in the underwater image directly acquired by AUVs. Although numerous existing methods aim to remove the color cast and restore image details, their effectiveness is still limited. This paper proposes a new method based on a deformable convolutional network and constrained by contrastive learning for underwater image enhancement. First, we propose a deformable convolutional residual block (DCRB) to achieve a more precise restoration of texture details by adaptively adjusting the convolution kernel shape. At the same time, we utilize the long-skip connection method of the U-Net architecture to preserve information that is prone to lose in shallow networks. Second, we propose a color contrastive loss function to compare the color difference between distorted images and the ground truth, resulting in a more realistic enhanced image. Finally, experimental results demonstrate that the proposed method outperforms the state-of-the-art methods regarding image quality and visual appeal. Xinran Guo, Weifeng Liu 0001, Dapeng Tao, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Class Centralized Dictionary Learning for Few-Shot Remote Sensing Scene ClassificationabstractRecently, few-shot scene classification has become an important task in the remote sensing (RS) field, mainly solving how to obtain better classification performance when there are insufficient labeled samples. The few-shot scene classification task includes the pretrain stage and meta-test stage. There is no category intersection between these two stages. Thus, the sample distribution of the training set and meta-test set is different, leading to the training model’s weak generalization or portability. To solve this problem, we propose a class-centralized dictionary learning (CCDL) method for the few-shot RS scene classification (FSRSSC). Specifically, in the pretraining stage, we adopt the model pretrained on a large natural images dataset and then fine-tune the network by the RS dataset. Using the pretrained model helps improve the model’s generalization ability. In the meta-test stage, we propose a CCDL classifier, which guarantees the sparse representations of different categories more distant and the same more concentrated. We experiment on several benchmark datasets and achieve superior performance, demonstrating the proposed method’s effectiveness. Lei Xing 0005, Lifei Zhao, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Double Discriminative Constraint-Based Affine Nonnegative Representation for Few-Shot Remote Sensing Scene ClassificationabstractRemote sensing scene classification (RSSC) has recently attracted more attention. However, due to restrictions in the imaging environment and equipment, it is difficult to get a large number of labeled images in remote sensing. This has led to the emergence of few-shot learning for RSSC, which aims to achieve better performance with few labeled samples. Remote sensing images’ large interclass similarity may cause classification confusion. To overcome this issue, this study proposes a double discriminative constraints-based affine nonnegative representation for few-shot RSSC. To be specific, we devise a novel representation-based classifier with two discriminative constraint terms in the objective function and utilize affine nonnegative constraints to restrict the learned parameters. These constraints reduce the correlation between classes and strengthen the class specificity of the learned parameters. Experiments on benchmark datasets demonstrate the effectiveness of our method. Tianhao Yuan, Weifeng Liu 0001, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Shared Dictionary Learning Via Coupled Adaptations for Cross-Domain Classification
Yuying Cai, Baodi Liu, Weijia Cao, Honglong Chen, Weifeng Liu 0001 |
Neural Process. Lett. | 3 |
| 2023 | GSA4FDA: Deep Geometric and Statistic Alignment for Fewer Labeled Domain Adaptation
Yuying Cai, Baodi Liu, Xinghao Yang, Dapeng Tao, Weifeng Liu 0001 |
Neural Process. Lett. | 2 |
| 2023 | Dynamic Feature Attention Network for Remote Sensing Image Dehazing
Wenzong Jiang, Weifeng Liu 0001, Weijia Cao, Baodi Liu |
Neural Process. Lett. | 5 |
| 2023 | Feature Fusion Based Parallel Graph Convolutional Neural Network for Image Annotation
Weifeng Liu 0001, Baodi Liu |
Neural Process. Lett. | 4 |
| 2023 | Cross-Domain Few-Shot classification via class-shared and class-specific dictionaries
Lei Xing 0005, Baodi Liu, Dapeng Tao, Weijia Cao, Weifeng Liu 0001 |
Pattern Recognit. | 3 |
| 2023 | Subspace prototype learning for few-Shot remote sensing scene classification
Wuli Wang, Lei Xing 0005, Peng Ren 0001, Yumeng Jiang, Baodi Liu |
Signal Process. | 6 |
| 2023 | Non-Contrastive Nearest Neighbor Identity-Guided Method for Unsupervised Object Re-IdentificationabstractRecently, self-paced contrastive learning has emerged as a promising method for unsupervised object re-identification. These methods generate pseudo labels, store centroid features in the memory bank, and periodically update them. However, affected by the performance of the clustering method, within each cluster exists inevitably noisy instances, and self-paced contrastive learning usually requires a large number of negative samples from various classes, where false-negative samples give rise to the class collision issue. These lead to performing incorrect model optimization. In this paper, we propose a non-contrastive nearest neighbor identity-guided (NNNI) method to overcome these challenges. The advantage of NNNI is to provide the model with a highly accurate prior. Specifically, this method relies on the random identity sampler commonly used in re-identification tasks to provide the network with a regression target of the nearest neighbors of the same identity within a mini-batch. It encodes more and more information in an iterative process through a Siamese network with an exponential moving average to train high-quality representations. NNNI alleviates the negative effects of noise instances and corrects class collision issues during training. Extensive experiments show that our method is effective on unsupervised object re-identification and achieves state-of-the-art performance on three large-scale person re-identification datasets and one large-scale vehicle re-identification dataset, which is competitive with even supervised methods. Dongchen Han, Weifeng Liu 0001, Mingchen Zou, Baodi Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Attention-Based Multi-View Feature Collaboration for Decoupled Few-Shot LearningabstractDecoupled Few-shot learning (FSL) is an effective methodology that deals with the problem of data-scarce. Its standard paradigm includes two phases: (1) Pre-train. Generating a CNN-based feature extraction model (FEM) via base data. (2) Meta-test. Employing the frozen FEM to obtain the novel data features, then classifying them. Obviously, one crucial factor, the category gap, prevents the development of FSL, i.e., it is challenging for the pre-trained FEM to adapt to the novel class flawlessly. Inspired by a common-sense theory: the FEMs based on different strategies focus on different priorities, we attempt to address this problem from the multi-view feature collaboration (MVFC) perspective. Specifically, we first denoise the multi-view features by subspace learning method, then design three attention blocks (loss-attention block, self-attention block and graph-attention block) to balance the representation between different views. The proposed method is evaluated on four benchmark datasets and achieves significant improvements of 0.9%-5.6% compared with SOTAs. Shuai Shao 0006, Lei Xing 0005, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Dynamic Adaptive Attention-Guided Self-Supervised Single Remote-Sensing Image DenoisingabstractOptical remote sensing images are widely used in many fields, and local complex texture details in images usually play a critical role in downstream tasks. However, noise interference will destroy the complex texture in the image, thus reducing the accuracy of downstream tasks. The current attention mechanism usually focuses on the global high-level features in the image, so it cannot effectively focus on the high-frequency information in the local complex texture in the remote sensing image, and obtaining clean remote sensing images to train neural networks is difficult. Therefore, applying the current depth learning based natural image denoising methods directly to optical remote sensing images is challenging. To solve these problems, we propose a dynamic adaptive attention guided self-supervised single remote sensing image denoising network (DAA-SSID). We construct a dynamic adaptive attention module (DAAM) by dynamically calculating the activation intensity of each neuron and combining the spatial feature information extracted from remote sensing images. It can effectively extract complex texture features from remote sensing images when only a single remote sensing image participates in training. And we use independent random Bernoulli sampling in the training and inference stages respectively to prevent over-fitting caused by single-image training. Therefore, compared with other self-supervised denoising methods, our proposed model can denoise remote sensing images with more complex textures when only a single image destroyed by noise is used as the training input. Experiments on synthetic additive gaussian noise data and authentic noise data have shown that the proposed model achieves satisfactory results. Minghao Liu 0016, Wenzong Jiang, Weifeng Liu 0001, Dapeng Tao, Baodi Liu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | RAN: Region-Aware Network for Remote Sensing Image Super-ResolutionabstractThe remote sensing (RS) image super-resolution (SR) algorithm aims to reconstruct a high-resolution (HR) image with rich texture details from a given low-resolution (LR) image, improving the spatial resolution. It has been widely concerned in remote sensing image processing and application. Most current deep learning-based methods rely on paired training datasets. However, most datasets are often based on bicubic degradation. This single construction way limits the performance of the pre-trained network. Moreover, SR is an ill-posed problem in that multiple SR images are constructed from a single LR input. This paper proposes a Region-Aware Network (RAN) for remote sensing image super-resolution to alleviate the above issues. First, we introduce the contrastive learning strategy to mine the latent degraded representation of the image and serve as the prior knowledge of the network. Considering the RS images are acquired in specific scenes that have apparent self-similarity. Then, we propose a Region-Aware Module (RAM) based on attention mechanisms and the graph neural network to explore region information and cross-patch self-similarity. Extensive experiments have demonstrated that the proposed RAN adapts to RS image super-resolution tasks with various degradations and performs better in constructing texture information. Baodi Liu, Lifei Zhao, Shuai Shao 0006, Weifeng Liu 0001, Dapeng Tao, Weijia Cao, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | A Lightweight Hybrid Convolutional Neural Network for Hyperspectral Image ClassificationabstractRecent studies have demonstrated the potential of hybrid convolutional models that combine 3D and 2D convolutional neural networks (CNNs) for hyperspectral image (HSI) classification. However, these models do not fully utilize the benefits of hybrid convolution due to inefficient connections between the two types of CNNs. Moreover, most CNNs, including hybrid models, require a significant number of parameters and computational resources for accurate classification, which increases the need for labeled samples and computational cost. Although the common lightweight strategies like depthwise separable convolution (DSC) can reduce parameters and computation compared to normal convolution (NC), they often compromise accuracy. To address these challenges, we propose a lightweight hybrid convolutional neural network (Lite-HCNet) for HSI classification with minimal model parameters and computational effort. Firstly, we design a novel channel attention module (NCAM) and combine it with a convolutional kernel decomposition (CKD) strategy to propose a lightweight and efficient DSC (LE-DSC) deployed in Lite-HCNet. The LE-DSC not only reduces the DSC volume further but also enhances its performance. Secondly, a lightweight and efficient hybrid convolutional layer (LE-HCL) is designed in Lite-HCNet to explore the efficient connection structure between 3D CNNs and 2D CNNs. Experiments show that the Lite-HCNet reduces the required computational cost and practical deployment difficulty while offering advanced performance with a small number of training samples. Furthermore, abundant ablation experiments confirm the superior performance of the designed LE-DSC. Xiaohu Ma, Xudong Kang, Huawei Qin, Wuli Wang, Guangbo Ren, Jianbu Wang, Baodi Liu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | ARPCNN: Auxiliary Review-Based Personalized Attentional CNN for Trustworthy RecommendationabstractConvolutional neural network (CNN)-based recommender systems are playing an increasingly significant role in the vigorous development of Industrial Internet of Things, and have made great contributions to analyzing and mining a large amount of data to provide various services for terminal users. However, as the lack of explainability in deep learning, users often have low trust in the system due to their incomprehension of recommendation results. In addition, recommender systems have been facing a serious sparsity problem, and relying only on sparse rating data to learn user preferences and similarities may face malicious recommendation attacks. The abovementioned problems have been hindering the further improvement of recommendation performance. Therefore, in order to effectively alleviate the sparsity problem and meanwhile enhance the trustworthiness, an auxiliary review-based personalized attentional CNN (ARPCNN) is proposed in this article. By applying the proposed personalized word-level attention mechanism and personalized review-level attention mechanism in parallel CNNs, critical words and informative reviews are given high attention weights. Moreover, a user auxiliary network is proposed, which regards the reviews written by kindred spirits who have a trust relationship with the user as auxiliary reviews, and effectively extracts the user’s auxiliary review features, thereby achieving more accurate user modeling to improve the recommendation performance. Extensive experiments are conducted on four real-world datasets, and the results show that the performance of the proposed model is better than that of baselines, which verifies the effectiveness of ARPCNN. Zhe Li 0026, Honglong Chen, Zhichen Ni, Xiaogang Deng, Baodi Liu, Weifeng Liu 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | Enrich Features for Few-Shot Point Cloud ClassificationabstractRecently, many existing fully supervised methods for point cloud classification have strongly promoted the development of point cloud learning. However, these methods require a lot of labeled data as support, which is challenging to obtain. To alleviate this problem, we propose a novel few-shot point cloud classification method to classify new categories given a few labeled samples. Specifically, we apply the feature supplement module to enrich the geometric information of points and then aggregate multi-scale features through the channel-wise attention module while reducing the computational complexity. Finally, we introduce a classifier to classify the point cloud features under the few-shot learning setup to predict its label. We carry out experimental verification on the benchmark dataset and achieve state-of-the-art performance. Hengxin Feng, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu |
ICASSP | 4 |
| 2022 | Agcyclegan: Attention-Guided Cyclegan for Single Underwater Image RestorationabstractUnderwater image restoration is a fundamental problem in image processing and computer vision. It has broad application prospects for underwater operations, especially underwater robot operations. The challenging work is how to keep the color authenticity of the captured underwater image. In this paper, we propose a novel network architecture based on CycleGAN. Specifically, in the generator part, we adopt the U-Net structure because the long skip connection of U-Net will obtain more detailed information. Besides, we append the pixel-level attention block to provide greater flexibility for detail structure modeling. It assigns different weights to each channel to pay more attention to the critical feature. We also verify its generalization performance on several benchmark datasets. The extensive experiments with comparisons to state-of-the-art approaches demonstrate the superiority of the proposed model. Zhenlong Wang, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu |
ICASSP | 4 |
| 2022 | MSL-FER: Mirrored Self-Supervised Learning for Facial Expression RecognitionabstractFacial Expression Recognition (FER) in the wild is a significant yet challenging classification task due to the inter-class similarities and intra-class variations. Recently, a large number of methods can extract expression features effectively. However, the intra-class variations mainly caused by various uncertainties (such as identity, pose, and occlusion) are difficult to capture in advance, and the cost of labeling these uncertainties is high. To tackle this challenge, we propose a novel Mirrored Self-supervised Learning FER (MSL-FER) method. The ground truth of self-supervised learning comes from the data itself rather than from human annotations, and horizontal inversion preserves emotional information without altering the facial structure. Specifically, MSL-FER introduces a binary classification task to recognize the 2D mirror operation in a self-supervised learning method. And we also combine our MSL-FER with an attention network to discriminate features along its dimensions selectively. Experiments on two public wild FER datasets show that our MSL-FER approach outperforms the baseline and other state-of-the-art methods with 87.92% on RAF-DB and 70.68% on FER2013. Xiangshuai Pan, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu |
ICIP | 5 |
| 2022 | Image Super-Resolution Based on Adaptive Feature Fusion Channel Attention
Qizhang Song, Baodi Liu, Weifeng Liu 0001 |
ICONIP (3) | 2 |
| 2022 | Virtual Try-on via Matching Relation with Landmark
Xingxing Yao, Baodi Liu, Weifeng Liu 0001 |
ICONIP (3) | 3 |
| 2022 | EMAS: Efficient Meta Architecture Search for Few-Shot LearningabstractWith the progress of few-shot learning, it has been scaled to many domains in the real world which have few labeled data, such as image classification and object detection. Many efforts for data embedding and feature combination have been made by designing a fixed neural architecture that can also be extended to variable and adaptive neural architectures for better performance. Recent works leverage neural architecture search technique to automatically design networks for few-shot learning but it requires vast computation costs and GPU memory requirements. This work introduces EMAS, an efficient method to speed up the searching process for few-shot learning. Specifically, we build a supernet to combine all candidate operations and then adopt gradient-based methods to search. Instead of training the whole supernet, we adopt Gumbel reparameterization technique to sample and activate a small subset of operations. EMAS handles a single path in a novel task adapted with just a few steps and time. A novel task only needs to learn fewer parameters and compute less content. During meta-testing, the task can well adapt to the network architecture although only with a few iterations. Empirical results show that EMAS yields a fair improvement in accuracy on the standard few-shot classification benchmark and is five times smaller in time. Dongkai Liu, Honglong Chen, Baodi Liu, Weifeng Liu 0001 |
ICTAI | 4 |
| 2022 | Automated Drawing Psychoanalysis via House-Tree-Person TestabstractThe increase of human psychological illness in today's fast paced and high stress world makes it essential to detect the warning signals of psychological problems. As the most representative drawing psychoanalysis method, House-Tree-Person (HTP) test is widely used in psychological assessment with the benefit of simplicity, non-verbal, and repeatability. HTP test can reveal the individual subconscious of the psychological state through the picture content of house, tree, and person drawn by the patient. Currently, HTP test is conducted by the therapist in person, which makes it time consuming and the results are mostly affected by the therapist's experience. Therefore, it is helpful and necessary to build an automated method to improve the objectivity, reliability, and efficiency of HTP test. In this paper, we propose an automated psychometric drawing screening method that forms the relationship between the psychological state and drawing feature. Specifically, we extract the key features including size, position, and shadow of the drawing, and then combine these features to construct a psychological state classifier. The proposed method can effectively screen out negative drawings for further diagnosis and treatment. Experiments are carried out on a builded dataset with the drawings from a psychological testing center of college. Experimental results demonstrate the effect and superiority of the proposed method. Baodi Liu, Weifeng Liu 0001 |
ICTAI | 3 |
| 2022 | MVFF: Multi-view Feature Fusion for Few-shot Remote Sensing Image Scene ClassificationabstractCompared to deep learning methods, few-shot learning methods do not need many labeled images. Therefore, few-shot remote sensing image scene classification has been studied extensively. However, obtaining effect information from the limited amount of labeled samples is a great challenge. Most methods only extract features from a single perspective of remote sensing images. Such information is scarce and even misleading. To address the problem, we propose a multi-view feature fusion (MVFF) method. Specifically, first, train two feature extractor networks on the original image dataset and remote sensing image dataset, respectively. And each model extracts features before and after average pooling, and we obtain four kinds of features for a remote sensing image. Second, we calculate fusion weights from support set features using Multi-Head Feature Collaboration (MHFC) method and four classifiers. Third, we utilize the weights to fuse predictive probability matrices and thus obtain the labels of query set samples. We implement experiments on three benchmark remote-sensing image datasets to validate the performance of our method. And the results demonstrate that our approach effectively handles few-shot remote sensing image scene classification. Anxun Han, Lei Xing 0005, Weifeng Liu 0001, Baodi Liu |
SMC | 4 |
| 2022 | Multi-task Facial Expression Recognition With Joint Gender LearningabstractFacial Expression Recognition (FER) in the wild is a significant yet challenging topic in computer vision due to the feature inconsistency caused by the individual specificity of facial expressions. In addition to variations of facial expressions caused by identity, pose, and occlusion, gender also affects the face of human emotions. Even though males, females, and infants share the same facial expressions, their characteristics are vastly different. To capture the effect of gender on facial expressions, we propose a novel multi-task FER method with joint gender learning. First, in addition to the original emotion labels of face images, we annotate gender labels, including male, female, and infant. Second, we introduce a gender-aware multi-task convolutional neural network for FER, which can learn the emotion and gender features of faces. Compared with single-task expression recognition methods, our proposed framework for introducing gender feature learning can significantly achieve higher performance on FER in the wild. Finally, we verify the effectiveness of our framework on two public wild FER datasets, RAF-DB and FER2013. And the results show that the gender learning auxiliary task is beneficial to the improvement of the performance of FER. Xiangshuai Pan, Qingtao Xie, Weifeng Liu 0001, Baodi Liu |
SMC | 4 |
| 2022 | Multi-relational Semantic Distillation for Few-Shot Object DetectionabstractWhile few-shot object detection(FSOD) has been developed to a certain extent, it is still a large margin from practical applications. Most existing methods use traditional object detection methods as the basic framework is improved to a limited extent. Previous methods often ignore the special characterization relationship between support and query images. This paper fully investigates the effect of support images on detection performance and proposes a new FSOD method called Multi-relational Semantic Distillation (MSD). Our approach aims to improve FSOD performance by building a multi-relational semantic representation model with support and query features. In addition, we propose a support enhancement (SE) module based on the self-attention mechanism to enhance the useful information in the support features to mitigate the negative impact of low-quality support images. To verify the effectiveness of MSD, we conduct sufficient experiments on Pascal VOC and MS-COCO datasets. Experiments show that MSD achieves competitive results at low shots compared to other state-of-the-art few-shot detectors. Qingtao Xie, Xiangshuai Pan, Weifeng Liu 0001, Baodi Liu |
SMC | 4 |
| 2022 | A graph convolutional neural network model with Fisher vector encoding and channel-wise spatial-temporal aggregation for skeleton-based action recognitionabstractAbstract Skeleton‐based action recognition is an inspired yet challenging task in computer vision. Recently, the latest graph convolutional network (GCN), which generalises well‐established convolutional neural networks to non‐Euclidean structures, is proven to be highly successful for action recognition from body skeleton data. However, the GCN architecture has not been fully studied. In this work, a Fisher vector (FV) encoding based GCN architecture (FV‐GCN) is proposed, which exceeds the limitations of existing GCN‐based methods by combining the GCN model with FV encoding. A channel‐wise spatial–temporal aggregation function to preserve spatial–temporal information in the whole action clip and integrate it into the FV‐GCN architecture is also presented. Since FV is different from the GCN structure, this hybrid architecture that incorporates the advantages of both algorithms can discover complementary information of feature representation effectively. On two challenging human action datasets, kinetics, and NTU‐RGBD, improved performance is demonstrated over the baseline method, and the FV‐GCN is better or comparable to some state‐of‐the‐art methods. Yanjiang Wang 0001, Sichao Fu, Baodi Liu, Weifeng Liu 0001 |
IET Image Process. | 4 |
| 2022 | Object re-identification with distribution corrected ranking list
Dongchen Han, Shuai Shao 0006, Weifeng Liu 0001, Baodi Liu |
Neurocomputing | 4 |
| 2022 | Multi-view learning for hyperspectral image classification: An overview
Baodi Liu, Kai Zhang 0029, Honglong Chen, Weijia Cao, Weifeng Liu 0001, Dapeng Tao |
Neurocomputing | 2 |
| 2022 | DLDL: Dynamic label dictionary learning via hypergraph regularization
Shuai Shao 0006, Rui Xu 0012, Zhenfang Wang, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu |
Neurocomputing | 6 |
| 2022 | Learning task-specific discriminative embeddings for few-shot image classification
Lei Xing 0005, Shuai Shao 0006, Weifeng Liu 0001, Anxun Han, Xiangshuai Pan, Baodi Liu |
Neurocomputing | 6 |
| 2022 | Adaptive graph convolutional collaboration networks for semi-supervised classification
Sichao Fu, Senlin Wang, Weifeng Liu 0001, Baodi Liu, Xinhua You, Qinmu Peng, Xiaoyuan Jing |
Inf. Sci. | 4 |
| 2022 | Adaptive multi-scale transductive information propagation for few-shot learning
Sichao Fu, Baodi Liu, Weifeng Liu 0001, Bin Zou 0002, Xinhua You, Qinmu Peng, Xiaoyuan Jing |
Knowl. Based Syst. | 2 |
| 2022 | Hyperspectral Image Classification Using CNN-Enhanced Multi-Level Haar Wavelet Features Fusion NetworkabstractConvolutional neural networks (CNNs) are widely utilized in hyperspectral image (HSI) classification due to their powerful capability to automatically learn features. However, ordinary CNN mainly captures the spatial characteristics of HSI and ignores the spectral information. To alleviate the issue, this work proposes a CNN-enhanced multi-level Haar wavelet features fusion network (CNN-MHWF2N), which combines the spatial features obtained through 2-D-CNN with the Haar wavelet decomposition features to obtain sufficient spectral–spatial features. Specifically, factor analysis is first used to reduce the HSI dimension. Then, four-level decomposition features are obtained through the Haar wavelet decomposition algorithm, which of them are, respectively, concatenated with four-layer convolution features for combining spatial with spectral information. In this way, spectral–spatial features achieve better information interaction. Besides, a double filtrating feature fusion module is designed, which is operated following each level spectral–spatial features to obtain finer characteristics. Finally, those recognizable features are merged via a fusion operator. The whole designed model is conducive to enhancing the final HSI classification performance. In addition, experiments also reveal that the designed model is superior on three benchmark databases compared with the state-of-the-art approaches. Wenhui Guo, Guixun Xu, Baodi Liu, Yanjiang Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | U-Shaped Attention Connection Network for Remote-Sensing Image Super-ResolutionabstractIn recent years, deep learning-based remote-sensing image super-resolution (SR) methods have made significant progress, and these methods require a large number of synthetic data for training. To obtain sufficient training data, researchers often generate synthetic data via fixed bicubic downsampling methods. However, the synthesized data cannot reflect the complex degradation process of real remote-sensing images. Thus, performance will dramatically reduce when these methods work in real low-resolution (LR) remote-sensing images. This letter proposes a U-shaped attention connection network (US-ACN) for remote-sensing image SR to solve this issue. Our US-ACN does not rely on any synthetic external dataset for training and merely requires one LR image to complete the training. The US-ACN utilizes remote-sensing images’ strong internal feature repetitiveness and fully learns this internal repetitive feature through a well-designed US-ACN to achieve the remote-sensing image SR. In addition, we design a 3-D attention module to generate effective 3-D weights by modeling channel and spatial attention weights, which is more helpful for the learning of internal features. Through the U-shaped connection among attention modules, context information propagation and attention weights learning are fully utilized. Many experiments show that our US-ACN adequately adapts to the remote-sensing image SR in various situations and performs advanced performance. Wenzong Jiang, Lifei Zhao, Yanjiang Wang 0001, Weifeng Liu 0001, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Multiorder Interaction Information Embedding-Based Multiview Fusion-Aided Hyperspectral Image ClassificationabstractHyperspectral images (HSI) are obtained from hyperspectral imaging sensors, which capture information in hundreds of spectral bands of objects. However, how to take full advantage of spatial and spectral information from many spectral bands to improve the performance of HSI classification remains an open question. Many HSI classification works have recently been reported by employing multi-view learning (MVL) algorithms that can fully use complementary information between different view features and thus have received widespread attention. This paper proposes a multi-view fusion network based on multi-order interaction information embedding for HSI classification. Firstly, the correlation matrix between spectral bands is used to divide the original data into multiple subsets as local views. The subset after the Segmented-PCA process is used as the global view. Secondly, the features of different views are extracted separately using a feature extraction network and mapped to the same dimension. Pre-fusion is achieved by multi-order interaction of various view features. Finally, loss-weighted fusion is applied to each view according to its contribution to the classification task. To evaluate the effectiveness of the proposed method, complete experiments were conducted on three commonly used HSI datasets, namely Pavia University, Houston 2013, and Houston 2018. The experimental results demonstrate that the proposed method improves the classification performance of existing feature extraction networks and is more competitive with other methods in the field. Weijia Cao, Kai Zhang 0029, Baodi Liu, Dapeng Tao, Weifeng Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Location Soft-Aggregation-Based Band Weighting for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) comprise hundreds of continuous spectral bands. How to effectively exploit the abundant spectral features of HSI to improve its classification accuracy is the focus of the research. Band weighting (BW) is extensively used due to its ability to emphasize usefully and suppress noisy bands adaptively. Most proposed works aggregate global information to construct band representation vectors in simple ways such as global averaging pooling. Those ways are not capable of retaining a more discriminating feature. Furthermore, modeling for interpixel positional relationships is something they have not considered. To address these problems, we propose a position embedding and importance aggregation BW module. The position embedding section encodes the position information by two 1-D features so that remote dependencies in one spatial direction can be obtained while retaining accurate position information in the other spatial direction. The importance aggregation section aggregates the global information. Finally, a group of weights is learned to recalibrate the raw input. Experiments on three public datasets of HSI demonstrate that our methods obtain competitive results compared to other methods. Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | A Hybrid CNN Based on Global Reasoning for Hyperspectral Image ClassificationabstractIn recent years, convolutional neural networks (CNNs) have been widely used in hyperspectral images (HSIs) classification. However, 2-D CNN, 3-D CNN, and even the newly emerged hybrid CNN (HCNN) all require multiple or deep CNN layers to obtain excellent classification performance, which inevitably results in the high complexity and the need for a large number of training samples. Moreover, as a local operator, convolution is challenging to fully use global information. To solve the above two issues, we design a HCNN based on global reasoning (GloRe-HCNN) for HSI classification. On the one hand, the GloRe-HCNN uses only one layer of 3-D CNN and one layer of 2-D CNN to jointly extract the spatial–spectral features of HSI. On the other hand, we contrive a spatial–spectral global reasoning unit (SS-GloRe-Unit) to take the place of stacked multilayer 3-D CNN for extracting global features fully. We select small training samples in three standard datasets and compare them with state-of-the-art CNN methods. Numerous experiments show that our GloRe-HCNN performs advanced performance. Wuli Wang, Xiaohu Ma, Linchun Leng, Yanjiang Wang 0001, Baodi Liu, Jinfeng Sun |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Rethinking Few-Shot Remote Sensing Scene Classification: A Good Embedding Is All You Need?abstractIn recent years, few-shot remote sensing scene classification (FSRSSC) has attracted more and more attention. For FSRSSC, most methods currently focus on designing a meta-learning algorithm, which obtains meta-knowledge from limited samples and then applies it to novel tasks. In this work, on the one hand, we optimize the training pipeline of the feature extractor; on the other hand, we apply a novel model fusion method further to optimize the feature extractor capability of the feature extractor. We show a novel few-shot remote sensing scene classification baseline: learning two feature representations through using two self-supervised methods on the meta-training set and then fusing the two representations into one. Then, training a linear classifier on this representation achieves state-of-the-art performance. It shows that training a good feature extractor can be more efficient than complex meta-learning algorithms for FSRSSC. We believe that our results can inspire a rethinking of few-shot remote sensing scene classification benchmarks. Lei Xing 0005, Yuteng Ma, Weijia Cao, Shuai Shao 0006, Weifeng Liu 0001, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Learning to Cooperate: Decision Fusion Method for Few-Shot Remote-Sensing Scene ClassificationabstractRecently, remote-sensing scene classification has become an essential primary research topic. Nowadays, scholars have proposed various few-shot remote-sensing scene classification methods to achieve superior performance with few labeled data. Most of the prior work utilized a meta-learning strategy, which suffered from too little data affecting performance. In this letter, we apply the pre-trained feature extractor for image embedding. Meanwhile, because of the negative transfer problem caused by the inadaptability of the pre-trained feature extractor to remote-sensing data, we propose to exploit two pre-trained models to classify the remote-sensing scene, respectively. Then we fuse the decision to obtain the final classification category. We design a decision attention module to automatically update combination weights for each decision. It comprehensively considers the contribution of various decisions and further improves the discrimination of features. We conduct comprehensive experiments to validate the method and achieve state-of-the-art performance on two benchmark remote-sensing scene datasets, namely NWPU-RESISC45 and UC Merced. Lei Xing 0005, Shuai Shao 0006, Yuteng Ma, Yanjiang Wang 0001, Weifeng Liu 0001, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Class Shared Dictionary Learning for Few-Shot Remote Sensing Scene ClassificationabstractIn the field of remote sensing, it is infeasible to collect a large number of labeled samples due to imaging equipment and the imaging environment. Few-Shot Learning (FSL) is the dominant method to alleviate this problem, which pursues quickly adapting to novel categories from a limited number of labeled samples. The few-shot Remote Sensing Scene Classification (RSSC) generally includes the pre-training and meta-test phases. However, a “negative transfer” problem exists that data categories in both phases are different. It causes the pre-trained feature extractor to be unable well-adapted to the novel data category. This paper proposes Class Shared Dictionary Learning for Few-Shot Remote Sensing Scene Classification (CSDL) to address this issue. Specifically, this paper designs the Mirror-based Feature Extractor (MFE) in the pre-training phase, constructing a self-supervised classification task to improve the feature extractor robustness. Furthermore, this paper proposes a Class Shared Dictionary classifier (CSD) based on dictionary learning. The CSD projects the novel data feature in meta-test into subspace to reconstruct more discriminative features and complete the classification task. Extensive experiments on remote sensing datasets have demonstrated that the proposed CSDL achieves the advanced classification performance. Lei Xing 0005, Lifei Zhao, Weijia Cao, Xinmin Ge, Weifeng Liu 0001, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | DMH-FSL: Dual-Modal Hypergraph for Few-Shot Learning
Rui Xu 0012, Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001 |
Neural Process. Lett. | 2 |
| 2022 | Co-Learning for Few-Shot Learning
Rui Xu 0012, Lei Xing 0005, Shuai Shao 0006, Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001 |
Neural Process. Lett. | 4 |
| 2022 | MDFM: Multi-Decision Fusing Model for Few-Shot LearningabstractIn recent years, researchers pay growing attention to the few-shot learning (FSL) task to address the data-scarce problem. A standard FSL framework is composed of two components: i) Pre-train. Employ the base data to generate a CNN-based feature extraction model (FEM). ii) Meta-test. Apply the trained FEM to the novel data (category is different from base data) to acquire the feature embeddings and recognize them. Although researchers have made remarkable breakthroughs in FSL, there still exists a fundamental problem. Since the trained FEM with base data usually cannot adapt to the novel class flawlessly, the novel data’s feature may lead to the distribution shift problem. To address this challenge, we hypothesize that even if most of the decisions based on different FEMs are viewed asweak decisions, which are not available for all classes, they still perform decent in some specific categories. Inspired by this assumption, we propose a novel method Multi-Decision Fusing Model (MDFM), which comprehensively considers the decisions based on multiple FEMs to enhance the efficacy and robustness of the model. MDFM is a simple, flexible, non-parametric method that can directly apply to the existing FEMs. Besides, we extend the proposed MDFM to two FSL settings (e.g., supervised and semi-supervised settings). We evaluate the proposed method on five benchmark datasets and achieve significant improvements of 3.4%-7.3% compared with state-of-the-arts. Shuai Shao 0006, Lei Xing 0005, Rui Xu 0012, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | GCT: Graph Co-Training for Semi-Supervised Few-Shot LearningabstractFew-shot learning (FSL), purposing to resolve the problem of data-scarce, has attracted considerable attention in recent years. A popular FSL framework contains two phases: (i) the pre-train phase employs the base data to train a CNN-based feature extractor. (ii) the meta-test phase applies the frozen feature extractor to novel data (novel data has different categories from base data) and designs a classifier for recognition. To correct few-shot data distribution, researchers propose Semi-Supervised Few-Shot Learning (SSFSL) by introducing unlabeled data. Although SSFSL has been proved to achieve outstanding performances in the FSL community, there still exists a fundamental problem: the pre-trained feature extractor cannot adapt to the novel data flawlessly due to the cross-category setting. Usually, large amounts of noises are introduced to the novel feature. We dub it as Feature-Extractor-Maladaptive (FEM) problem. To tackle FEM, we make two efforts in this paper. First, we propose a novel label prediction method, Isolated Graph Learning (IGL). IGL introduces the Laplacian operator to encode the raw data to graph space, which helps reduce the dependence on features when classifying, and then project graph representation to label space for prediction. The key point is that: IGL can weaken the negative influence of noise from the feature representation perspective, and is also flexible to independently complete training and testing procedures, which is suitable for SSFSL. Second, we propose Graph Co-Training (GCT) to tackle this challenge from a multi-modal fusion perspective by extending the proposed IGL to the co-training framework. GCT is a semi-supervised method that exploits the unlabeled samples with two modal features to crossly strengthen the IGL classifier. We estimate our method on five benchmark few-shot learning datasets and achieve outstanding performances compared with other state-of-the-art methods. It demonstrates the effectiveness of our GCT. Rui Xu 0012, Lei Xing 0005, Shuai Shao 0006, Lifei Zhao, Baodi Liu, Weifeng Liu 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Accurately Modeling the Resting Brain Functional Correlations Using Wave Equation With Spatiotemporal Varying Hypergraph LaplacianabstractHow spontaneous brain neural activities emerge from the underlying anatomical architecture, characterized by structural connectivity (SC), has puzzled researchers for a long time. Over the past decades, much effort has been directed toward the graph modeling of SC, in which the brain SC is generally considered as relatively invariant. However, the graph representation of SC is unable to directly describe the connections between anatomically unconnected brain regions and fail to model the negative functional correlations. Here, we extend the static graph model to a spatiotemporal varying hypergraph Laplacian diffusion (STV-HGLD) model to describe the propagation of the spontaneous neural activity in human brain by incorporating the Laplacian of the hypergraph representation of the structural connectome ( h SC) into the regular wave equation. Theoretical solution shows that the dynamic functional couplings between brain regions fluctuate in the form of an exponential wave regulated by the spatiotemporal varying Laplacian of h SC. Empirical study suggests that the cortical wave might give rise to resonance with SC during the self-organizing interplay between excitation and inhibition among brain regions, which orchestrates the cortical waves propagating with harmonics emanating from the h SC while being bound by the natural frequencies of SC. Besides, the average statistical dependencies between brain regions, normally defined as the functional connectivity (FC), arises just at the moment before the cortical wave reaches the steady state after the wave spreads across all the brain regions. Comprehensive tests on four extensively studied empirical brain connectome datasets with different resolutions confirm our theory and findings. Yanjiang Wang 0001, Jichao Ma, Baodi Liu |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Leveraging GANs via Non-local Features
Xuyang Peng, Weifeng Liu 0001, Baodi Liu, Kai Zhang 0029, Yicong Zhou |
ICANN (2) | 3 |
| 2021 | Adaptive Multi-Feature Fusion for Robust Object TrackingabstractIn this paper, in order to better describe the object, an adaptive multi-feature fusion method is proposed, which makes full use of the advantages of various features. Firstly, hierarchical convolution features and two hand-crafted features are fused linearly, and the weights of different features are adjusted adaptively to obtain the optimal object representation in the tracking process. Secondly, a translation filter and a scale filter are adopted to estimate the object’s exact position and scale, respectively. Finally, in the model update stage, an efficient adaptive model update strategy is used to improve the performance, which can significantly alleviate the model noises. Extensive experimental results on well-known benchmark datasets show that the proposed algorithm performs favorably against the state-of-the-art tracking methods. Yujuan Qi, Yanjiang Wang 0001, Baodi Liu |
ICIP | 4 |
| 2021 | Linked Attention-Based Dynamic Graph Convolution Module for Point Cloud ClassificationabstractWith the rapid development of 3D technology, point cloud data is becoming more and more popular, which arouses researchers’ interest. But its properties – irregularity and disorder – make it difficult to analyze. In this work, we combine the attention module with the dynamic graph convolutional neural network to pay attention to the target’s critical part. Then, the modules are densely connected to guarantee that each layer is fully utilized. Finally, we carry out experiments on several benchmark datasets to verify the proposed model and achieve state-of-the-art performance. Baodi Liu, Weifeng Liu 0001, Kai Zhang 0029 |
ICIP | 2 |
| 2021 | OPS-Net: Over-Parameterized Sharing Networks for Video Frame InterpolationabstractThe video frame interpolation algorithm can improve temporal resolution by inserting non-existent frames in the video sequence. With the help of skip connections, many kernel-based methods train deep neural networks to accurately establish the complicated spatiotemporal relationship among pixels in adjacent frames. Still, these connections are only performed in the feature dimension. To this end, we introduce the Over-Parameterized Sharing Networks (OPS-Net) to implement weight sharing under different layers, capable of integrating deep and shallow features more directly. Specifically, we over-parameterize each convolutional layer to capture movement information efficiently, where the additional trainable weights from distinct ones will be shared. After the training, the additional weights will be fused into the conventional convolutional layer and do not increase the test phase’s computation. Experimental results show that the proposed method can generate favorable frames compared with several state-of-the-art approaches. Zhenfang Wang, Yanjiang Wang 0001, Shuai Shao 0006, Baodi Liu |
ICIP | 4 |
| 2021 | Affine Non-Negative Collaborative Representation for Deep Metric Learning
Baodi Liu, Weifeng Liu 0001, Kai Zhang 0029 |
ICIP | 2 |
| 2021 | SSDL: Self-Supervised Dictionary LearningabstractThe label-embedded dictionary learning (DL) algorithms generate influential dictionaries by introducing discriminative information. However, there exists a limitation: All the label-embedded DL methods rely on the labels due that this way merely achieves ideal performances in supervised learning. While in semi-supervised and unsupervised learning, it is no longer sufficient to be effective. Inspired by the concept of self-supervised learning (e.g., setting the pretext task to generate a universal model for the downstream task), we propose a Self-Supervised Dictionary Learning (SSDL) framework to address this challenge. Specifically, we first design a p-Laplacian Attention Hypergraph Learning (pAHL) block as the pretext task to generate pseudo soft labels for DL. Then, we adopt the pseudo labels to train a dictionary from a primary label-embedded DL method. We evaluate our SSDL on two human activity recognition datasets. The comparison results with other state-of-the-art methods have demonstrated the efficiency of SSDL. Shuai Shao 0006, Lei Xing 0005, Wei Yu 0004, Rui Xu 0012, Yanjiang Wang 0001, Baodi Liu |
ICME | 6 |
| 2021 | Collaborative Representation for Deep Meta Metric LearningabstractMost metric learning methods utilize all training data to construct a single metric, and it is usually over-fitting on the "salient" feature. To overcome this issue, we propose a deep meta metric learning method based on collaborative representation. We construct multiple episodes from the original training data to train a general metric, where each episode consists of a query set and a support set. Then, we introduce a collaborative representation method, which fits the query sample with the support samples per class. We predict the query sample's label via the optimal fitness among the query sample and the support samples in each specific class. Besides, we adopt a hard mining strategy to learn a more discriminative metric according to increasing the training tasks' difficulty. Experiments verify that our method achieves state-of-the-art results on three re-ID benchmark datasets. Weifeng Liu 0001, Kai Zhang 0029, Baodi Liu |
ICMR | 6 |
| 2021 | MHFC: Multi-Head Feature Collaboration for Few-Shot LearningabstractFew-shot learning (FSL) aims to address the data-scarce problem. A standard FSL framework is composed of two components: (1) Pre-train. Employ the base data to generate a CNN-based feature extraction model (FEM). (2) Meta-test. Apply the trained FEM to acquire the novel data's features and recognize them. FSL relies heavily on the design of the FEM. However, various FEMs have distinct emphases. For example, several may focus more attention on the contour information, whereas others may lay particular emphasis on the texture information. The single-head feature is only a one-sided representation of the sample. Besides the negative influence of cross-domain (e.g., the trained FEM can not adapt to the novel class flawlessly), the distribution of novel data may have a certain degree of deviation compared with the ground truth distribution, which is dubbed as distribution-shift-problem (DSP). To address the DSP, we propose Multi-Head Feature Collaboration (MHFC) algorithm, which attempts to project the multi-head features (e.g., multiple features extracted from a variety of FEMs) to a unified space and fuse them to capture more discriminative information. Typically, first, we introduce a subspace learning method to transform the multi-head features to aligned low-dimensional representations. It corrects the DSP via learning the feature with more powerful discrimination and overcomes the problem of inconsistent measurement scales from different head features. Then, we design an attention block to update combination weights for each head feature automatically. It comprehensively considers the contribution of various perspectives and further improves the discrimination of features. We evaluate the proposed method on five benchmark datasets (including cross-domain experiments) and achieve significant improvements of 2.1%-7.8% compared with state-of-the-arts. Shuai Shao 0006, Lei Xing 0005, Yan Wang 0076, Rui Xu 0012, Yanjiang Wang 0001, Baodi Liu |
ACM Multimedia | 7 |
| 2021 | LDAnet: a discriminant subspace for metric-based few-shot learningabstractDeep neural networks have surpassed humans in some cases, such as image recognition and image classification, with numerous labeled training samples. However, multiple tasks cannot provide enough labeled samples, training a neural network with a limited number of labeled samples is challenging. Therefore, meta-learning emerges. It aims to summarize from few-shot tasks and can quickly adapt to new categories. However, it is challenging to ensure that the trained feature extractor in the meta-train can adapt to the novel class in the meta-test. In this paper, we propose a subspace learning module to deal with the feature-mismatch problem. Specifically, we embed the linear discriminant analysis (LDA) module into the few-shot learning framework. It guarantees the feature embeddings in each few-shot task are more discriminative via increasing the inter-class distance and reducing the intra-class variances. We conduct experiments on four benchmark few-shot learning datasets, namely mini-Imagenet, CIFAR-FS, tiered-ImageNet, and CUB, to demonstrate the effectiveness of the proposed module. Experimental results show that this method has better performance than the state-of-the-art approaches. Dalei Chen, Baodi Liu |
SMC | 2 |
| 2021 | Adaptive Eigenmodes for Robust Object TrackingabstractDiscriminative correlation filters based algorithms have attracted extensive attention due to their strong tracking capability. However, object tracking still faces many challenges due to object appearance variations, background clutter, occlusion, plane rotation, etc. In this paper, to better express the object, multiple features are integrated to make full use of the advantage of different features. Furthermore, the adaptive eigen-decomposition and reconstruction ("eigenmodes") method is applied to carry out the integrated-feature decomposition, and optimal expression of the object is established through simple eigen-relationships. It has proved experimentally that the predicted value after eigenmodes is closer to the groundtruth than before. Furthermore, to solve the tracking failure caused by interfering objects or background clutters and improve the tracking accuracy, the average peak-correlation energy (APCE) method is utilized as an optimized update strategy in this paper. A large number of experimental results on the known reference datasets indicate that our algorithm has good performance compared to the existing tracking methods. Yujuan Qi, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001 |
SMC | 4 |
| 2021 | From Objects to a Whole PaintingabstractStyle image painting is the process of using some stylized strokes to redraw a reference image purposefully and meaningfully. It is a kind of style image generation. In recent years, the application of GAN has greatly improved the quality of generated images for style image generation. However, those methods which use GAN are usually non-serialized. To solve this problem, reinforcement learning and RNN methods are applied to the generation of style images based on strokes. To speed up training, stroke-based style image painting using CNN framework is also proposed. But none of the existing style image painting methods takes into account the content distribution in the reference image, which makes the painting process lack a clear order. We extract the feature of the object contained in the image through the content acquisition module and use the optimal transmission theory to construct a compound loss function to integrate the content information into the painting process. As more advanced macro content information is added to the painting process, our painting method can draw images in a more orderly way. We have conducted a lot of experiments to prove that our method is superior to the current state-of-the-art style image painting method. Fei Wang 0032, Baodi Liu, Weifeng Liu 0001 |
SMC | 2 |
| 2021 | CNN-combined graph residual network with multilevel feature fusion for hyperspectral image classificationabstractAbstract The application of graph convolutional networks (GCN) in hyperspectral image (HSI) classification has become a promising method, thanks to its flexible convolution operation in any irregular image region. For the classification of HSI, GCN can extract more superpixel‐level features with a topological structure, in comparison to the traditional convolutional neural networks (CNNs) using fixed square kernels distilling pixel‐level features. To fully leverage the different levels of features, this study proposes a novel deep network referred to as a CNN‐combined graph residual network (GRN), which integrates the multilevel graph residual module and spectral‐spatial features continuous learning module. During the extraction of topology information using the former module, HSI pixels are divided into superpixels and served as input nodes of the module to reduce the computational complexity and obtain the multilevel spatial relevance between adjacent superpixels. Besides, for the latter module, the spectral‐spatial features are learnt continuously, which could obtain the finer pixel‐level features. Finally, the captured spectral‐spatial features of different levels are concatenated. This strategy could not only adequately utilize the correlation and difference of adjacent spatial but also obtain the finer and more valuable spectral‐spatial information, which makes a significant boost in the HSI classification. Additionally, the experiment results demonstrate the superiority and availability of the GRN on three benchmark datasets of HSI, compared with the state‐of‐the‐art methods for the classification of HSI. Wenhui Guo, Guixun Xu, Weifeng Liu 0001, Baodi Liu, Yanjiang Wang 0001 |
IET Comput. Vis. | 4 |
| 2021 | Learn from Object Counting: Crowd Counting with Meta-learningabstractAbstract The objective of crowd counting is to learn a counter that can estimate the number of people in a single image. So far, most of the proposed work evaluates the crowd density by fitting the constructed density map corresponding to the sample. The performance of those algorithms depends on a large amount of carefully prepared data. However, a significant problem with crowd data sets is the difficulty of labeling. To address such a situation, utilizing object counting data in few‐shot scenes is considered and an efficient algorithm to extract the meta‐information is proposed, thus improving the accuracy and convergence rate of the crowd counting tasks. Specifically, the counting network is trained with only object counting tasks constructed on different domains during the meta‐training phase. Then, the meta‐counter is testing on crowd counting tasks in the meta‐testing stage. Experimentally, it is demonstrated that the above way improves the converge rate and accuracy of crowd counting tasks on three crowd counting datasets when meta‐training on ten‐type object counting tasks. Changtong Zan, Baodi Liu, Weili Guan, Kai Zhang 0029, Weifeng Liu 0001 |
IET Image Process. | 2 |
| 2021 | Accurately modeling the human brain functional correlations with hypergraph Laplacian
Jichao Ma, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001 |
Neurocomputing | 3 |
| 2021 | Classification of Remotely Sensed Images Using an Ensemble of Improved Convolutional NetworkabstractIn the last few years, the deep learning methods, especially the residual neural network, have achieved impressive performance in remote sensing image recognition tasks. However, there are still specific problems that need to be addressed. It is well known that the first several layers of the network provide much discriminative information, and the ResNet reduces the size of the feature map so quickly that it failed to fully learn the information beneficial to classification in the early stage. Second, insufficient labeling data in remote sensing database may easily lead to overfitting and affect the final classification accuracy. Third, the optimal results cannot be achieved by relying solely on transfer learning. To overcome the problems mentioned earlier, we propose an enhanced residual neural network (ERNet) to improve the classification performance on remote sensing images. We moderately broadened the first several layers of the network, changed the size of the convolution filters, and made it learn more information of image features. Second, we add dropout layer to each residual unit of the proposed network to improve the accuracy and generalization power of ERNet. Finally, an ensemble of learning methods based on ERNet was introduced to improve the classification performance by fusing features of other baseline methods. Extensive experimental results on several benchmark data sets of remote sensing images demonstrate the superior performance of our proposed algorithm. Li Wang 0040, Yanjiang Wang 0001, Yaqian Zhao, Baodi Liu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Unified Cross-domain Classification via Geometric and Statistical Adaptations
Weifeng Liu 0001, Baodi Liu, Weili Guan, Yicong Zhou, Changsheng Xu |
Pattern Recognit. | 3 |
| 2020 | Local structure alignment guided domain adaptation with few source samplesabstractDomain adaptation has received lots of attention for its high efficiency in dealing with cross-domain learning tasks. Most existing domain adaptation methods adopt the strategies relying on large amounts of source label information, which limits their applications in the real world where only a few label samples are available. We exploit the local geometric connections to tackle this problem and propose a Local Structure Alignment (LSA) guided domain adaptation method in this paper. LSA leverages the Nyström method to describe the distribution difference from the geometric perspective and then perform the distribution alignment between domains. Specifically, LSA constructs a domain-invariant Hessian matrix to locally connect the data of the two domains through minimizing the Nyström approximation error. And then it integrates the domain-invariant Hessian matrix with the semi-supervised learning and finally builds an adaptive semi-supervised model. Extensive experimental results validate that the proposed LSA outperforms the traditional domain adaptation methods especially when only sparse source label information is available. Yuying Cai, Baodi Liu, Weifeng Liu 0001, Kai Zhang 0029, Changsheng Xu |
MMAsia | 3 |
| 2020 | Label embedded dictionary learning for image classification
Shuai Shao 0006, Rui Xu 0012, Weifeng Liu 0001, Baodi Liu, Yanjiang Wang 0001 |
Neurocomputing | 4 |
| 2020 | Class specific or shared? A cascaded dictionary learning framework for image classification
Yanjiang Wang 0001, Shuai Shao 0006, Rui Xu 0012, Weifeng Liu 0001, Baodi Liu |
Signal Process. | 5 |
| 2019 | Laplacian Eigenmaps Regularized Feature Mapping for Image AnnotationabstractIn the past two decades, researchers have shown great interest in automatic image annotation. However, most existing research methods do not consider similarities among samples or do not obtain the suitable manifold information. Those methods also require the adequate and precise label sets. Considering above mentioned challenges, we propose a method, called laplacian eigenmaps regularized feature mapping for image annotation, which construct a laplacian matrix with all data in the training set (include labeled data and unlabeled data) and embed the laplacian matrix into feature mapping. Experimental results conducted on several benchmark image annotation datasets, such as Corel5K and ESP Game, demon-strate the effectiveness of the proposed method. Qianqian Shao, Baodi Liu |
SMC | 2 |
| 2018 | Biological modeling of human visual system for object recognition using GLoP filters and sparse coding on multi-manifolds
Limiao Deng, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001, Yujuan Qi |
Mach. Vis. Appl. | 3 |
| 2018 | Ranking-Preserving Low-Rank Factorization for Image Annotation With Missing LabelsabstractAutomatic image annotation has been extensively studied in the recent decades. Nevertheless, existing methods usually assume a properly labeled training set, which greatly inhibits their application to real-world datasets with incomplete labels. Due to their lack of special treatments for noisy data, most existing methods simply consider the missing labels as strictly negative ones, leading to the degradation of tagging accuracy. In light of such challenges, we propose a novel model in this paper, called ranking-preserving low-rank factorization. Specifically, we construct a local training set for each test image, and conduct low-rank matrix factorization on the model coefficient matrix, to simultaneously capture the label dependency and reduce the model complexity. Furthermore, to alleviate the ambiguity introduced by missing labels, the prediction model is learnt via tag ranking regularized by sample similarities and tag correlations, and both regularization terms are incorporated into our factorization scheme. By assembling all the aforementioned components together, our method obviates the need for making binary decisions based on unreliable data, and thus is more robust towards missing labels. Extensive empirical evaluations conducted on four datasets demonstrate the effectiveness of the proposed method. Xue Li 0005, Bin Shen 0002, Baodi Liu, Yu-Jin Zhang |
IEEE Trans. Multim. | 3 |
| 2017 | Class specific centralized dictionary learning for face recognition
Baodi Liu, Liangke Gui, Yu-Xiong Wang, Bin Shen 0002, Xue Li 0005, Yanjiang Wang 0001 |
Multim. Tools Appl. | 1 |
| 2016 | Saliency-context two-stream convnets for action recognitionabstractRecently, very deep two-stream ConvNets have achieved great discriminative power for video classification, which is especially the case for the temporal ConvNets when trained on multi-frame optical flow. However, action recognition in videos often fall prey to the wild camera motion, which poses challenges on the extraction of reliable optical flow for human body. In light of this, we propose a novel method to remove the global camera motion, which explicitly calculates a homography between two consecutive frames without human detection. Given the estimated homography due to camera motion, background motion can be canceled out from the warped optical flow. We take this a step further and design a new architecture called Saliency-Context two-stream ConvNets, where the context two-stream ConvNets are employed to recognize the entire scene in video frames, whilst the saliency streams are trained on salient human motion regions that are detected from the warped optical flow. Finally, the Saliency-Context two-stream ConvNets allow us to capture complementary information and achieve state-of-the-art performance on UCF101 dataset. Quan-Qi Chen, Xue Li 0005, Baodi Liu, Yu-Jin Zhang |
ICIP | 4 |
| 2016 | Class specific dictionary learning based kernel collaborative representation for fine-grained image classificationabstractRecently, dictionary learning based sparse representation algorithm has been widely adopted and achieved satisfying performance in image classification. However, sparse representation based classification (SRC) as well as collaborative representation based classification (CRC) always result in high residual error due to their basic assumption that considers training samples as dictionary directly for each category. And conventional class specific dictionary learning algorithm usually operates in the Euclidean space and fails to capture nonlinear information. To deal with these problems, we propose a classification algorithm which is called class specific dictionary learning based kernel collaborative representation (CSDL-KCRC) to enhance the classification accuracy. Extensive experimental results operated on three fine-grained image datasets, such as Oxford 102-Flowers dataset, Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset and Stanford Dogs dataset, demonstrate the effectiveness of CSDL-KCRC in image classification. Xiaojie Feng, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001 |
SMC | 3 |
| 2016 | Low-rank image tag completion with dual reconstruction structure preserved
Xue Li 0005, Yu-Jin Zhang, Bin Shen 0002, Baodi Liu |
Neurocomputing | 4 |
| 2016 | Face recognition using class specific dictionary learning for sparse representation and collaborative representation
Baodi Liu, Bin Shen 0002, Liangke Gui, Yu-Xiong Wang, Xue Li 0005, Yanjiang Wang 0001 |
Neurocomputing | 1 |
| 2016 | Blockwise coordinate descent schemes for efficient and effective dictionary learning
Baodi Liu, Yu-Xiong Wang, Bin Shen 0002, Xue Li 0005, Yu-Jin Zhang, Yanjiang Wang 0001 |
Neurocomputing | 1 |
| 2016 | Elastic net regularized dictionary learning for image classification
Bin Shen 0002, Baodi Liu, Qifan Wang 0001 |
Multim. Tools Appl. | 2 |
| 2016 | A Locality Sensitive Low-Rank Model for Image Tag CompletionabstractMany visual applications have benefited from the outburst of web images, yet the imprecise and incomplete tags arbitrarily provided by users, as the thorn of the rose, may hamper the performance of retrieval or indexing systems relying on such data. In this paper, we propose a novel locality sensitive low-rank model for image tag completion, which approximates the global nonlinear model with a collection of local linear models. To effectively infuse the idea of locality sensitivity, a simple and effective pre-processing module is designed to learn suitable representation for data partition, and a global consensus regularizer is introduced to mitigate the risk of overfitting. Meanwhile, low-rank matrix factorization is employed as local models, where the local geometry structures are preserved for the low-dimensional representation of both tags and samples. Extensive empirical evaluations conducted on three datasets demonstrate the effectiveness and efficiency of the proposed method, where our method outperforms pervious ones by a large margin. Xue Li 0005, Bin Shen 0002, Baodi Liu, Yu-Jin Zhang |
IEEE Trans. Multim. | 3 |
| 2015 | SP-SVM: Large Margin Classifier for Data on Multiple ManifoldsabstractAs one of the most important state-of-the-art classification techniques, Support Vector Machine (SVM) has been widely adopted in many real-world applications, such as object detection, face recognition, text categorization, etc., due to its competitive practical performance and elegant theoretical interpretation. However, it treats all samples independently, and ignores the fact that, in many real situations especially when data are in high dimensional space, samples typically lie on low dimensional manifolds of the feature space and thus a sample can be related to its neighbors by being represented as a linear combination of other samples on the same manifold. This linear representation, which is usually sparse, reflects the structure of underlying manifolds. It has been extensively explored in the recent literature and proven to be critical for the performance of classification. To benefit from both the underlying low dimensional manifold structure and the large margin classifier, this paper proposes a novel method called Sparsity Preserving Support Vector Machine(SP-SVM), which explicitly considers the sparse representation of samples while maximizing the margin between different classes. Consequently, SP-SVM inherits both the discriminative power of support vector machine and the merits of sparsity. A set of experiments on real-world benchmark data sets show that SP-SVM achieves significantly higher precision on recognition task than various competitive baselines including the traditional SVM, the sparse representation based method and the classical nearest neighbor classifier. Bin Shen 0002, Baodi Liu, Qifan Wang 0001, Yi Fang 0008, Jan P. Allebach |
AAAI | 2 |
| 2015 | A Locality Preserving Approach for Kernel PCA
Bin Shen 0002, Baodi Liu, Yu-Jin Zhang |
ICIG (1) | 5 |
| 2015 | Locality sensitive dictionary learning for image classificationabstractIn this paper, motivated by the superior performance of sparse representation based dictionary learning for application of image classification and the usage of nonlinearity property in improving performance of image representation, we propose a locality sensitive dictionary learning algorithm with global consistency and smoothness constraint to overcome the restriction of linearity at relatively low cost. Specifically, the image features are partitioned into several groups in a locality sensitive way and a global consistency regularizer is embedded into locality sensitive dictionary learning algorithm. The proposed algorithm is efficient to capture complex nonlinear structure. Experimental results on several benchmark data sets demonstrate the efficiency of our proposed locality sensitive dictionary learning algorithm. Baodi Liu, Bin Shen 0002, Xue Li 0005 |
ICIP | 1 |
| 2014 | Self-explanatory Sparse Representation for Image Classification
Baodi Liu, Yu-Xiong Wang, Bin Shen 0002, Yu-Jin Zhang, Martial Hebert |
ECCV (2) | 1 |
| 2014 | Blockwise coordinate descent schemes for sparse representationabstractThe current sparse representation framework is to decouple it as two subproblems, i.e., alternate sparse coding and dictionary learning using different optimizers, treating elements in bases and codes separately. In this paper, we treat elements both in bases and codes ho-mogenously. The original optimization is directly decoupled as several blockwise alternate subproblems rather than above two. Hence, sparse coding and bases learning optimizations are coupled together. And the variables involved in the optimization problems are partitioned into several suitable blocks with convexity preserved, making it possible to perform an exact block coordinate descent. For each separable subproblem, based on the convexity and monotonic property of the parabolic function, a closed-form solution is obtained. Thus the algorithm is simple, efficient and effective. Experimental results show that our algorithm significantly accelerates the learning process. Baodi Liu, Yu-Xiong Wang, Bin Shen 0002, Yu-Jin Zhang, Yanjiang Wang 0001 |
ICASSP | 1 |
| 2014 | Image tag completion by low-rank factorization with dual reconstruction structure preservedabstractA novel tag completion algorithm is proposed in this paper, which is designed with the following features: 1) Low-rank and error s-parsity: the incomplete initial tagging matrix D is decomposed into the complete tagging matrix A and a sparse error matrix E. However, instead of minimizing its nuclear norm, A is further factorized into a basis matrix U and a sparse coefficient matrix V, i.e. D = UV + E. This low-rank formulation encapsulating sparse coding enables our algorithm to recover latent structures from noisy initial data and avoid performing too much denoising; 2) Local reconstruction structure consistency: to steer the completion of D, the local linear reconstruction structures in feature space and tag space are obtained and preserved by U and V respectively. Such a scheme could alleviate the negative effect of distances measured by low-level features and incomplete tags. Thus, we can seek a balance between exploiting as much information and not being mislead to suboptimal performance. Experiments conducted on Corel5k dataset and the newly issued Flickr30Concepts dataset demonstrate the effectiveness and efficiency of the proposed method. Xue Li 0005, Yu-Jin Zhang, Bin Shen 0002, Baodi Liu |
ICIP | 4 |
| 2014 | TISVM: Large margin classifier for misaligned image classificationabstractSupport vector machine is one of the most successful machine learning methods in image processing and computer vision in the past decades. However, its performance strongly depends on the training data, which are sometimes expensive and of low quality. Specifically, in many real applications, such as face recognition, the images are rarely perfectly aligned, thus the misalignment between training and testing data impairs the performance. In this paper, we propose a strategy to compensate the misalignment between images while learning the classifier without looking at the testing samples. Specifically, some certain critical transformations are inferred and applied to training samples to alleviate the effect of the worst case of possible misalignment. The resulted large margin classifier generalizes better than traditional SVM, especially when there is misalignment. Experimental results on real image data sets show the efficacy of the proposed algorithm. Bin Shen 0002, Baodi Liu, Jan P. Allebach |
ICIP | 2 |
| 2014 | Robust nonnegative matrix factorization via L1 norm regularization by multiplicative updating rulesabstractNonnegative Matrix Factorization (NMF) is a widely used technique in many applications such as face recognition, motion segmentation, etc. It approximates the nonnegative data in an original high dimensional space with a linear representation in a low dimensional space by using the product of two nonnegative matrices. In many applications data are often partially corrupted with large additive noise. When the positions of noise are known, some existing variants of N-MF can be applied by treating these corrupted entries as missing values. However, the positions are often unknown in many real world applications, which prevents the usage of traditional NMF or other existing variants of NMF. This paper proposes a Robust Nonnegative Matrix Factorization (RobustNMF) algorithm that explicitly models the partial corruption as large additive noise without requiring the information of positions of noise. In particular, the proposed method jointly approximates the clean data matrix with the product of two nonnegative matrices and estimates the positions and values of outliers/noise. An efficient iterative optimization algorithm with a solid theoretical justification has been proposed to learn the desired matrix factorization. Experimental results demonstrate the advantages of the proposed algorithm. Bin Shen 0002, Baodi Liu, Qifan Wang 0001, Rongrong Ji |
ICIP | 2 |
| 2014 | Class specific subspace learning for collaborative representationabstractCollaborative representation based classification (CRC) has been successfully used for visual recognition and showed impressive performance recently. However, it directly uses the training samples from each class as the subspaces to calculate the minimum residual error for a given testing sample. This leads to high residual error and instability, which is critical especially for a small number of training samples in each class. In this paper, we propose a class specific subspace learning algorithm for collaborative representation. By introducing the dual form of subspace learning, it presents an explicit relationship between the basis vectors and the original image features, and thus enhances the interpretability. Lagrange multipliers are then applied to optimize the corresponding objective function, i.e., learning the weights used in constructing the subspaces. Extensive experimental results demonstrate that the proposed algorithm has achieved superior performance in several visual recognition tasks. Baodi Liu, Bin Shen 0002, Yu-Xiong Wang, Weifeng Liu 0001, Yanjiang Wang 0001 |
SMC | 1 |
| 2013 | Self-Explanatory Convex Sparse Representation for Image ClassificationabstractSparse representation technique has been widely used in various areas of computer vision over the last decades. Unfortunately, in the current formulations, there are no explicit relationship between the learned dictionary and the original data. By tracing back and connecting sparse representation with the K-means algorithm, a novel variation scheme termed as self-explanatory convex sparse representation (SCSR) has been proposed in this paper. To be specific, the basis vectors of the dictionary are refined as convex combination of the data points. The atoms now would capture a notion of centroids similar to K-means, leading to enhanced interpretability. Sparse representation and K-means are thus unified under the same framework in this sense. Besides, an appealing property also emerges that the weight and code matrices both tend to be naturally sparse without additional constraints. Compared with the standard formulations, SCSR is easier to be extended into the kernel space. To solve the corresponding sparse coding sub problem and dictionary learning sub problem, block-wise coordinate descent and Lagrange multipliers are proposed accordingly. To validate the proposed algorithm, it is implemented in image classification, a successful applications of sparse representation. Experimental results on several benchmark data sets, such as UIUC-Sports, Scene 15, and Caltech-256 demonstrate the effectiveness of our proposed algorithm. Baodi Liu, Yu-Xiong Wang, Bin Shen 0002, Yu-Jin Zhang, Yanjiang Wang 0001, Weifeng Liu 0001 |
SMC | 1 |
| 2013 | Learning dictionary on manifolds for image classification
Baodi Liu, Yu-Xiong Wang, Yu-Jin Zhang, Bin Shen 0002 |
Pattern Recognit. | 1 |
| 2012 | Discriminant sparse coding for image classificationabstractRecently, dictionary learned by sparse coding has been widely adopted in image classification and has achieved competitive performance. Sparse coding is capable of reducing the reconstruction error in transforming low-level descriptors into compact mid-level features. Nevertheless, dictionary learned by sparse coding does not have the ability to distinguish different classes. That is to say, it is not the optimum dictionary for the classification task. In this paper, based on the global image statistics, a novel discriminant dictionary learning method combining linear discriminant analysis with sparse coding is proposed to obtain a more discriminative dictionary while preserving its descriptive abilities and a block coordinate descent algorithm is proposed to solve the optimization problem. Experimental results show that our algorithm has capabilities to learn dictionary with more discriminative power and achieves superior performance. Baodi Liu, Yu-Xiong Wang, Yu-Jin Zhang |
ICASSP | 1 |
| 2012 | Action recognition in still images using a combination of human pose and context informationabstractIn this work, a novel method is proposed for recognizing human actions in still images, which incorporates both pose and context information. Poselet-based action classifiers are learned using Poselet Activation Vector as features, which contain pose information for each action. And context-based action classifiers for each action are learned on contextual information, which is obtained by sparse coding on foreground and background. The confidences of an image belonging to each action are obtained through summing up the probability outputs of the poselet-based and the context-based classifiers. The contribution of this work is three folded. Firstly, sparse coding is adopted to find compact patterns of the original features. Secondly, a block coordinate descent algorithm is proposed for sparse coding, which can be performed very fast in practice. Thirdly, both pose and context information are taken into consideration for action recognition. The experimental results show the proposed method achieves the state-of-the-art performance on several benchmarks. Yu-Jin Zhang, Xue Li 0005, Baodi Liu |
ICIP | 4 |
| 2011 | Robust Moving Cast Shadows Detection and Removal with Reliability Checking of the Color PropertyabstractIn this paper, we propose a novel method that is capable of detecting and removing moving cast shadows robustly. Four properties of the moving cast shadows are observed. To make full use of the four properties, pixel level processing is adopted to extract the color, texture and luminance features of the moving foreground, and region level processing is introduced to detect shadows. Another innovation is that we also propose a method to check the reliability of the color feature that will be used only when it is reliable. Furthermore, by introducing the geometric center detection as a prejudge step, the efficiency of the method gets improved. The experimental results prove that the proposed method is effective and robust. Xue Li 0005, Yu-Jin Zhang, Baodi Liu |
ICIG | 3 |