Weifeng Liu 0001

dblp:23/1112-1 · DBLP profile ↗
← Back
194ranked-venue papers
21as first author
151since 2021 · last 2026
0000-0002-5388-9080ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 77 · 13 first-author · 59 since 2021Graphics, computer vision, multimedia, augmented reality and games · 73 · 9 first-author · 57 since 2021Applied, interdisciplinary, general and emerging computing · 40 · 1 first-author · 32 since 2021Databases, data management, data science and information retrieval · 12 · 9 since 2021Human-computer interaction and ubiquitous computing · 11 · 6 since 2021Computer networks · 7 · 7 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models
abstract
3D understanding has drawn significant attention recently, leveraging Vision-Language Models (VLMs) to enable multi-modal reasoning between point cloud and text data. Current 3D-VLMs directly embed the 3D point clouds into 3D tokens, following large 2D-VLMs with powerful reasoning capabilities. However, this framework has a great computational cost limiting its application, where we identify that the bottleneck lies in processing all 3D tokens in the Large Language Model (LLM) part. This raises the question: how can we reduce the computational overhead introduced by 3D tokens while preserving the integrity of their essential information? To address this question, we introduce Hierarchical Compensatory Compression (HCC-3D) to efficiently compress 3D tokens while maintaining critical detail retention. Specifically, we first propose a global structure compression (GSC), in which we design global queries to compress all 3D tokens into a few key tokens while keeping overall structural information. Then, to compensate for the information loss in GSC, we further propose an adaptive detail mining (ADM) module that selectively recompresses salient but under-attended features through complementary scoring. Extensive experiments demonstrate that HCC-3D not only achieves extreme compression ratios (approximately 98%) compared to previous 3D VLMs, but also achieves new state-of-the-art performance, showing the great improvements on both efficiency and performance.
Liheng Zhang, Bingfeng Zhang, Weifeng Liu 0001
AAAI5
2026 Enhancing the Transferability of Jailbreak Attacks on Large Language Models via Exploiting Reparameterization Invariance
abstract
Jailbreak attacks serve as a pivotal technique for evaluating the safety alignment of Large language models.Current token-level attacks have shown remarkable efficacy on open-source models by leveraging gradient-based optimization.However, these attacks suffer from poor cross-model transferability, severely limiting their utility on proprietary ones.To address this limitation, we propose Reparameterization Invariance Gradient-based Jailbreak (RIGJ), a natural gradient based framework designed to improve cross-model transferability.Unlike prior token-level methods whose optimization paths are constrained by model-specific Euclidean geometry, RIGJ defines update directions according to differences in output distributions rather than parameter-space distances.Since language models are trained to capture similar dependency structures of natural language, their output distributions share common geometry across architectures, yielding intrinsically model-agnostic optimization trajectories and substantially stronger jailbreak transferability.Extensive experiments demonstrate superior performance, increasing the cross-model Attack Success Rate and Average Harmfulness Score by 14.9% and 1.23, respectively.Our code is provided in https://github.com/nohuma/AISafety_ transfer_jailbreak_RIGJ_2026.
Xinghao Yang, Yongshun Gong, Wei Liu 0007, Bao-di Liu, Weifeng Liu 0001
ACL (1)6
2026 Improving zero-shot translation with the navigation ability-enhanced language tags
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yibin Lei, Yibing Zhan, Weifeng Liu 0001
Eng. Appl. Artif. Intell.6
2026 Universal Attack based Focal Enhancement for Bearing Fault Diagnosis
Puhua Jia, Xinghao Yang, Yongwei Tang, Baodi Liu, Wei Li 0032, Weifeng Liu 0001
Knowl. Based Syst.6
2026 Context-aware feature complementary screening network for mass segmentation in whole mammograms
Qingkun Guo, Luhao Sun, Chao Li 0075, Wenzong Jiang, Weifeng Liu 0001, Baodi Liu
Multim. Syst.6
2026 Dynamic frequency-band filtering domain generalization for mammogram classification
Shenxiao Li, Yunqi Huang, Wenzong Jiang, Chao Li 0075, Weifeng Liu 0001, Xiongbin Wang, Baodi Liu
Multim. Syst.6
2026 SCM: Semantic Segmentation with Dual-stream Semantic Synergy under Adverse Weather Conditions
Shuochen Tian, Jian Pang, Bingfeng Zhang, Weifeng Liu 0001
Multim. Syst.5
2026 Cross-difference-driven dual-stream contrast multi-view network for mammogram classification
Ruijia Tian, Chenteng Zhang, Wenzong Jiang, Chao Li 0075, Weifeng Liu 0001, Xiongbin Wang, Baodi Liu
Multim. Syst.6
2026 Diagnosis-driven hard sample generation: low-frequency attenuation supervised contrastive learning for mammogram classification
Changchao Wang, Wenzong Jiang, Chao Li 0075, Weifeng Liu 0001, Xiongbin Wang, Baodi Liu
Multim. Syst.5
2026 A pressure-conditioned generative adversarial network for efficient temperature field visualization in combustion simulations
Haoran Yu 0005, Baodi Liu, Weifeng Liu 0001
Multim. Syst.6
2026 Channel-guided dual-pooling multi-scale spatial attention network for mass segmentation in whole mammograms
Wenzong Jiang, Weifeng Liu 0001, Xiongbin Wang, Baodi Liu
Multim. Syst.4
2026 Causal-guided strength differential independence sample weighting for out-of-distribution generalization
Haoran Yu 0005, Weifeng Liu 0001, Yingjie Wang 0007, Baodi Liu, Dapeng Tao, Honglong Chen
Pattern Recognit.2
2026 Joint subgraph independence for graph out-of-distribution generalization
Weifeng Liu 0001, Baodi Liu, Dapeng Tao, Honglong Chen
Pattern Recognit.2
2026 Cross-Hierarchical Decoding With SAM for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised medical image segmentation (SSMIS) mainly leverages valuable information from unlabeled data to complement the limited labeled guidance. Incorporating SAM, with its excellent generalization capabilities, enhances the learning process from unlabeled data, as demonstrated by existing methods. However, most current SAM-based methods in SSMIS focus on unique prompt design, while the prompts generated for unlabeled data through pseudo-labels unavoidably introduce noise, limiting the following decoding process. In this paper, we propose a Cross-Hierarchical Decoding (CHD) process for SAM, which removes explicit prompts (e.g., point or box) and thus mitigates the influence of inaccurate pseudo labels. Specifically, our CHD is a two-stage decoder. The first stage uses the original decoder in SAM to generate probability masks, which are combined with a learnable mask interaction module in the second stage to achieve more fine-grained segmentation. Meanwhile, to remove the restriction that the original SAM can only segment foreground-background categories, we design a cross-class correlation module in CHD to capture class-wise interrelationships between different classes, thus achieving multi-class segmentation. Extensive experiments show that CHD achieves new state-of-the-art performance for SSMIS, significantly improving different baselines.
Hanyang Chi, Xuru Gao, Guixun Luo, Bingfeng Zhang, Weifeng Liu 0001
IEEE Trans. Circuits Syst. Video Technol.7
2026 CO3+: Improved Collaborative Consortium of Foundation Models for Open-World Few-Shot Learning
abstract
Open-World Few-Shot Learning (OFSL) is a critical research domain focused on accurately identifying target samples under conditions where data is scarce and labels are unreliable. This field is highly relevant to real-world scenarios, holding significant practical implications. Currently, the field has only a few solutions, primarily relying on conventional methods such as metric learning and feature aggregation. However, these methods often struggle in more complex scenarios. Recent breakthroughs in foundation models such as CLIP and DINO have demonstrated their strong representational capabilities, even in resource-limited environments. These advancements have led to a shift from “training model from scratch” towards “exploiting the extensive capabilities and expertise of these pre-trained foundation models for OFSL”. Inspired by this shift, we introduce the Improved Collaborative Consortium of Foundation Models (CO+3), an extension of CO3, first presented in AAAI 2024. CO+3significantly improves the accuracy of OFSL by integrating the strengths of four foundational models. It includes three decoupled blocks: (1) The Label Correction Block (LC-Block) rectifies unreliable labels, (2) the Data Augmentation Block (DA-Block) enriches the available data, and (3) the Text-guided Fusion Adapter (TeFu-Adapter) merges various features and reduces the impact of noisy labels through semantic constraints. We evaluate CO+3across eleven benchmark datasets, comparing it against recent state-of-the-art methods. Our thorough evaluations demonstrate that the proposed CO+3consistently surpasses existing methods by a substantial margin, particularly in high-noise scenarios.
Shuai Shao 0006, Rui Xu 0012, Bingfeng Zhang, Baodi Liu, Weifeng Liu 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.5
2026 MATE: A D2D-Enhanced Multi-Bitrate Video Caching Strategy for Cloud-Edge-Device Collaborative Networks
abstract
Edge caching alleviates backhaul pressure and enhances video service quality by deploying video content near user devices. However, the limited storage capacity of edge servers struggles to cope with the exponential growth of video data, challenging the delivery of high-quality video services. While both Device-to-Device (D2D) caching and multi-bitrate video technology are promising solutions to relieve the pressure on edge servers, existing research suffers from a key limitation: studies on multi-bitrate caching are predominantly focused on the edge layer, while D2D caching is often limited to single-bitrate scenarios. This isolation neglects the significant benefits of integrating D2D caching with multi-bitrate technology and fails to develop a cross-layer caching strategy for multi-bitrate videos. To address this limitation, we propose a D2D-enhanced Multi-bitrate video cAching straTEgy (MATE) for cloud-edge-device collaborative networks. We formulate a joint service latency and caching replacement cost optimization problem, which can be modeled as a mixed-integer programming problem. To overcome the coupling between caching strategies at the edge layer and device layer, we employ an alternating iterative optimization approach to decouple the original problem into two subproblems. We design an edge-device double-layer joint caching strategy, i.e., a device-layer caching strategy based on greedy algorithm and Lagrange multipliers, and an edge-layer caching strategy based on multi-agent twin delayed deep deterministic policy gradient algorithm. Extensive simulations are conducted to demonstrate the effectiveness of the proposed MATE.
Honglong Chen, Xinglong Fan, Zhichen Ni, Liantao Wu, Peng Sun 0003, Weifeng Liu 0001
IEEE Trans. Mob. Comput.7
2026 Lesion Asymmetry Screening Assisted Global Awareness Multi-View Network for Mammogram Classification
abstract
Mammography is a primary method for early screening, and developing deep learning-based computer-aided systems is of great significance. However, current deep learning models typically treat each image as an independent entity for diagnosis, rather than integrating images from multiple views to diagnose the patient. These methods do not fully consider and address the complex interactions between different views, resulting in poor diagnostic performance and interpretability. To address this issue, this paper proposes a novel end-to-end framework for breast cancer diagnosis: lesion asymmetry screening assisted global awareness multi-view network (LAS-GAM). More than just the most common image-level diagnostic model, LAS-GAM operates at the patient level, simulating the workflow of radiologists analyzing mammographic images. The framework processes the four views of a patient and revolves around two key modules: a global module and a lesion screening module. The global module simulates the comprehensive assessment by radiologists, integrating complementary information from the craniocaudal (CC) and mediolateral oblique (MLO) views of both breasts to generate global features that represent the patient's overall condition. The lesion screening module mimics the process of locating lesions by comparing symmetric regions in contralateral views, identifying potential lesion areas and extracting lesion-specific features using a lightweight model. By combining the global features and lesion-specific features, LAS-GAM simulates the diagnostic process, making patient-level predictions. Moreover, it is trained using only patient-level labels, significantly reducing data annotation costs. Experiments on the Digital Database for Screening Mammography (DDSM) and In-house datasets validate LAS-GAM, achieving AUCs of 0.817 and 0.894, respectively.
Xinchuan Liu, Luhao Sun, Chao Li 0075, Bowen Han 0001, Wenzong Jiang, Tianhao Yuan, Weifeng Liu 0001, Zhaoyun Liu, Baodi Liu
IEEE Trans. Medical Imaging7
2026 Unbiased Semantic Decoding With Vision Foundation Models for Few-Shot Segmentation
abstract
Few-shot segmentation (FSS) has garnered significant attention. Many recent approaches attempt to introduce the segment anything model (SAM) to handle this task. With the strong generalization ability and rich object-specific extraction ability of the SAM model, such a solution shows great potential in FSS. However, the decoding process of SAM highly relies on accurate and explicit prompts, making previous approaches mainly focus on extracting prompts from the support set, which is insufficient to activate the generalization ability of SAM, and this design is easy to result in a biased decoding process when adapting to the unknown classes. In this work, we propose an unbiased semantic decoding (USD) strategy integrated with SAM, which extracts target information from both the support and query set simultaneously to perform consistent predictions guided by the semantics of the contrastive language-image pretraining (CLIP) model. Specifically, to enhance the unbiased semantic discrimination of SAM, we design two feature enhancement strategies that leverage the semantic alignment capability of CLIP to enrich the original SAM features, mainly including a global supplement at the image level to provide a generalize category indicate with support image and a local guidance at the pixel level to provide a useful target location with query image. Besides, to generate target-focused prompt embeddings, a learnable visual-text target prompt generator (VTPG) is proposed by interacting target text embeddings and clip visual features. Without requiring retraining of the vision foundation models, the features with semantic discrimination draw attention to the target region through the guidance of prompt with rich target information. Experiments on both the PASCAL- $5^{i}$ and COCO- $20^{i}$ show that our proposed method outperforms the existing approaches by a clear margin and achieves new state-of-the-art performances.
Bingfeng Zhang, Jian Pang, Weifeng Liu 0001, Baodi Liu, Honglong Chen
IEEE Trans. Neural Networks Learn. Syst.4
2025 Modeling All Response Surfaces in One for Conditional Search Spaces
abstract
Bayesian Optimization (BO) is a sample-efficient black-box optimizer commonly used in search spaces where hyperparameters are independent. However, in many practical AutoML scenarios, there will be dependencies among hyperparameters, forming a conditional search space, which can be partitioned into structurally distinct subspaces. The structure and dimensionality of hyperparameter configurations vary across these subspaces, challenging the application of BO. Some previous BO works have proposed solutions to develop multiple Gaussian Process models in these subspaces. However, these approaches tend to be inefficient as they require a substantial number of observations to guarantee each GP's performance and cannot capture relationships between hyperparameters across different subspaces. To address these issues, this paper proposes a novel approach to model the response surfaces of all subspaces in one, which can model the relationships between hyperparameters elegantly via a self-attention mechanism. Concretely, we design a structure-aware hyperparameter embedding to preserve the structural information. Then, we introduce an attention-based deep feature extractor, capable of projecting configurations with different structures from various subspaces into a unified feature space, where the response surfaces can be formulated using a single standard Gaussian Process. The empirical results on a simulation function, various real-world tasks, and HPO-B benchmark demonstrate that our proposed approach improves the efficacy and efficiency of BO within conditional search spaces.
Wei Liu 0005, Chao Xue 0003, Yibing Zhan, Xiaoxing Wang, Weifeng Liu 0001, Dacheng Tao
AAAI6
2025 Excluding the Impossible for Open Vocabulary Semantic Segmentation
abstract
Open vocabulary semantic segmentation is a hot topic in research, focusing on segmenting and recognizing a diverse array of categories in varied environments, including those previously unknown, thereby holding significant practical value. Mainstream studies utilize the CLIP model for direct semantic segmentation (denoted as “forward methods”), which often struggles to represent underrepresented categories effectively. To address this issue, this paper introduces a novel approach Excluding the ImpossibLe Semantic Segmentation Network (ELSE-Net) based on reverse thinking. By excluding improbable categories, ELSE-Net narrows the selection range for forward methods, significantly reducing the risk of misclassification. In implementation, we initially draw on leading research to design the General Processing Block (GP-Block), which generates inclusion probabilities (the likelihood of belonging to a category) by using the CLIP model cooperated with a Mask Proposal Network (MPN). We then present the EXcluding the ImPossible Block (EXP-Block), which computes exclusion probabilities (the likelihood of not belonging to a category) through the CLIPN model and a custom-designed Reverse Retrieval Adapter (R2-Adapter). These exclusion probabilities are subsequently used to refine the inclusion probabilities, which are ultimately employed to annotate class-agnostic masks. Moreover, the core component of our EXP-Block is model-agnostic, enabling it to enhance the capabilities of existing frameworks. Experimental results from four benchmark datasets validate the effectiveness of ELSE-Net and underscore the seamless model-agnostic functionality of the EXP-Block.
Shiyuan Zhao, Baodi Liu, Weifeng Liu 0001, Shuai Shao 0006
AAAI4
2025 Disentangled Information Bottleneck for Adversarial Text Defense
abstract
Adversarial text defense is a significant strategy to protect modern NLP models from being attacked.Typical text defense methods usually enhance the model's robustness by model retraining or equipping it with a data preprocessing step, aiming to eliminate the non-robust features and preserve the robust ones.Although some efforts have been made to recognize the robust features, e.g., by the information bottleneck (IB) technique, how to fully disentangle the robust and non-robust representation remains a big challenge.To alleviate this problem, we propose a novel text defense method, named Disentangled Information Bottleneck (DisIB), with two major merits.Firstly, we separate the robust features and non-robust features with a disentangled two-line framework rather than the one-line compression network in IB.This prevents the loss of robust features caused by information compression and produces complete robust features.Secondly, we design a discriminator network to approximate the minimum mutual information of the two lines, which sufficiently disentangles robust and non-robust features.To validate the effectiveness of our DisIB, we conduct a total of 96 defense experiments on four datasets by defending four popular attack methods.Experimental results elaborate that our method significantly outperforms six baselines, with accuracy improvements ranging from 3.8% to 20.7%.
Yidan Xu, Xinghao Yang, Bao-di Liu, Weifeng Liu 0001
EMNLP5
2025 Directional Denoising Diffusion Model for Defect Reconstruction Using Alternating Current Field Measurement
Yuhong Lu, Wei Li 0032, Weifeng Liu 0001
ICIC (16)5
2025 Domain Generalization for Mammogram Classification by Suppressing Domain-Specific Features
Jiqun Chen, Luhao Sun, Wenzong Jiang, Weifeng Liu 0001, Chao Li 0075, Baodi Liu
MICCAI (7)4
2025 Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
abstract
Multimodal Large Language Models (MLLMs) have demonstrated exceptional performance across various objective multimodal perception tasks, yet their application to subjective, emotionally nuanced domains, such as psychological analysis, remains largely unexplored. In this paper, we introduce PICK, a multi-step framework designed for Psychoanalytical Image Comprehension through hierarchical analysis and Knowledge injection with MLLMs, specifically focusing on the House-Tree-Person (HTP) Test, a psychological assessment test. First, we decompose drawings containing multiple instances into semantically meaningful sub-drawings, constructing a hierarchical representation that captures spatial structure and content across three levels: single-object level, multi-object level, and whole level. Next, we analyze these sub-drawings at each level with a targeted focus, extracting psychological or emotional insights from their visual cues. We also introduce an HTP knowledge base and design a feature extraction module, trained with reinforcement learning, to generate a psychological profile for single-object level analysis. This profile captures both holistic stylistic features and dynamic object-specific features (such as those of the house, tree, or person), correlating them with psychological states. Finally, we integrate these multi-faceted information to produce a well-informed assessment that aligns with expert-level reasoning. Our approach bridges the gap between MLLMs and specialized expert domains, offering a structured and interpretable framework for understanding human mental states through visual expression. Experimental results demonstrate that the proposed PICK significantly enhances the capability of MLLMs in psychological analysis. It is further validated as a general framework through extensions to emotion understanding tasks. Codes are released at https://github.com/YanbeiJiang/PICK.
Xueqi Ma, Yanbei Jiang, Sarah M. Erfani, James Bailey 0001, Weifeng Liu 0001, Krista A. Ehinger, Jey Han Lau
ACM Multimedia5
2025 FFCBA: Feature-based Full-target Clean-label Backdoor Attacks
Yangxu Yin, Honglong Chen, Yudong Gao, Peng Sun 0003, Liantao Wu, Zhe Li 0026, Weifeng Liu 0001
ACM Multimedia7
2025 Noise-Robust Few-Shot Classification via Variational Adversarial Data Augmentation
abstract
Few-shot classification models trained with clean samples poorly classify samples from the real world with various scales of noise. To enhance the model for recognizing noisy samples, researchers usually utilize data augmentation or use noisy samples generated by adversarial training for model training. However, existing methods still have problems: (i) The effects of data augmentation on the robustness of the model are limited. (ii) The noise generated by adversarial training usually causes overfitting and reduces the generalization ability of the model, which is very significant for few-shot classification. (iii) Most existing methods cannot adaptively generate appropriate noise. Given the above three points, this paper proposes a noise-robust few-shot classification algorithm, VADA—Variational Adversarial Data Augmentation. Unlike existing methods, VADA utilizes a variational noise generator to generate an adaptive noise distribution according to different samples based on adversarial learning, and optimizes the generator by minimizing the expectation of the empirical risk. Applying VADA during training can make few-shot classification more robust against noisy data, while retaining generalization ability. In this paper, we utilize FEAT and ProtoNet as baseline models, and accuracy is verified on several common few-shot classification datasets, including MiniImageNet, TieredImageNet, and CUB. After training with VADA, the classification accuracy of the models increases for samples with various scales of noise.
Baodi Liu, Kai Zhang 0029, Honglong Chen, Dapeng Tao, Weifeng Liu 0001
Comput. Vis. Media6
2025 Code-switching finetuning: Bridging multilingual pretrained language models for enhanced cross-lingual performance
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yu Cao 0014, Weifeng Liu 0001
Eng. Appl. Artif. Intell.5
2025 Hypnos: A domain-specific large language model for anesthesiology
Zhonghai Wang, Yibing Zhan, Bohao Zhou, Chong Zhang 0013, Baosheng Yu, Liang Ding 0006, Weifeng Liu 0001
Neurocomputing12
2025 Building accurate translation-tailored large language models with language-aware instruction tuning
abstract
Large language models (LLMs) exhibit remarkable capabilities in various natural language processing tasks, such as machine translation. However, the large number of LLM parameters incurs significant costs during inference. Previous studies have attempted to train translation-tailored LLMs with moderately sized models by fine-tuning them on the translation data. Nevertheless, when performing translations in zero-shot directions that are absent from the fine-tuning data, the problem of ignoring instructions and thus producing translations in the wrong language (i.e., the off-target translation issue) remains unresolved. In this work, we design a two-stage fine-tuning algorithm to improve the instruction-following ability of translation-tailored LLMs, particularly for maintaining accurate translation directions. We first fine-tune LLMs on the translation data to elicit basic translation capabilities. At the second stage, we construct instruction-conflicting samples by randomly replacing the instructions with the incorrect ones. Then, we introduce an extra unlikelihood loss to reduce the probability assigned to those samples. Experiments on two benchmarks using the LLaMA 2 and LLaMA 3 models, spanning 16 zero-shot directions, demonstrate that, compared to the competitive baseline—translation-finetuned LLaMA, our method could effectively reduce the off-target translation ratio (up to −62.4 percentage points), thus improving translation quality (up to +9.7 bilingual evaluation understudy). Analysis shows that our method can preserve the model’s performance on other tasks, such as supervised translation and general tasks. Code is released at https://github.com/alphadl/LanguageAware_Tuning .
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yibing Zhan, Xinghao Yang, Weifeng Liu 0001
Frontiers Inf. Technol. Electron. Eng.6
2025 WMANet:Weighted multiple adaptive feature attention for self-supervised single remote-sensing image denoising
Weifeng Liu 0001, Dapeng Tao, Baodi Liu, Yanjiang Wang 0001
Knowl. Based Syst.4
2025 PPBU: Progressive Pixel Bank Updating Strategy for Single Remote Sensing Image Denoising
Baodi Liu, Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001
IEEE Geosci. Remote. Sens. Lett.3
2025 GSF-GAN: A Global-Aware Selective Fusion Generative Adversarial Network for Multitemporal Cloud Removal
abstract
Remote sensing images are essential for surface observation, yet cloud cover often leads to a lack of spectral information, significantly disrupting the continuity of these images. Therefore, effective de-clouding techniques are crucial. Among these, multi-temporal de-clouding approaches offer notable advantages, as they utilize complementary images from different time periods and avoid the complexity of multimodal data fusion. However, maintaining image continuity and preserving fine details remains a major challenge due to the complex surface features, high detail recovery demands, and subtle differences between multi-temporal images. To address these challenges, we propose a novel de-clouding framework, the Global-aware Selective Fusion Generative Adversarial Network (GSF-GAN). GSF-GAN tackles the problem by introducing a Triplet Weight Selection Module (TWSM), which efficiently filters high-quality features to preserve more surface details, while the RescaleNorm-ReLU Swin Transformer (ReSwin) captures temporal variation patterns through global context modeling, further enhancing the de-clouding effect. Through experimental validation on STGAN and Sen2_MTC datasets, our method shows significant advantages in the de-cloud effect, and the PSNR and SSIM indexes are better than the control model, which verifies its effectiveness and superiority.
Aozhe Dou, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.4
2025 Multi-scale region selection network in deep features for full-field mammogram classification
abstract
Early diagnosis and treatment of breast cancer can effectively reduce mortality. Since mammogram is one of the most commonly used methods in the early diagnosis of breast cancer, the classification of mammogram images is an important work of computer-aided diagnosis (CAD) systems. With the development of deep learning in CAD, deep convolutional neural networks have been shown to have the ability to complete the classification of breast cancer tumor patches with high quality, which makes most previous CNN-based full-field mammography classification methods rely on region of interest (ROI) or segmentation annotation to enable the model to locate and focus on small tumor regions. However, the dependence on ROI greatly limits the development of CAD, because obtaining a large number of reliable ROI annotations is expensive and difficult. Some full-field mammography image classification algorithms use multi-stage training or multi-feature extractors to get rid of the dependence on ROI, which increases the computational amount of the model and feature redundancy. In order to reduce the cost of model training and make full use of the feature extraction capability of CNN, we propose a deep multi-scale region selection network (MRSN) in deep features for end-to-end training to classify full-field mammography without ROI or segmentation annotation. Inspired by the idea of multi-example learning and the patch classifier, MRSN filters the feature information and saves only the feature information of the tumor region to make the performance of the full-field image classifier closer to the patch classifier. MRSN first scores different regions under different dimensions to obtain the location information of tumor regions. Then, a few high-scoring regions are selected by location information as feature representations of the entire image, allowing the model to focus on the tumor region. Experiments on two public datasets and one private dataset prove that the proposed MRSN achieves the most advanced performance.
Luhao Sun, Bowen Han 0001, Wenzong Jiang, Weifeng Liu 0001, Baodi Liu, Dapeng Tao, Chao Li 0075
Medical Image Anal.4
2025 Parentheses insertion based sentence-level text adversarial attack
Xinghao Yang, Baodi Liu, Honglong Chen, Dapeng Tao, Weifeng Liu 0001
Multim. Syst.6
2025 Correction: Parentheses insertion based sentence-level text adversarial attack
Xinghao Yang, Baodi Liu, Honglong Chen, Dapeng Tao, Weifeng Liu 0001
Multim. Syst.6
2025 Feature aggregation and connectivity for object re-identification
Dongchen Han, Baodi Liu, Shuai Shao 0006, Weifeng Liu 0001, Yicong Zhou
Pattern Recognit.4
2025 IW-ViT: Independence-Driven Weighting Vision Transformer for out-of-distribution generalization
Weifeng Liu 0001, Haoran Yu 0005, Yingjie Wang 0007, Baodi Liu, Dapeng Tao, Honglong Chen
Pattern Recognit.1
2025 A Triple Stealthy Backdoor: Hidden in Spatial, Frequency, and Feature Domains
abstract
Backdoor attacks pose significant security risks to deep neural networks (DNNs). These attacks involve models that make intentionally incorrect (and potentially targeted) predictions on poisoned inputs containing carefully crafted triggers, while operating normally with clean inputs. Prior studies have investigated the invisibility of backdoor triggers to improve attack stealthiness. However, they primarily concentrate on achieving invisibility solely in the spatial domain, ignoring the generation of invisible triggers in the frequency and feature domains. This constraint makes the poisoned images vulnerable to detection by recent defense mechanisms. To tackle this problem, we introduce a Triple stealthy BAckdoor attack approach, termed TriBA, which simultaneously ensures the invisibility of triggers in all the spatial, frequency, and feature domains, to achieve desirable attack performance, while ensuring strong stealthiness. Specifically, we initially utilize Wavelet Transform to embed the high-frequency information from the trigger image into the clean image to ensure effective attack performance. Then, to achieve strong stealthiness across both spatial and frequency domains, we integrate Fourier Transform and Cosine Transform to blend the poisoned image and clean image in the frequency domain. Furthermore, TriBA adopts an attack strategy to make the backdoor features similar to clean features in the feature space, which guarantees trigger invisibility in the feature domain while maintaining attack effectiveness. We theoretically prove the effectiveness of this strategy. Finally, TriBA has been comprehensively evaluated on four datasets against popular image classifiers, demonstrating a marked improvement over existing state-of-the-art backdoor attacks in terms of both attack success rate and stealthiness.
Yudong Gao, Honglong Chen, Peng Sun 0003, Junjian Li, Yangxu Yin, Zhibo Wang 0001, Weifeng Liu 0001
IEEE Trans. Dependable Secur. Comput.7
2025 MSC-GAN: A Multistream Complementary Generative Adversarial Network With Grouping Learning for Multitemporal Cloud Removal
abstract
Optical remote sensing images have extensive application value, but cloud contamination greatly limits their potential use in the field of geographic information. Cloud removal aims to restore clear, unobstructed images from cloud-covered ones for subsequent in-depth analysis. Due to severe cloud cover problems such as thick clouds in some areas of remote sensing images, cloud removal tasks have become challenging. Recently, many methods have attempted to incrementally fill in obscured regions by fusing cloud-free information from multitemporal data. However, most of these methods fail to effectively utilize the interaction among different temporal data, and some information of data is easily lost in the process of deep transmission, this causes problems such as inadequate cloud removal and blurred recovery of ground under the clouds. Therefore, we propose a multistream complementary generative adversarial network (MSC-GAN) for cloud removal using multitemporal data. First, it employs a multistream complementary (MSC) architecture in the down-sampling feature encoding stage to effectively promote the interaction of feature information across multitemporal data, alleviating information loss as network depth increases. Second, to reduce the feature blur, we design a group feature reweighting (GFR) module as a complementary connection of long-distance information, in which the grouping learning and multidimensional parallel architecture can cost-effectively enhance semantic fusion between low-level and high-level features. Moreover, a channel enhancement method is introduced to assist in processing the underlying transition information, minimizing the interference of invalid information. Experimental results on multiple benchmark datasets under a series of image quality assessment metrics demonstrate the effectiveness of the proposed method.
Yanjiang Wang 0001, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Trans. Geosci. Remote. Sens.3
2025 See Degraded Objects: A Physics-Guided Approach for Object Detection in Adverse Environments
abstract
In adverse environments, the detector often fails to detect degraded objects because they are almost invisible and their features are weakened by the environment. Common approaches involve image enhancement to support detection, but they inevitably introduce human-invisible noise that negatively impacts the detector. In this work, we propose a physics-guided approach for object detection in adverse environments, which gives a straightforward solution that injects the physical priors into the detector, enabling it to detect poorly visible objects. The physical priors, derived from the imaging mechanism and image property, include environment prior and frequency prior. The environment prior is generated from the physical model, e.g., the atmospheric model, which reflects the density of environmental noise. The frequency prior is explored based on an observation that the amplitude spectrum could highlight object regions from the background. The proposed two priors are complementary in principle. Furthermore, we present a physics-guided loss that incorporates a novel weight item, which is estimated by applying the membership function on physical priors and could capture the extent of degradation. By backpropagating the physics-guided loss, physics knowledge is injected into the detector to aid in locating degraded objects. We conduct experiments in synthetic foggy environment, real foggy environment, and real underwater scenario. The results demonstrate that our method is effective and achieves state-of-the-art performance. The code is available at https://github.com/PangJian123/See-Degraded-Objects.
Weifeng Liu 0001, Jian Pang, Bingfeng Zhang, Baodi Liu, Dapeng Tao
IEEE Trans. Image Process.1
2025 Black-Box Adversarial Defense Based on Image Decomposition and Reconstruction
abstract
Adversarial attacks have challenged the security of deep neural networks (DNNs) recently. The most prominent adversarial attack methods include backdoor attacks, adversarial examples, etc. These attack methods inject triggers or perturbations into images, leading to extremely dangerous security vulnerability in deep learning domain. The various forms of adversarial attacks can contaminate DNNs with their distinct characteristics. The complexity of adversarial attack poses a great challenge to designing a general defense strategy. In this paper, we propose a novel defense method against most of adversarial attacks through Image Decomposition and Reconstruction (IDR). Our method can be applied to poisoned images without the need for internal information about the model or any prior knowledge of the clean/poisoned images. We apply a linear transformation on the poisoned image to destroy the perturbations or triggers and deploy a pre-trained diffusion model to reconstruct the original information. In particular, we propose a novel reverse process that utilizes the consistency of range-null space decomposition to guide the generation of purified images. The decomposition of the range-null space can guarantee the retrieval of image information, which enhances the robustness of our method and contributes to the reliable purification of poisoned images. We assess the effectiveness of our proposed IDR against various prevalent backdoor attacks, adversarial examples and Image-Scaling attack methods. The experimental results highlight the outstanding defensive capabilities of our proposed IDR, demonstrating an exceptionally high defense success rate.
Jimiao Yu, Honglong Chen, Junjian Li, Linghan Chen, Yudong Gao, Weifeng Liu 0001
IEEE Trans. Multim.6
2024 A Dual Stealthy Backdoor: From Both Spatial and Frequency Perspectives
abstract
Backdoor attacks pose serious security threats to deep neural networks (DNNs). Backdoored models make arbitrarily (targeted) incorrect predictions on inputs containing well-designed triggers, while behaving normally on clean inputs. Prior researches have explored the invisibility of backdoor triggers to enhance attack stealthiness. However, most of them only focus on the invisibility in the spatial domain, neglecting the generation of invisible triggers in the frequency domain. This limitation renders the generated poisoned images easily detectable by recent defense methods. To address this issue, we propose a DUal stealthy BAckdoor attack method named DUBA, which simultaneously considers the invisibility of triggers in both the spatial and frequency domains, to achieve desirable attack performance, while ensuring strong stealthiness. Specifically, we first use Wavelet Transform to embed the high-frequency information of the trigger image into the clean image to ensure attack effectiveness. Then, to attain strong stealthiness, we incorporate Fourier Transform and Cosine Transform to mix the poisoned image and clean image in the frequency domain. Moreover, DUBA adopts a novel attack strategy, training the model with weak triggers and attacking with strong triggers to further enhance attack performance and stealthiness. DUBA is evaluated extensively on four datasets against popular image classifiers, showing significant superiority over state-of-the-art backdoor attacks in attack success rate and stealthiness.
Yudong Gao, Honglong Chen, Peng Sun 0003, Junjian Li, Anqing Zhang, Zhibo Wang 0001, Weifeng Liu 0001
AAAI7
2024 Adaptive Bidirectional Displacement for Semi-Supervised Medical Image Segmentation
abstract
Consistency learning is a central strategy to tackle unlabeled data in semi-supervised medical image segmentation (SSMIS), which enforces the model to produce consistent predictions under the perturbation. However, most current approaches solely focus on utilizing a specific single perturbation, which can only cope with limited cases, while employing multiple perturbations simultaneously is hard to guarantee the quality of consistency learning. In this paper, we propose an Adaptive Bidirectional Displacement (ABD) approach to solve the above challenge. Specifically, we first design a bidirectional patch displacement based on reliable prediction confidence for unlabeled data to generate new samples, which can effectively suppress uncontrollable regions and still retain the influence of input perturbations. Meanwhile, to enforce the model to learn the potentially uncontrollable content, a bidirectional displacement operation with inverse confidence is proposed for the labeled images, which generates samples with more unreliable information to facilitate model learning. Extensive experiments show that ABD achieves new state-of-the-art performances for SSMIS, significantly improving different base-lines. Source code is available at https://github.com/chy-upclABD.
Hanyang Chi, Jian Pang, Bingfeng Zhang, Weifeng Liu 0001
CVPR4
2024 Rethinking Prior Information Generation with CLIP for Few-Shot Segmentation
abstract
Few-shot segmentation remains challenging due to the limitations of its labeling information for unseen classes. Most previous approaches rely on extracting high-level fea-ture maps from the frozen visual encoder to compute the pixel- wise similarity as a key prior guidance for the decoder. However, such a prior representation suffers from coarse granularity and poor generalization to new classes since these high-level feature maps have obvious category bias. In this work, we propose to replace the visual prior representation with the visual-text alignment capacity to capture more reliable guidance and enhance the model generalization. Specifically, we design two kinds of trainingfree prior information generation strategy that attempts to utilize the semantic alignment capability of the Contrastive Language-Image Pre-training model (CLIP) to locate the target class. Besides, to acquire more accurate prior guidance, we build a high-order relationship of attention maps and utilize it to refine the initial prior information. Experiments on both the PASCAL-5i and COCO-20i datasets show that our method obtains a clearly substantial improvement and reaches the new state-of-the-art performance. The code is available on the project website11https://github.com/vangjin/PI-CLIP.
Bingfeng Zhang, Jian Pang, Honglong Chen, Weifeng Liu 0001
CVPR5
2024 Adaptive Immune-based Sound-Shape Code Substitution for Adversarial Chinese Text Attacks
abstract
Adversarial textual examples reveal the vulnerability of natural language processing (NLP) models.Most existing text attack methods are designed for English text, while the robust implementation of the second popular language, i.e., Chinese with 1 billion users, is greatly underestimated.Although several Chinese attack methods have been presented, they either directly transfer from English attacks or adopt simple greedy search to optimize the attack priority, usually leading to unnatural sentences.To address these issues, we propose an adaptive Immune-based Sound-Shape Code (ISSC) algorithm for adversarial Chinese text attacks.Firstly, we leverage the Sound-Shape code to generate natural substitutions, which comprehensively integrate multiple Chinese features.Secondly, we employ adaptive immune algorithm (IA) to determine the replacement order, which can reduce the duplication of population to improve the search ability.Extensive experimental results validate the superiority of our ISSC in producing high-quality Chinese adversarial texts.
Xinghao Yang, Baodi Liu, Weifeng Liu 0001
EMNLP5
2024 HTPSeg: A Semantic Segmentation Database for House-Tree-Person Psychological Test
abstract
The House Tree Person (HTP) test is widely recommended for clinical application of mental illness. Traditionally, therapists assess a patient's mental state by analyzing the content of their HTP drawings manually, which is time-consuming and susceptible to the therapist's subjective influences. Recently, using intelligent diagnostic models to tackle HTP test has attracted much attention. However, most existing models attempt to make a binary classification for HTP drawings, i.e., positive or negative psychological states, making details that can reflect the psychological state of the patient lost. To address the above challenge, in this paper, we introduce the semantic segmentation task into the HTP test for the first time to generate pixel-level semantic information. We first construct a semantic segmentation dataset about HTP psychological diagnosis named HTPSeg. Subsequently, to overcome the domain gap between HTP drawings and natural images, we propose to use the Low-Rank Adaptation (LoRA) fine-tuning strategy to adapt the Segment Anything Model (SAM) to the task of HTP drawing analysis. Specifically, we integrate the rank decomposition matrices into the projection layer of the transformer block in SAM's image encoder for fine-tuning. Additionally, we freeze the Mask Decoder and fine-tune the Prompt Encoder using default embeddings. Extensive experiments indicate the effectiveness and efficiency of the proposed method. The dataset and code will be available at https://github.com/clown06/HTPSeg.
Bingfeng Zhang, Weifeng Liu 0001
ICTAI5
2024 Discriminative Representation-Based Classifier for Few-Shot Remote Sensing Classification
Tianhao Yuan, Weifeng Liu 0001, Yingjie Wang 0007, Baodi Liu
PRCV (13)2
2024 Ensembling Multi-View Discriminative Semantic Feature for Few-Shot Classification
Rui Xu 0012, Shuai Shao 0006, Lei Xing 0005, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001
Eng. Appl. Artif. Intell.6
2024 Feedback-Irrelevant Mapping: An evaluation method for decoupled few-shot classification
Rui Xu 0012, Shuai Shao 0006, Lei Xing 0005, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001
Eng. Appl. Artif. Intell.6
2024 Adaptive Gradient-based Word Saliency for adversarial text attacks
Yupeng Qi, Xinghao Yang, Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001
Neurocomputing5
2024 MWLN: Multilevel Wavelet Learning Network for Continuous-Scale Remote-Sensing Image Super-Resolution
abstract
Remote-sensing image super-resolution (SR) reconstructs high resolution (HR) with texture from the input low resolution (LR). It has been widely used and applied in image-processing tasks. However, most algorithms focus on designing more complex structures to enhance performance, ignoring learning frequency information. Moreover, existing methods are designed for SR tasks with specific scales, such as scales of 2 and 4. It limits the network performance in applications. To alleviate the above issues, this letter designs a multilevel wavelet learning network (MWLN) for continuous-scale remote-sensing image SR. MWLN achieves continuous magnification remote-sensing image SR tasks without training at different scales multiple times through multilevel wavelet feature aggregation (MWFA) and self-learning implicit representation (SLIR). MWFA extracts hierarchical features and applies discrete wavelet transforms (DWTs), capturing high-frequency information while avoiding information loss. Moreover, this letter cascades a multidimensional attention mechanism model channel and spatial features and enhances features’ interaction. SLIR maps the image coordinates and red, green, and blue (RGB) value through self-learning, realizing the continuous-scale reconstruction. Extensive experimental results demonstrate that MWLN outperforms the compared methods in quantitative and qualitative results on specific and continuous-scale remote-sensing image SR tasks.
Baodi Liu, Lifei Zhao, Weifeng Liu 0001
IEEE Geosci. Remote. Sens. Lett.3
2024 SELM: Self-Motivated Ensemble Learning Model for Cross-Domain Few-Shot Classification in Hyperspectral Images
abstract
Hyperspectral image (HSI) classification is a common task in remote sensing that often faces challenges due to limited samples and cross-domain discrepancies between training and test data. This particular problem is termed as HSI Cross-domain Few-Shot Classification (HSI-CFSC). To solve this problem, we propose a Self-motivated Ensemble Learning Model (SELM). Building upon source pre-trained representations, our end-to-end approach comprises a self-training paradigm to iteratively refine target representations independent of direct source supervision. Moreover, an ensemble classifier suite leveraging diverse decision boundaries is optimized to excavate comprehensive classification cues from limited labeled target data. The OA, AA and Kappa of SELM in UP, PC and Salinas data sets are respectively 86.55%, 82.27%, 82.10%, 98.07%, 94.20%, 97.30% and 91.33%, 94.96%, 90.37%, which achieve the state-of-art performance compared with other classical methods.
Shiyuan Zhao, Shuai Shao 0006, Weifeng Liu 0001, Xinmin Ge, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.4
2024 Remote sensing image cloud removal based on multi-scale spatial information perception
Aozhe Dou, Weifeng Liu 0001, Zhenzhong Wang, Baodi Liu
Multim. Syst.3
2024 Design of integrated interactive system for pre-diagnosis of breast cancer pathological images based on CNN and PyQt5
Yunkai Yang, Qijia Yang, Weifeng Liu 0001, Baodi Liu
Multim. Syst.3
2024 Target Oriented Dynamic Adaption for Cross-Domain Few-Shot Learning
abstract
Abstract Few-shot learning has achieved satisfactory progress over the years, but these methods implicitly hypothesize that the data in the base (source) classes and novel (target) classes are sampled from the same data distribution (domain), which is often invalid in reality. The purpose of cross-domain few-shot learning (CD-FSL) is to successfully identify novel target classes with a small quantity of labeled instances on the target domain under the circumstance of domain shift between the source domain and the target domain. However, in CD-FSL, the knowledge learned by the network on the source domain often suffers from the situation of inadaptation when it is transferred to the target domain, since the instances on the source and target domains do not obey the same data distribution. To surmount this problem, we propose a Target Oriented Dynamic Adaption (TODA) model, which uses a tiny amount of target data to orient the network to dynamically adjust and adapt during training. Specifically, this work proposes a domain-specific adapter to ameliorate the network inadaptability issues in transfer to the target domain. The domain-specific adapter can make the extracted features more specific to the tasks in the target domain and reduce the impact of tasks in the source domain by combining them with the mainstream backbone network. In addition, we propose an adaptive optimization method in the network optimization process, which assigns different weights according to the importance of different optimization tasks. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our TODA method.
Xinyi Chang, Chunyu Du, Xinjing Song, Weifeng Liu 0001, Yanjiang Wang 0001
Neural Process. Lett.4
2024 Central Attention with Multi-Graphs for Image Annotation
abstract
Abstract In recent decades, the development of multimedia and computer vision has sparked significant interest among researchers in the field of automatic image annotation. However, much of the research has primarily focused on using a single graph for annotating images in semi-supervised learning. Conversely, numerous approaches have explored the integration of multi-view or image segmentation techniques to create multiple graph structures. Yet, relying solely on a single graph proves to be challenging, as it struggles to capture the complete manifold of structural information. Furthermore, the computational complexity of building multiple graph structures based on multi-view or image segmentation is substantial and time-consuming. To address these issues, we propose a novel method called "Central Attention with Multi-graphs for Image Annotation." Our approach emphasizes the critical role of the central image region in the annotation process. Remarkably, we demonstrate that impressive performance can be achieved by leveraging just two graph structures, composed of central and overall features, in semi-supervised learning. To validate the effectiveness of our proposed method, we conducted a series of experiments on benchmark datasets, including Corel5K, ESPGame, and IAPRTC12. These experiments provide empirical evidence of our method’s capabilities.
Baodi Liu, Qianqian Shao, Weifeng Liu 0001
Neural Process. Lett.4
2024 Simplified Multi-head Mechanism for Few-Shot Remote Sensing Image Classification
abstract
Abstract The study of few-shot remote sensing image classification has received significant attention. Although meta-learning-based algorithms have been the primary focus of recent examination, feature fusion methods stress feature extraction and representation. Nonetheless, current feature fusion methods, like the multi-head mechanism, are restricted by their complicated network structure and challenging training process. This manuscript presents a simplified multi-head mechanism for obtaining multiple feature representations from a single sample. Furthermore, we perform specific fundamental transformations on remote-sensing images to obtain more suitable features for information representation. Specifically, we reduce multiple feature extractors of the multi-head mechanism to a single one and add an image transformation module before the feature extractor. After transforming the image, the features are extracted resulting in multiple features for each sample. The feature fusion stage is integrated with the classification prediction stage, and multiple linear classifiers are combined for multi-decision fusion to complete feature fusion and classification. By combining image transformation with feature decision fusion, we compare our results with other methods through validation tests and demonstrate that our algorithm simplifies the multi-head mechanism while maintaining or improving classification performance.
Xujian Qiao, Lei Xing 0005, Anxun Han, Weifeng Liu 0001, Baodi Liu
Neural Process. Lett.4
2024 Few-shot image classification via hybrid representation
Baodi Liu, Shuai Shao 0006, Lei Xing 0005, Weifeng Liu 0001, Weijia Cao, Yicong Zhou
Pattern Recognit.5
2024 MCNet: Magnitude consistency network for domain adaptive object detection under inclement environments
Jian Pang, Weifeng Liu 0001, Bingfeng Zhang, Xinghao Yang, Baodi Liu, Dapeng Tao
Pattern Recognit.2
2024 Enhanced online CAM: Single-stage weakly supervised semantic segmentation via collaborative guidance
Bingfeng Zhang, Xuru Gao, Siyue Yu, Weifeng Liu 0001
Pattern Recognit.4
2024 Weight Saliency search with Semantic Constraint for Neural Machine Translation attacks
Wen Han, Xinghao Yang, Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001
Pattern Recognit. Lett.5
2024 FADS: Fourier-Augmentation Based Data-Shunting for Few-Shot Classification
abstract
Collecting a substantial number of labeled samples is infeasible in many real-world scenarios, thereby bringing out challenges for supervised classification. The research on Few-Shot Classification (FSC) aims to address this issue. Current FSC methods mainly leverage ideas such as meta-learning, self-supervised learning, and data augmentation. Among them, data augmentation appears to be an extremely efficient approach to alleviate the aforementioned data-deficiency problem. Here, we propose a novel data augmentation based FSC method termed Fourier-Augmentation based Data-Shunting (FADS). FADS mainly contains two operations, namely Fourier-based data augmentation (FDA) and data shunting. (i) Fourier transform has a desirable property for classification tasks: the image’s phase and amplitude components in the frequency domain correspond to its high-level structure (i.e., semantic) and low-level style (i.e., statistic) information, which do not interfere with each other. Inspired by this observation, we design the FDA operation, which changes the amplitude spectrum of the to-be-augmented images to obtain new images of the same category. (ii) Then we design the data shunting operation to cooperate with the FDA to accomplish FSC. Specifically, it splits the augmented data into different groups to get independent, weak decisions and then fuses them to obtain a unified, strong decision. We conduct experiments on four benchmark datasets. Results show that utilizing our method brings a performance gain of 0.3%-2% in terms of classification accuracy, compared with the classical methods.
Shuai Shao 0006, Yan Wang 0076, Bin Liu 0021, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu
IEEE Trans. Circuits Syst. Video Technol.4
2024 EME: Energy-Based Multiexpert Model for Long-Tailed Remote Sensing Image Classification
abstract
The distribution of remote sensing scene images often follows a long-tailed pattern, where there is an abundance of samples in a few dominant classes and a scarcity of samples in most other classes. This presents two major challenges when it comes to identifying this type of data: Head-Dominance: Models trained on such data tend to prioritize the dominant classes, overlooking the tail classes and resulting in poor performance when it comes to recognizing them. Tail-Interference: The presence of tail classes disrupts the learned representations for the head classes, acting as noise that negatively impacts the recognition accuracy of the head data. To address these challenges, we propose an innovative solution called the energy-based multiexpert (EME) model. The core concept behind this approach is to utilize energy-based discriminators (EDors) to separate the data into head and tail categories. Subsequently, we design multiple experts to classify the head and tail data separately, ensuring that the significant differences in data volume between these categories do not interfere with each other. Experimental results obtained by applying the EME model to three remote sensing datasets demonstrate its efficiency, outperforming current state-of-the-art (SOTA) methods. These findings underscore the effectiveness of our proposed approach in addressing the challenges posed by the long-tailed distribution in remote sensing scene images.
Shuai Shao 0006, Shiyuan Zhao, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Trans. Geosci. Remote. Sens.4
2024 Investigating the Backdoor on DNNs Based on Recolorization and Reconstruction: From a Multi-Channel Perspective
abstract
Recently, backdoor attacks have become a serious security threat to Deep Neural Networks (DNNs). Backdoor attacks involve embedding a hidden backdoor into a DNN model, compelling it to correctly classify benign images while erroneously classifying images with backdoor triggers as the target label. However, both current backdoor attacks and defenses have their limitations. In backdoor attacks, they are either non-stealthy or vulnerable to well-designed backdoor defense strategies. As for backdoor defenses, they often rely heavily on additional assumptions (such as determined extra clean images) and are not universally applicable, which may become impractical in the face of the latest backdoor attacks. To address the above problems, in this paper, we investigate the backdoor attack and defense strategies from a multi-channel perspective. Specifically, in terms of attacks, we propose a recolorization based attack method (RC-Attack) to generate triggers in color ab channels, which is more stealthy and effective. In terms of defenses, we propose a reconstruction-based defense method (RC-Defense) to reconstruct the color AB channels and lightness channel respectively, thus making the triggers in the reconstructed images ineffective, which is a more practical solution. Extensive experiments are conducted to demonstrate the superior performance of the proposed RC-Attack in terms of effectiveness, stealthiness and defense-resistance, and also to validate the effectiveness of the proposed RC-Defense.
Honglong Chen, Yudong Gao, Anqing Zhang, Peng Sun 0003, Nan Jiang 0013, Weifeng Liu 0001
IEEE Trans. Inf. Forensics Secur.6
2024 Call White Black: Enhanced Image-Scaling Attack in Industrial Artificial Intelligence Systems
abstract
The increasing prevalence of deep neural networks (DNNs) in industrial artificial intelligence systems (IAISs) promotes the development of industrial automation. However, the growing employment of DNNs also exposes them to various attacks. Recent studies have shown that the data preprocessing process of DNNs is vulnerable to image-scaling attack. Such attacks can craft an attack image, which looks like a given source image but becomes a different target image after being scaled to the target size. The attack images generated by existing image-scaling attacks are easily perceivable to the human visual system, significantly degrading the attack's stealthiness. In this paper, we investigate image-scaling attack from the perspective of signal processing. We unearth that the root cause of the weak deceiving effects of existing image-scaling attack images lies in the introduction of additional high-frequency signals during their construction. Thus, we propose an enhanced image-scaling attack (EIS), which employs adversarial images crafted based on the source (“clean”) images as the target images. Those adversarial images preserve the “clean” pixel information of source images, thereby significantly mitigating the emergence of additional high-frequency signals in the attack images. Specifically, we consider three realistic threat models covering deep models' training and inference phases. Correspondingly, we design three strategies tailored to generate adversarial images with vicious patterns. These patterns are subsequently integrated into the attack images, which can mislead a model with target input size after the necessary scaling operation. Extensive experiments validate the superior performance of the proposed image-scaling attack compared to the original one.
Junjian Li, Honglong Chen, Peng Sun 0003, Zhibo Wang 0001, Zhichen Ni, Weifeng Liu 0001
IEEE Trans. Ind. Informatics6
2024 Deep Location Soft-Embedding-Based Network With Regional Scoring for Mammogram Classification
abstract
Early detection and treatment of breast cancer can significantly reduce patient mortality, and mammogram is an effective method for early screening. Computer-aided diagnosis (CAD) of mammography based on deep learning can assist radiologists in making more objective and accurate judgments. However, existing methods often depend on datasets with manual segmentation annotations. In addition, due to the large image sizes and small lesion proportions, many methods that do not use region of interest (ROI) mostly rely on multi-scale and multi-feature fusion models. These shortcomings increase the labor, money, and computational overhead of applying the model. Therefore, a deep location soft-embedding-based network with regional scoring (DLSEN-RS) is proposed. DLSEN-RS is an end-to-end mammography image classification method containing only one feature extractor and relies on positional embedding (PE) and aggregation pooling (AP) modules to locate lesion areas without bounding boxes, transfer learning, or multi-stage training. In particular, the introduced PE and AP modules exhibit versatility across various CNN models and improve the model's tumor localization and diagnostic accuracy for mammography images. Experiments are conducted on published INbreast and CBIS-DDSM datasets, and compared to previous state-of-the-art mammographic image classification methods, DLSEN-RS performed satisfactorily.
Bowen Han 0001, Luhao Sun, Chao Li 0075, Wenzong Jiang, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Trans. Medical Imaging6
2024 AdvST: Generating Unrestricted Adversarial Images via Style Transfer
abstract
Recent years have witnessed extensive applications of Deep Neural Networks (DNNs) in various vision tasks. However, DNNs are vulnerable to adversarial images crafted by introducing perturbations into inputs to induce incorrect predictions. Unlike$L_{p}$-norm restricted adversarial attacks, many unrestricted attacks have been proposed by modifying attributes of the image (e.g., edge, color), while the critical components of the image are preserved. However, most existing unrestricted attacks easily introduce unnatural distortions, colors, stains and schemes, in the generated adversarial images. This paper proposes a novel unrestricted attack (named AdvST) to create stylized, natural-looking, and high-transferability adversarial images. The basic idea of AdvST is to embed adversarial perturbations when transferring the style from the reference image onto the original image (i.e., rendering the original image's semantic contents into the reference image's style). To further improve the image quality of generated adversarial images, we refine two kinds of reference images (i.e., photographs and artworks) based on different attractive styles and design two attacks accordingly. For photorealistic attack, we incorporate semantic information obtained from segmentation maps to improve the photo realism of adversarial images. For artistic attack, we propose integrating edge information extracted by the Laplace operator to preserve the structural integrity of the original image. Extensive experimental results validate the superior performance of AdvST in terms of adversarial image quality and black-box transferability compared to benchmark methods.
Honglong Chen, Peng Sun 0003, Junjian Li, Anqing Zhang, Weifeng Liu 0001, Nan Jiang 0013
IEEE Trans. Multim.6
2023 Tensor Canonical Correlation Analysis Networks for Multi-view Remote Sensing Scene Recognition (Extended Abstract)
abstract
Remote sensing (RS) images are frequently observed from multiviews. In this paper, we propose the tensor canonical correlation analysis network (TCCANet) to tackle the multiview RS recognition problem. Particularly, TCCANet learns filter banks by simultaneously maximizing arbitrary number of views with high-order-correlation and solves the optimization problem by decomposing a covariance tensor. After the convolutional stage, we utilize binarization and block-wise histogram strategies to generate the final feature. Furthermore, we also develop a Multiple Scale version of TCCANet, i.e., MS-TCCANet, to extract enriched representation of the RS data by incorporating all previous convolutional layers. Numerical experiment results on RSSCN7 and SAT-6 datasets demonstrate the advantages of TCCANet and MS-TCCANet for RS scene recognition.
Xinghao Yang, Weifeng Liu 0001, Wei Liu 0007
ICDE2
2023 Attribute Space Analysis for Image Editing
Shuqi Yang, Baodi Liu, Weifeng Liu 0001
ICIG (2)4
2023 Self-Compensating Learning for Few-Shot Segmentation
abstract
Few-shot segmentation (FSS) has witnessed rapid development. Most existing approaches extract prototypes from support images to segment query images. However, the integrity and validity of these support prototypes cannot be guaranteed. To solve the above drawbacks, we propose a self-compensating strategy, aiming to provide query-aware support information, to build more effective matching between support information and query images. Specifically, we design a prototype compensating module to mine useful information from the query prediction, to update original support prototypes as new query-aware support prototypes. Then the updated prototypes are utilized to perform the second matching with query features. In addition, we also compensate the information of original prior masks on the second matching phase, to improve the quality of prior masks. With improved prototype representations and prior knowledge, our approach can directly improve the performance of different approaches with new state-of-the-art performances.
Bingfeng Zhang, Weifeng Liu 0001, Baodi Liu, Siyue Yu
ICIP3
2023 Tilted Sparse Additive Models
abstract
Additive models have been burgeoning in data analysis due to their flexible representation and desirable interpretability. However, most existing approaches are constructed under empirical risk minimization (ERM), and thus perform poorly in situations where average performance is not a suitable criterion for the problems of interest, e.g., data with complex non-Gaussian noise, imbalanced labels or both of them. In this paper, a novel class of sparse additive models is proposed under tilted empirical risk minimization (TERM), which addresses the deficiencies in ERM by imposing tilted impact on individual losses, and is flexibly capable of achieving a variety of learning objectives, e.g., variable selection, robust estimation, imbalanced classification and multiobjective learning. On the theoretical side, a learning theory analysis which is centered around the generalization bound and function approximation error bound (under some specific data distributions) is conducted rigorously. On the practical side, an accelerated optimization algorithm is designed by integrating Prox-SVRG and random Fourier acceleration technique. The empirical assessments verify the competitive performance of our approach on both synthetic and real data.
Yingjie Wang 0007, Hong Chen 0004, Weifeng Liu 0001, Fengxiang He, Tieliang Gong, Youcheng Fu, Dacheng Tao
ICML3
2023 Annealing Genetic-based Preposition Substitution for Text Rubbish Example Generation
abstract
Modern Natural Language Processing (NLP) models expose under-sensitivity towards text rubbish examples. The text rubbish example is the heavily modified input text which is nonsensical to humans but does not change the model’s prediction. Prior work crafts rubbish examples by iteratively deleting words and determining the deletion order with beam search. However, the produced rubbish examples usually cause a reduction in model confidence and sometimes deliver human-readable text. To address these problems, we propose an Annealing Genetic based Preposition Substitution (AGPS) algorithm for text rubbish sample generation with two major merits. Firstly, the AGPS crafts rubbish text examples by substituting input words with meaningless prepositions instead of directly removing them, which brings less degradation to the model’s confidence. Secondly, we design an Annealing Genetic algorithm to optimize the word replacement priority, which allows the Genetic Algorithm (GA) to jump out the local optima with probabilities. This is significant in achieving better objectives, i.e., a high word modification rate and a high model confidence. Experimental results on five popular datasets manifest the superiority of AGPS compared with the baseline and expose the fact: the NLP models can not really understand the semantics of sentences, as they give the same prediction with even higher confidence for the nonsensical preposition sequences.
Xinghao Yang, Baodi Liu, Weifeng Liu 0001, Honglong Chen
IJCAI4
2023 Prompt-Learning for Cross-Lingual Relation Extraction
abstract
Relation Extraction (RE) is a crucial task in Information Extraction, which entails predicting relationships between entities within a given sentence. However, extending pre-trained RE models to other languages is challenging, particularly in real-world scenarios where Cross-Lingual Relation Extraction (XRE) is required. Despite recent advancements in Prompt-Learning, which involves transferring knowledge from Multilingual Pre-trained Language Models (PLMs) to diverse downstream tasks, there is limited research on the effective use of multilingual PLMs with prompts to improve XRE. In this paper, we present a novel XRE algorithm based on Prompt-Tuning, referred to as Prompt-Xre. To evaluate its effectiveness, we design and implement several prompt templates, including hard, soft, and hybrid prompts, and empirically test their performance on competitive multilingual PLMs, specifically mBART. Our extensive experiments, conducted on the low-resource ACE05 benchmark across multiple languages, demonstrate that our Prompt-Xre algorithm significantly outperforms both vanilla multilingual PLMs and other existing models, achieving state-of-the-art performance in XRE. To further show the generalization of our Prompt-XRE on larger data scales, we construct and release a new XRE dataset-WMTI7-EnZh XRE, containing 0.9M English-Chinese pairs extracted from WMT 2017 parallel corpus. Experiments on WMTI7-EnZh XRE also show the effectiveness of our Prompt-XRE against other competitive baselines. The code and newly constructed dataset are freely available at httus://2ithub.com/HSU-CHIA-MING/Promut-XRE.
Chiaming Hsu, Changtong Zan, Liang Ding 0006, Longyue Wang, Weifeng Liu 0001, Wenbin Hu 0001
IJCNN6
2023 Generating New Paintings by Semantic Guidance
Fei Wang 0032, Junzhou Xie, Weifeng Liu 0001
MMM (2)4
2023 A Stable Vision Transformer for Out-of-Distribution Generalization
Haoran Yu 0005, Baodi Liu, Yingjie Wang 0007, Kai Zhang 0029, Dapeng Tao, Weifeng Liu 0001
PRCV (8)6
2023 Deep Positional-Representation-Based Local Information Retention Networks for Mammography Classification
abstract
Early diagnosis of breast cancer is challenging because in the most common mammogram images, the tumor usually occupies only a very small part of the entire image, which often makes deep learning models lose attention to the tumor area. In previous work, most models solved this problem by using ROI labeling to train models, which was expensive and difficult to widely apply. Some recent ROI-free methods use multi-scale features or multi-stage training, which gets rid of the model's dependence on ROI but greatly increases the computational complexity and deployment difficulty, limiting the potential of deep neural networks. Therefore, a deep positional-representation-based local information retention networks (PR-LIR) was proposed. PR-LIR is a lightweight, end-to-end mammogram classification model, which uses positional representation (PR) and multi-scale regional pooling (MRP) modules to locate tumor regions and retain regional semantic information of small target tumors at different scales, without ROI labeling and multi-stage training, and almost no increasement in parameters and computational complexity. In particular, the proposed PR and MRP modules have good generalization performance, which can be applied to most CNN models and improve the classification accuracy of mammography images. Experimental results on two publicly available datasets show that PR-LIR achieves the best AUC and satisfactory accuracy compared to the previous state-of-the-art mammogram classification method.
Bowen Han 0001, Luhao Sun, Chao Li 0075, Wenzong Jiang, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
SMC6
2023 CSN: Component supervised network for few-shot classification
Rui Xu 0012, Shuai Shao 0006, Lei Xing 0005, Yujun Wei, Weifeng Liu 0001, Baodi Liu, Yanjiang Wang 0001
Eng. Appl. Artif. Intell.5
2023 Generation-based parallel particle swarm optimization for adversarial text attacks
Xinghao Yang, Yupeng Qi, Honglong Chen, Baodi Liu, Weifeng Liu 0001
Inf. Sci.5
2023 Selecting Information Fusion Generative Adversarial Network for Remote-Sensing Image Cloud Removal
abstract
The multi-temporal remote sensing cloud removal method has improved performance, but it lacks a screening mechanism during feature fusion, simply summing and fusing features from different temporal states. This results in the inclusion of unwanted clouds and redundant feature information, hindering the restoration of the landscape under the clouds. To address this, we propose a selective information fusion generative adversarial network (SIF-GAN) for remote sensing image cloud removal. SIF-GAN incorporates channel attention during feature extraction to capture important information in different channels and uses the selective information fusion network to assign weights to the feature information from other temporal states, selecting the crucial features for fusion. The feature of cloud-free regions in different temporal states is utilized maximally by the selection process to recover the image features under clouds. The results of the experiments show that SIF-GAN achieves superior cloud removal performance compared to other methods.
Wenzong Jiang, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.3
2023 Deformable Convolutional Network Constrained by Contrastive Learning for Underwater Image Enhancement
abstract
Autonomous underwater vehicles (AUVs) based on remote sensing technology have been widely applied in various underwater tasks. However, the complex underwater environment leads to challenges such as color distortion, blurred details, and fog effects in the underwater image directly acquired by AUVs. Although numerous existing methods aim to remove the color cast and restore image details, their effectiveness is still limited. This paper proposes a new method based on a deformable convolutional network and constrained by contrastive learning for underwater image enhancement. First, we propose a deformable convolutional residual block (DCRB) to achieve a more precise restoration of texture details by adaptively adjusting the convolution kernel shape. At the same time, we utilize the long-skip connection method of the U-Net architecture to preserve information that is prone to lose in shallow networks. Second, we propose a color contrastive loss function to compare the color difference between distorted images and the ground truth, resulting in a more realistic enhanced image. Finally, experimental results demonstrate that the proposed method outperforms the state-of-the-art methods regarding image quality and visual appeal.
Xinran Guo, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.3
2023 Double Discriminative Constraint-Based Affine Nonnegative Representation for Few-Shot Remote Sensing Scene Classification
abstract
Remote sensing scene classification (RSSC) has recently attracted more attention. However, due to restrictions in the imaging environment and equipment, it is difficult to get a large number of labeled images in remote sensing. This has led to the emergence of few-shot learning for RSSC, which aims to achieve better performance with few labeled samples. Remote sensing images’ large interclass similarity may cause classification confusion. To overcome this issue, this study proposes a double discriminative constraints-based affine nonnegative representation for few-shot RSSC. To be specific, we devise a novel representation-based classifier with two discriminative constraint terms in the objective function and utilize affine nonnegative constraints to restrict the learned parameters. These constraints reduce the correlation between classes and strengthen the class specificity of the learned parameters. Experiments on benchmark datasets demonstrate the effectiveness of our method.
Tianhao Yuan, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.2
2023 Shared Dictionary Learning Via Coupled Adaptations for Cross-Domain Classification
Yuying Cai, Baodi Liu, Weijia Cao, Honglong Chen, Weifeng Liu 0001
Neural Process. Lett.6
2023 GSA4FDA: Deep Geometric and Statistic Alignment for Fewer Labeled Domain Adaptation
Yuying Cai, Baodi Liu, Xinghao Yang, Dapeng Tao, Weifeng Liu 0001
Neural Process. Lett.6
2023 Dynamic Feature Attention Network for Remote Sensing Image Dehazing
Wenzong Jiang, Weifeng Liu 0001, Weijia Cao, Baodi Liu
Neural Process. Lett.3
2023 Feature Fusion Based Parallel Graph Convolutional Neural Network for Image Annotation
Weifeng Liu 0001, Baodi Liu
Neural Process. Lett.3
2023 Cross-Domain Few-Shot classification via class-shared and class-specific dictionaries
Lei Xing 0005, Baodi Liu, Dapeng Tao, Weijia Cao, Weifeng Liu 0001
Pattern Recognit.6
2023 Dynamic Pruning of Regions for Image-Sentence Matching
abstract
Image–sentence matching is becoming increasingly essential in the integrated understanding of vision and language. Prior approaches apply a pre-trained detection model to extract region features and explore fine-grained relationships between image and sentence by aggregating the similarities of all region–word pairs. However, all images are represented by the same number of regions, regardless of their respective semantic complexity, which results in a large number of redundant regions interfering with semantic inference and bringing additional computational burden. To address the lack of flexibility in image representation and information redundancy, a novel method named Dynamic Pruning of Regions for Image–Sentence Matching (DPRM) is proposed to efficiently capture relationships between text and image. In particular, a dynamic region pruning module is presented to dynamically select the appropriate number of regions according to the semantic complexity of each image, thus pruning redundant regions and reducing superfluous computations. Moreover, an inter-modality refinement module is designed to refine the fine-grained relationships of region–word pairs by retaining meaningful interaction features and suppressing interference from redundant alignments, which learns the more accurate semantic correspondences. Extensive experiments on MSCOCO and Flickr30K datasets prove the superiority of DPRM compared with previous approaches.
Jie Wu 0033, Weifeng Liu 0001, Leiquan Wang, Xiuxuan Shen, Chunlei Wu
Signal Process. Image Commun.2
2023 MSCET: A Multi-Scenario Offloading Schedule for Biomedical Data Processing and Analysis in Cloud-Edge-Terminal Collaborative Vehicular Networks
abstract
With the rapid development of Artificial Intelligence (AI) and Internet of Things (IoTs), an increasing number of computation intensive or delay sensitive biomedical data processing and analysis tasks are produced in vehicles, bringing more and more challenges to the biometric monitoring of drivers. Edge computing is a new paradigm to solve these challenges by offloading tasks from the resource-limited vehicles to Edge Servers (ESs) in Road Side Units (RSUs). However, most of the traditional offloading schedules for vehicular networks concentrate on the edge, while some tasks may be too complex for ESs to process. To this end, we consider a collaborative vehicular network in which the cloud, edge and terminal can cooperate with each other to accomplish the tasks. The vehicles can offload the computation intensive tasks to the cloud to save the resource of edge. We further construct the virtual resource pool which can integrate the resource of multiple ESs since some regions may be covered by multiple RSUs. In this paper, we propose a Multi-Scenario offloading schedule for biomedical data processing and analysis in Cloud-Edge-Terminal collaborative vehicular networks called MSCET. The parameters of the proposed MSCET are optimized to maximize the system utility. We also conduct extensive simulations to evaluate the proposed MSCET and the results illustrate that MSCET outperforms other existing schedules.
Zhichen Ni, Honglong Chen, Zhe Li 0026, Na Yan 0003, Weifeng Liu 0001, Feng Xia 0001
IEEE ACM Trans. Comput. Biol. Bioinform.6
2023 Non-Contrastive Nearest Neighbor Identity-Guided Method for Unsupervised Object Re-Identification
abstract
Recently, self-paced contrastive learning has emerged as a promising method for unsupervised object re-identification. These methods generate pseudo labels, store centroid features in the memory bank, and periodically update them. However, affected by the performance of the clustering method, within each cluster exists inevitably noisy instances, and self-paced contrastive learning usually requires a large number of negative samples from various classes, where false-negative samples give rise to the class collision issue. These lead to performing incorrect model optimization. In this paper, we propose a non-contrastive nearest neighbor identity-guided (NNNI) method to overcome these challenges. The advantage of NNNI is to provide the model with a highly accurate prior. Specifically, this method relies on the random identity sampler commonly used in re-identification tasks to provide the network with a regression target of the nearest neighbors of the same identity within a mini-batch. It encodes more and more information in an iterative process through a Siamese network with an exponential moving average to train high-quality representations. NNNI alleviates the negative effects of noise instances and corrects class collision issues during training. Extensive experiments show that our method is effective on unsupervised object re-identification and achieves state-of-the-art performance on three large-scale person re-identification datasets and one large-scale vehicle re-identification dataset, which is competitive with even supervised methods.
Dongchen Han, Weifeng Liu 0001, Mingchen Zou, Baodi Liu
IEEE Trans. Circuits Syst. Video Technol.2
2023 Attention-Based Multi-View Feature Collaboration for Decoupled Few-Shot Learning
abstract
Decoupled Few-shot learning (FSL) is an effective methodology that deals with the problem of data-scarce. Its standard paradigm includes two phases: (1) Pre-train. Generating a CNN-based feature extraction model (FEM) via base data. (2) Meta-test. Employing the frozen FEM to obtain the novel data features, then classifying them. Obviously, one crucial factor, the category gap, prevents the development of FSL, i.e., it is challenging for the pre-trained FEM to adapt to the novel class flawlessly. Inspired by a common-sense theory: the FEMs based on different strategies focus on different priorities, we attempt to address this problem from the multi-view feature collaboration (MVFC) perspective. Specifically, we first denoise the multi-view features by subspace learning method, then design three attention blocks (loss-attention block, self-attention block and graph-attention block) to balance the representation between different views. The proposed method is evaluated on four benchmark datasets and achieves significant improvements of 0.9%-5.6% compared with SOTAs.
Shuai Shao 0006, Lei Xing 0005, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.5
2023 Self-Paced Hard Task-Example Mining for Few-Shot Classification
abstract
In recent years, researchers have commonly employed assistant tasks to enhance the training phase of the few-shot classification models. Several methods have been proposed to exploit and optimize the training tasks, such as Curriculum Learning (CL) and Hard Example Mining (HEM). However, most of the existing strategies can not elaborately leverage the training tasks and share some common drawbacks, including 1) the ignorance of the target tasks’ properties, and 2) the neglect of sample relationships. In this work, we propose a Self-Paced Hard tAsk-Example Mining (SP-HAEM) method to solve these problems. Specifically, the SP-HAEM automatically chooses hard examples via the similarity between training and target tasks to optimize the support set. To represent the property of target tasks, SP-HAEM obtains a representation of the dataset, called “meta-task”. No need to apply an additional model to measure difficulty and choose hard examples like other HEM methods, SP-HAEM selects the tasks with large optimal transport distance to the meta-task as hard tasks. Thus, training with such hard tasks can not only enhances the generalization ability of the model but also eliminate the negative effect of redundancy tasks. To evaluate the effectiveness of SP-HAEM, we conduct extensive experiments on a variety of datasets, including MiniImageNet, TieredImageNet, and FC100. The results of the experiments show that SP-HAEM can achieve higher accuracy compared with the typical few-shot classification models, e.g., Prototypical Network, MAML, FEAT, and MTL.
Xinghao Yang, Xingxing Yao, Dapeng Tao, Weijia Cao, Weifeng Liu 0001
IEEE Trans. Circuits Syst. Video Technol.7
2023 Dynamic Adaptive Attention-Guided Self-Supervised Single Remote-Sensing Image Denoising
abstract
Optical remote sensing images are widely used in many fields, and local complex texture details in images usually play a critical role in downstream tasks. However, noise interference will destroy the complex texture in the image, thus reducing the accuracy of downstream tasks. The current attention mechanism usually focuses on the global high-level features in the image, so it cannot effectively focus on the high-frequency information in the local complex texture in the remote sensing image, and obtaining clean remote sensing images to train neural networks is difficult. Therefore, applying the current depth learning based natural image denoising methods directly to optical remote sensing images is challenging. To solve these problems, we propose a dynamic adaptive attention guided self-supervised single remote sensing image denoising network (DAA-SSID). We construct a dynamic adaptive attention module (DAAM) by dynamically calculating the activation intensity of each neuron and combining the spatial feature information extracted from remote sensing images. It can effectively extract complex texture features from remote sensing images when only a single remote sensing image participates in training. And we use independent random Bernoulli sampling in the training and inference stages respectively to prevent over-fitting caused by single-image training. Therefore, compared with other self-supervised denoising methods, our proposed model can denoise remote sensing images with more complex textures when only a single image destroyed by noise is used as the training input. Experiments on synthetic additive gaussian noise data and authentic noise data have shown that the proposed model achieves satisfactory results.
Minghao Liu 0016, Wenzong Jiang, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Trans. Geosci. Remote. Sens.3
2023 RAN: Region-Aware Network for Remote Sensing Image Super-Resolution
abstract
The remote sensing (RS) image super-resolution (SR) algorithm aims to reconstruct a high-resolution (HR) image with rich texture details from a given low-resolution (LR) image, improving the spatial resolution. It has been widely concerned in remote sensing image processing and application. Most current deep learning-based methods rely on paired training datasets. However, most datasets are often based on bicubic degradation. This single construction way limits the performance of the pre-trained network. Moreover, SR is an ill-posed problem in that multiple SR images are constructed from a single LR input. This paper proposes a Region-Aware Network (RAN) for remote sensing image super-resolution to alleviate the above issues. First, we introduce the contrastive learning strategy to mine the latent degraded representation of the image and serve as the prior knowledge of the network. Considering the RS images are acquired in specific scenes that have apparent self-similarity. Then, we propose a Region-Aware Module (RAM) based on attention mechanisms and the graph neural network to explore region information and cross-patch self-similarity. Extensive experiments have demonstrated that the proposed RAN adapts to RS image super-resolution tasks with various degradations and performs better in constructing texture information.
Baodi Liu, Lifei Zhao, Shuai Shao 0006, Weifeng Liu 0001, Dapeng Tao, Weijia Cao, Yicong Zhou
IEEE Trans. Geosci. Remote. Sens.4
2023 ARPCNN: Auxiliary Review-Based Personalized Attentional CNN for Trustworthy Recommendation
abstract
Convolutional neural network (CNN)-based recommender systems are playing an increasingly significant role in the vigorous development of Industrial Internet of Things, and have made great contributions to analyzing and mining a large amount of data to provide various services for terminal users. However, as the lack of explainability in deep learning, users often have low trust in the system due to their incomprehension of recommendation results. In addition, recommender systems have been facing a serious sparsity problem, and relying only on sparse rating data to learn user preferences and similarities may face malicious recommendation attacks. The abovementioned problems have been hindering the further improvement of recommendation performance. Therefore, in order to effectively alleviate the sparsity problem and meanwhile enhance the trustworthiness, an auxiliary review-based personalized attentional CNN (ARPCNN) is proposed in this article. By applying the proposed personalized word-level attention mechanism and personalized review-level attention mechanism in parallel CNNs, critical words and informative reviews are given high attention weights. Moreover, a user auxiliary network is proposed, which regards the reviews written by kindred spirits who have a trust relationship with the user as auxiliary reviews, and effectively extracts the user’s auxiliary review features, thereby achieving more accurate user modeling to improve the recommendation performance. Extensive experiments are conducted on four real-world datasets, and the results show that the performance of the proposed model is better than that of baselines, which verifies the effectiveness of ARPCNN.
Zhe Li 0026, Honglong Chen, Zhichen Ni, Xiaogang Deng, Baodi Liu, Weifeng Liu 0001
IEEE Trans. Ind. Informatics6
2023 Egocentric Early Action Prediction via Adversarial Knowledge Distillation
abstract
Egocentric early action prediction aims to recognize actions from the first-person view by only observing a partial video segment, which is challenging due to the limited context information of the partial video. In this article, to tackle the egocentric early action prediction problem, we propose a novel multi-modal adversarial knowledge distillation framework. In particular, our approach involves a teacher network to learn the enhanced representation of the partial video by considering the future unobserved video segment, and a student network to mimic the teacher network to produce the powerful representation of the partial video and based on that predicting the action label. To promote the knowledge distillation between the teacher and the student network, we seamlessly integrate adversarial learning with latent and discriminative knowledge regularizations encouraging the learned representations of the partial video to be more informative and discriminative toward the action prediction. Finally, we devise a multi-modal fusion module toward comprehensively predicting the action label. Extensive experiments on two public egocentric datasets validate the superiority of our method over the state-of-the-art methods. We have released the codes and involved parameters to benefit other researchers. 1
Xuemeng Song, Weifeng Liu 0001, Yan Yan 0002, Liqiang Nie
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Semantic-Preserving Adversarial Text Attacks
abstract
Deep learning models are known immensely brittle to adversarial text examples. Existing text adversarial attack strategies can be roughly divided into character-level, word-level, and sentence-level attacks. Despite the success brought by recent text attack methods, how to induce misclassification with minimal text modifications while keeping the lexical correctness, syntactic soundness, and semantic consistency is still a challenge. In this paper, we devise a Bigram and Unigram-based adaptive Semantic Preservation Optimization (BU-SPO) approach which attacks text documents not only at a unigram word level but also at a bigram level to avoid generating meaningless sentences. We also present a hybrid attack strategy that collects substitution words from both synonyms and sememe candidates, to enrich the potential candidate set. Besides, a Semantic Preservation Optimization (SPO) method is devised to determine the word substitution priority and reduce the perturbation cost. Furthermore, we constrain the SPO with a semantic Filter (dubbed SPOF) to improve the semantic similarity. To estimate the effectiveness of our proposed methods, BU-SPO and BU-SPOF, we attack four victim deep learning models trained on three text datasets. Experimental results demonstrate that our approaches accomplish the highest semantics consistency and attack success rates by making minimal word modifications compared with competitive methods.
Xinghao Yang, Yongshun Gong, Weifeng Liu 0001, James Bailey 0001, Dacheng Tao, Wei Liu 0007
IEEE Trans. Sustain. Comput.3
2022 On the Complementarity between Pre-Training and Random-Initialization for Resource-Rich Machine Translation
abstract
Pre-Training (PT) of text representations has been successfully applied to low-resource Neural Machine Translation (NMT). However, it usually fails to achieve notable gains (some- times, even worse) on resource-rich NMT on par with its Random-Initialization (RI) counterpart. We take the first step to investigate the complementarity between PT and RI in resource-rich scenarios via two probing analyses, and find that: 1) PT improves NOT the accuracy, but the generalization by achieving flatter loss landscapes than that of RI; 2) PT improves NOT the confidence of lexical choice, but the negative diversity by assigning smoother lexical probability distributions than that of RI. Based on these insights, we propose to combine their complementarities with a model fusion algorithm that utilizes optimal transport to align neurons between PT and RI. Experiments on two resource-rich translation benchmarks, WMT’17 English-Chinese (20M) and WMT’19 English-German (36M), show that PT and RI could be nicely complementary to each other, achieving substantial improvements considering both translation accuracy, generalization, and negative diversity. Probing tools and code are released at: https://github.com/zanchangtong/PTvsRI.
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yu Cao 0014, Weifeng Liu 0001, Dacheng Tao
COLING5
2022 Enrich Features for Few-Shot Point Cloud Classification
abstract
Recently, many existing fully supervised methods for point cloud classification have strongly promoted the development of point cloud learning. However, these methods require a lot of labeled data as support, which is challenging to obtain. To alleviate this problem, we propose a novel few-shot point cloud classification method to classify new categories given a few labeled samples. Specifically, we apply the feature supplement module to enrich the geometric information of points and then aggregate multi-scale features through the channel-wise attention module while reducing the computational complexity. Finally, we introduce a classifier to classify the point cloud features under the few-shot learning setup to predict its label. We carry out experimental verification on the benchmark dataset and achieve state-of-the-art performance.
Hengxin Feng, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu
ICASSP2
2022 Agcyclegan: Attention-Guided Cyclegan for Single Underwater Image Restoration
abstract
Underwater image restoration is a fundamental problem in image processing and computer vision. It has broad application prospects for underwater operations, especially underwater robot operations. The challenging work is how to keep the color authenticity of the captured underwater image. In this paper, we propose a novel network architecture based on CycleGAN. Specifically, in the generator part, we adopt the U-Net structure because the long skip connection of U-Net will obtain more detailed information. Besides, we append the pixel-level attention block to provide greater flexibility for detail structure modeling. It assigns different weights to each channel to pay more attention to the critical feature. We also verify its generalization performance on several benchmark datasets. The extensive experiments with comparisons to state-of-the-art approaches demonstrate the superiority of the proposed model.
Zhenlong Wang, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu
ICASSP2
2022 MSL-FER: Mirrored Self-Supervised Learning for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) in the wild is a significant yet challenging classification task due to the inter-class similarities and intra-class variations. Recently, a large number of methods can extract expression features effectively. However, the intra-class variations mainly caused by various uncertainties (such as identity, pose, and occlusion) are difficult to capture in advance, and the cost of labeling these uncertainties is high. To tackle this challenge, we propose a novel Mirrored Self-supervised Learning FER (MSL-FER) method. The ground truth of self-supervised learning comes from the data itself rather than from human annotations, and horizontal inversion preserves emotional information without altering the facial structure. Specifically, MSL-FER introduces a binary classification task to recognize the 2D mirror operation in a self-supervised learning method. And we also combine our MSL-FER with an attention network to discriminate features along its dimensions selectively. Experiments on two public wild FER datasets show that our MSL-FER approach outperforms the baseline and other state-of-the-art methods with 87.92% on RAF-DB and 70.68% on FER2013.
Xiangshuai Pan, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu
ICIP2
2022 Image Super-Resolution Based on Adaptive Feature Fusion Channel Attention
Qizhang Song, Baodi Liu, Weifeng Liu 0001
ICONIP (3)3
2022 Virtual Try-on via Matching Relation with Landmark
Xingxing Yao, Baodi Liu, Weifeng Liu 0001
ICONIP (3)5
2022 EMAS: Efficient Meta Architecture Search for Few-Shot Learning
abstract
With the progress of few-shot learning, it has been scaled to many domains in the real world which have few labeled data, such as image classification and object detection. Many efforts for data embedding and feature combination have been made by designing a fixed neural architecture that can also be extended to variable and adaptive neural architectures for better performance. Recent works leverage neural architecture search technique to automatically design networks for few-shot learning but it requires vast computation costs and GPU memory requirements. This work introduces EMAS, an efficient method to speed up the searching process for few-shot learning. Specifically, we build a supernet to combine all candidate operations and then adopt gradient-based methods to search. Instead of training the whole supernet, we adopt Gumbel reparameterization technique to sample and activate a small subset of operations. EMAS handles a single path in a novel task adapted with just a few steps and time. A novel task only needs to learn fewer parameters and compute less content. During meta-testing, the task can well adapt to the network architecture although only with a few iterations. Empirical results show that EMAS yields a fair improvement in accuracy on the standard few-shot classification benchmark and is five times smaller in time.
Dongkai Liu, Honglong Chen, Baodi Liu, Weifeng Liu 0001
ICTAI6
2022 Automated Drawing Psychoanalysis via House-Tree-Person Test
abstract
The increase of human psychological illness in today's fast paced and high stress world makes it essential to detect the warning signals of psychological problems. As the most representative drawing psychoanalysis method, House-Tree-Person (HTP) test is widely used in psychological assessment with the benefit of simplicity, non-verbal, and repeatability. HTP test can reveal the individual subconscious of the psychological state through the picture content of house, tree, and person drawn by the patient. Currently, HTP test is conducted by the therapist in person, which makes it time consuming and the results are mostly affected by the therapist's experience. Therefore, it is helpful and necessary to build an automated method to improve the objectivity, reliability, and efficiency of HTP test. In this paper, we propose an automated psychometric drawing screening method that forms the relationship between the psychological state and drawing feature. Specifically, we extract the key features including size, position, and shadow of the drawing, and then combine these features to construct a psychological state classifier. The proposed method can effectively screen out negative drawings for further diagnosis and treatment. Experiments are carried out on a builded dataset with the drawings from a psychological testing center of college. Experimental results demonstrate the effect and superiority of the proposed method.
Baodi Liu, Weifeng Liu 0001
ICTAI4
2022 Where Does the Performance Improvement Come From?: - A Reproducibility Concern about Image-Text Retrieval
abstract
This article aims to provide the information retrieval community with some reflections on recent advances in retrieval learning by analyzing the reproducibility of image-text retrieval models. Due to the increase of multimodal data over the last decade, image-text retrieval has steadily become a major research direction in the field of information retrieval. Numerous researchers train and evaluate image-text retrieval algorithms using benchmark datasets such as MS-COCO and Flickr30k. Research in the past has mostly focused on performance, with multiple state-of-the-art methodologies being suggested in a variety of ways. According to their assertions, these techniques provide improved modality interactions and hence more precise multimodal representations. In contrast to previous works, we focus on the reproducibility of the approaches and the examination of the elements that lead to improved performance by pretrained and nonpretrained models in retrieving images and text.
Jun Rao, Fei Wang 0032, Liang Ding 0006, Shuhan Qi, Yibing Zhan, Weifeng Liu 0001, Dacheng Tao
SIGIR6
2022 MVFF: Multi-view Feature Fusion for Few-shot Remote Sensing Image Scene Classification
abstract
Compared to deep learning methods, few-shot learning methods do not need many labeled images. Therefore, few-shot remote sensing image scene classification has been studied extensively. However, obtaining effect information from the limited amount of labeled samples is a great challenge. Most methods only extract features from a single perspective of remote sensing images. Such information is scarce and even misleading. To address the problem, we propose a multi-view feature fusion (MVFF) method. Specifically, first, train two feature extractor networks on the original image dataset and remote sensing image dataset, respectively. And each model extracts features before and after average pooling, and we obtain four kinds of features for a remote sensing image. Second, we calculate fusion weights from support set features using Multi-Head Feature Collaboration (MHFC) method and four classifiers. Third, we utilize the weights to fuse predictive probability matrices and thus obtain the labels of query set samples. We implement experiments on three benchmark remote-sensing image datasets to validate the performance of our method. And the results demonstrate that our approach effectively handles few-shot remote sensing image scene classification.
Anxun Han, Lei Xing 0005, Weifeng Liu 0001, Baodi Liu
SMC3
2022 Multi-task Facial Expression Recognition With Joint Gender Learning
abstract
Facial Expression Recognition (FER) in the wild is a significant yet challenging topic in computer vision due to the feature inconsistency caused by the individual specificity of facial expressions. In addition to variations of facial expressions caused by identity, pose, and occlusion, gender also affects the face of human emotions. Even though males, females, and infants share the same facial expressions, their characteristics are vastly different. To capture the effect of gender on facial expressions, we propose a novel multi-task FER method with joint gender learning. First, in addition to the original emotion labels of face images, we annotate gender labels, including male, female, and infant. Second, we introduce a gender-aware multi-task convolutional neural network for FER, which can learn the emotion and gender features of faces. Compared with single-task expression recognition methods, our proposed framework for introducing gender feature learning can significantly achieve higher performance on FER in the wild. Finally, we verify the effectiveness of our framework on two public wild FER datasets, RAF-DB and FER2013. And the results show that the gender learning auxiliary task is beneficial to the improvement of the performance of FER.
Xiangshuai Pan, Qingtao Xie, Weifeng Liu 0001, Baodi Liu
SMC3
2022 Multi-relational Semantic Distillation for Few-Shot Object Detection
abstract
While few-shot object detection(FSOD) has been developed to a certain extent, it is still a large margin from practical applications. Most existing methods use traditional object detection methods as the basic framework is improved to a limited extent. Previous methods often ignore the special characterization relationship between support and query images. This paper fully investigates the effect of support images on detection performance and proposes a new FSOD method called Multi-relational Semantic Distillation (MSD). Our approach aims to improve FSOD performance by building a multi-relational semantic representation model with support and query features. In addition, we propose a support enhancement (SE) module based on the self-attention mechanism to enhance the useful information in the support features to mitigate the negative impact of low-quality support images. To verify the effectiveness of MSD, we conduct sufficient experiments on Pascal VOC and MS-COCO datasets. Experiments show that MSD achieves competitive results at low shots compared to other state-of-the-art few-shot detectors.
Qingtao Xie, Xiangshuai Pan, Weifeng Liu 0001, Baodi Liu
SMC3
2022 A graph convolutional neural network model with Fisher vector encoding and channel-wise spatial-temporal aggregation for skeleton-based action recognition
abstract
Abstract Skeleton‐based action recognition is an inspired yet challenging task in computer vision. Recently, the latest graph convolutional network (GCN), which generalises well‐established convolutional neural networks to non‐Euclidean structures, is proven to be highly successful for action recognition from body skeleton data. However, the GCN architecture has not been fully studied. In this work, a Fisher vector (FV) encoding based GCN architecture (FV‐GCN) is proposed, which exceeds the limitations of existing GCN‐based methods by combining the GCN model with FV encoding. A channel‐wise spatial–temporal aggregation function to preserve spatial–temporal information in the whole action clip and integrate it into the FV‐GCN architecture is also presented. Since FV is different from the GCN structure, this hybrid architecture that incorporates the advantages of both algorithms can discover complementary information of feature representation effectively. On two challenging human action datasets, kinetics, and NTU‐RGBD, improved performance is demonstrated over the baseline method, and the FV‐GCN is better or comparable to some state‐of‐the‐art methods.
Yanjiang Wang 0001, Sichao Fu, Baodi Liu, Weifeng Liu 0001
IET Image Process.5
2022 Object re-identification with distribution corrected ranking list
Dongchen Han, Shuai Shao 0006, Weifeng Liu 0001, Baodi Liu
Neurocomputing3
2022 Multi-view learning for hyperspectral image classification: An overview
Baodi Liu, Kai Zhang 0029, Honglong Chen, Weijia Cao, Weifeng Liu 0001, Dapeng Tao
Neurocomputing6
2022 DLDL: Dynamic label dictionary learning via hypergraph regularization
Shuai Shao 0006, Rui Xu 0012, Zhenfang Wang, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu
Neurocomputing4
2022 Learning task-specific discriminative embeddings for few-shot image classification
Lei Xing 0005, Shuai Shao 0006, Weifeng Liu 0001, Anxun Han, Xiangshuai Pan, Baodi Liu
Neurocomputing3
2022 Adaptive graph convolutional collaboration networks for semi-supervised classification
Sichao Fu, Senlin Wang, Weifeng Liu 0001, Baodi Liu, Xinhua You, Qinmu Peng, Xiaoyuan Jing
Inf. Sci.3
2022 Adaptive multi-scale transductive information propagation for few-shot learning
Sichao Fu, Baodi Liu, Weifeng Liu 0001, Bin Zou 0002, Xinhua You, Qinmu Peng, Xiaoyuan Jing
Knowl. Based Syst.3
2022 U-Shaped Attention Connection Network for Remote-Sensing Image Super-Resolution
abstract
In recent years, deep learning-based remote-sensing image super-resolution (SR) methods have made significant progress, and these methods require a large number of synthetic data for training. To obtain sufficient training data, researchers often generate synthetic data via fixed bicubic downsampling methods. However, the synthesized data cannot reflect the complex degradation process of real remote-sensing images. Thus, performance will dramatically reduce when these methods work in real low-resolution (LR) remote-sensing images. This letter proposes a U-shaped attention connection network (US-ACN) for remote-sensing image SR to solve this issue. Our US-ACN does not rely on any synthetic external dataset for training and merely requires one LR image to complete the training. The US-ACN utilizes remote-sensing images’ strong internal feature repetitiveness and fully learns this internal repetitive feature through a well-designed US-ACN to achieve the remote-sensing image SR. In addition, we design a 3-D attention module to generate effective 3-D weights by modeling channel and spatial attention weights, which is more helpful for the learning of internal features. Through the U-shaped connection among attention modules, context information propagation and attention weights learning are fully utilized. Many experiments show that our US-ACN adequately adapts to the remote-sensing image SR in various situations and performs advanced performance.
Wenzong Jiang, Lifei Zhao, Yanjiang Wang 0001, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.4
2022 Multiorder Interaction Information Embedding-Based Multiview Fusion-Aided Hyperspectral Image Classification
abstract
Hyperspectral images (HSI) are obtained from hyperspectral imaging sensors, which capture information in hundreds of spectral bands of objects. However, how to take full advantage of spatial and spectral information from many spectral bands to improve the performance of HSI classification remains an open question. Many HSI classification works have recently been reported by employing multi-view learning (MVL) algorithms that can fully use complementary information between different view features and thus have received widespread attention. This paper proposes a multi-view fusion network based on multi-order interaction information embedding for HSI classification. Firstly, the correlation matrix between spectral bands is used to divide the original data into multiple subsets as local views. The subset after the Segmented-PCA process is used as the global view. Secondly, the features of different views are extracted separately using a feature extraction network and mapped to the same dimension. Pre-fusion is achieved by multi-order interaction of various view features. Finally, loss-weighted fusion is applied to each view according to its contribution to the classification task. To evaluate the effectiveness of the proposed method, complete experiments were conducted on three commonly used HSI datasets, namely Pavia University, Houston 2013, and Houston 2018. The experimental results demonstrate that the proposed method improves the classification performance of existing feature extraction networks and is more competitive with other methods in the field.
Weijia Cao, Kai Zhang 0029, Baodi Liu, Dapeng Tao, Weifeng Liu 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 Location Soft-Aggregation-Based Band Weighting for Hyperspectral Image Classification
abstract
Hyperspectral images (HSIs) comprise hundreds of continuous spectral bands. How to effectively exploit the abundant spectral features of HSI to improve its classification accuracy is the focus of the research. Band weighting (BW) is extensively used due to its ability to emphasize usefully and suppress noisy bands adaptively. Most proposed works aggregate global information to construct band representation vectors in simple ways such as global averaging pooling. Those ways are not capable of retaining a more discriminating feature. Furthermore, modeling for interpixel positional relationships is something they have not considered. To address these problems, we propose a position embedding and importance aggregation BW module. The position embedding section encodes the position information by two 1-D features so that remote dependencies in one spatial direction can be obtained while retaining accurate position information in the other spatial direction. The importance aggregation section aggregates the global information. Finally, a group of weights is learned to recalibrate the raw input. Experiments on three public datasets of HSI demonstrate that our methods obtain competitive results compared to other methods.
Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Rethinking Few-Shot Remote Sensing Scene Classification: A Good Embedding Is All You Need?
abstract
In recent years, few-shot remote sensing scene classification (FSRSSC) has attracted more and more attention. For FSRSSC, most methods currently focus on designing a meta-learning algorithm, which obtains meta-knowledge from limited samples and then applies it to novel tasks. In this work, on the one hand, we optimize the training pipeline of the feature extractor; on the other hand, we apply a novel model fusion method further to optimize the feature extractor capability of the feature extractor. We show a novel few-shot remote sensing scene classification baseline: learning two feature representations through using two self-supervised methods on the meta-training set and then fusing the two representations into one. Then, training a linear classifier on this representation achieves state-of-the-art performance. It shows that training a good feature extractor can be more efficient than complex meta-learning algorithms for FSRSSC. We believe that our results can inspire a rethinking of few-shot remote sensing scene classification benchmarks.
Lei Xing 0005, Yuteng Ma, Weijia Cao, Shuai Shao 0006, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.5
2022 Learning to Cooperate: Decision Fusion Method for Few-Shot Remote-Sensing Scene Classification
abstract
Recently, remote-sensing scene classification has become an essential primary research topic. Nowadays, scholars have proposed various few-shot remote-sensing scene classification methods to achieve superior performance with few labeled data. Most of the prior work utilized a meta-learning strategy, which suffered from too little data affecting performance. In this letter, we apply the pre-trained feature extractor for image embedding. Meanwhile, because of the negative transfer problem caused by the inadaptability of the pre-trained feature extractor to remote-sensing data, we propose to exploit two pre-trained models to classify the remote-sensing scene, respectively. Then we fuse the decision to obtain the final classification category. We design a decision attention module to automatically update combination weights for each decision. It comprehensively considers the contribution of various decisions and further improves the discrimination of features. We conduct comprehensive experiments to validate the method and achieve state-of-the-art performance on two benchmark remote-sensing scene datasets, namely NWPU-RESISC45 and UC Merced.
Lei Xing 0005, Shuai Shao 0006, Yuteng Ma, Yanjiang Wang 0001, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.5
2022 Class Shared Dictionary Learning for Few-Shot Remote Sensing Scene Classification
abstract
In the field of remote sensing, it is infeasible to collect a large number of labeled samples due to imaging equipment and the imaging environment. Few-Shot Learning (FSL) is the dominant method to alleviate this problem, which pursues quickly adapting to novel categories from a limited number of labeled samples. The few-shot Remote Sensing Scene Classification (RSSC) generally includes the pre-training and meta-test phases. However, a “negative transfer” problem exists that data categories in both phases are different. It causes the pre-trained feature extractor to be unable well-adapted to the novel data category. This paper proposes Class Shared Dictionary Learning for Few-Shot Remote Sensing Scene Classification (CSDL) to address this issue. Specifically, this paper designs the Mirror-based Feature Extractor (MFE) in the pre-training phase, constructing a self-supervised classification task to improve the feature extractor robustness. Furthermore, this paper proposes a Class Shared Dictionary classifier (CSD) based on dictionary learning. The CSD projects the novel data feature in meta-test into subspace to reconstruct more discriminative features and complete the classification task. Extensive experiments on remote sensing datasets have demonstrated that the proposed CSDL achieves the advanced classification performance.
Lei Xing 0005, Lifei Zhao, Weijia Cao, Xinmin Ge, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.5
2022 DMH-FSL: Dual-Modal Hypergraph for Few-Shot Learning
Rui Xu 0012, Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001
Neural Process. Lett.5
2022 Co-Learning for Few-Shot Learning
Rui Xu 0012, Lei Xing 0005, Shuai Shao 0006, Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001
Neural Process. Lett.6
2022 MDFM: Multi-Decision Fusing Model for Few-Shot Learning
abstract
In recent years, researchers pay growing attention to the few-shot learning (FSL) task to address the data-scarce problem. A standard FSL framework is composed of two components: i) Pre-train. Employ the base data to generate a CNN-based feature extraction model (FEM). ii) Meta-test. Apply the trained FEM to the novel data (category is different from base data) to acquire the feature embeddings and recognize them. Although researchers have made remarkable breakthroughs in FSL, there still exists a fundamental problem. Since the trained FEM with base data usually cannot adapt to the novel class flawlessly, the novel data’s feature may lead to the distribution shift problem. To address this challenge, we hypothesize that even if most of the decisions based on different FEMs are viewed asweak decisions, which are not available for all classes, they still perform decent in some specific categories. Inspired by this assumption, we propose a novel method Multi-Decision Fusing Model (MDFM), which comprehensively considers the decisions based on multiple FEMs to enhance the efficacy and robustness of the model. MDFM is a simple, flexible, non-parametric method that can directly apply to the existing FEMs. Besides, we extend the proposed MDFM to two FSL settings (e.g., supervised and semi-supervised settings). We evaluate the proposed method on five benchmark datasets and achieve significant improvements of 3.4%-7.3% compared with state-of-the-arts.
Shuai Shao 0006, Lei Xing 0005, Rui Xu 0012, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu
IEEE Trans. Circuits Syst. Video Technol.4
2022 GCT: Graph Co-Training for Semi-Supervised Few-Shot Learning
abstract
Few-shot learning (FSL), purposing to resolve the problem of data-scarce, has attracted considerable attention in recent years. A popular FSL framework contains two phases: (i) the pre-train phase employs the base data to train a CNN-based feature extractor. (ii) the meta-test phase applies the frozen feature extractor to novel data (novel data has different categories from base data) and designs a classifier for recognition. To correct few-shot data distribution, researchers propose Semi-Supervised Few-Shot Learning (SSFSL) by introducing unlabeled data. Although SSFSL has been proved to achieve outstanding performances in the FSL community, there still exists a fundamental problem: the pre-trained feature extractor cannot adapt to the novel data flawlessly due to the cross-category setting. Usually, large amounts of noises are introduced to the novel feature. We dub it as Feature-Extractor-Maladaptive (FEM) problem. To tackle FEM, we make two efforts in this paper. First, we propose a novel label prediction method, Isolated Graph Learning (IGL). IGL introduces the Laplacian operator to encode the raw data to graph space, which helps reduce the dependence on features when classifying, and then project graph representation to label space for prediction. The key point is that: IGL can weaken the negative influence of noise from the feature representation perspective, and is also flexible to independently complete training and testing procedures, which is suitable for SSFSL. Second, we propose Graph Co-Training (GCT) to tackle this challenge from a multi-modal fusion perspective by extending the proposed IGL to the co-training framework. GCT is a semi-supervised method that exploits the unlabeled samples with two modal features to crossly strengthen the IGL classifier. We estimate our method on five benchmark few-shot learning datasets and achieve outstanding performances compared with other state-of-the-art methods. It demonstrates the effectiveness of our GCT.
Rui Xu 0012, Lei Xing 0005, Shuai Shao 0006, Lifei Zhao, Baodi Liu, Weifeng Liu 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.6
2022 Tensor Canonical Correlation Analysis Networks for Multi-View Remote Sensing Scene Recognition
abstract
Convolutional neural network (CNN) has been proven an effective way to extract high-level features from remote sensing (RS) images automatically. Many variants of the CNN model have been proposed, including principal component analysis network (PCANet), canonical correlation analysis network (CCANet), multiple scale CCANet (MS-CCANet) and multiview CCANet (MCCANet). The PCANet is specialized for single view feature abstraction, while in many real-world practices, the RS data are frequently observed from many more views. Although CCANet, MS-CCANet and MCCANet can be applied to two or more view data, they consider only the pair-wise correlation by calculating a series oftwo-ordercovariance matrices. However, the high-order consistence, which can only be explored by collectively and simultaneously examining all views, remains undiscovered. In this paper, we propose the tensor canonical correlation analysis network (TCCANet) to tackle this problem. Particularly, TCCANet learns filter banks by simultaneously maximizing arbitrary number of views with high-order-correlation and solves the optimization problem by decomposing a covariance tensor. After the convolutional stage, we utilize binarization and block-wise histogram strategies to generate the final feature. Furthermore, we also develop a Multiple Scale version of TCCANet, i.e., MS-TCCANet, to extract enriched representation of the RS data by incorporating all previous convolutional layers. Numerical experiment results on RSSCN7 and SAT-6 datasets demonstrate the advantages of TCCANet and MS-TCCANet for RS scene recognition.
Xinghao Yang, Weifeng Liu 0001, Wei Liu 0007
IEEE Trans. Knowl. Data Eng.2
2022 Learning Representation on Optimized High-Order Manifold for Visual Classification
abstract
Graph convolutional networks (GCNs) and graph neural networks (GNNs) have demonstrated convincing performance on many tasks by learning the intrinsic structure of the data. However, it is still valuable and challenging to consider the complex and complete correlations of objects, i.e., high-order manifold structures, for representation learning. In this paper, we present a novel representation learning method that utilizes the optimized high-order manifold of the data for classification tasks of nonstructural data and graph-structure data. In the method, we fully explore the complicated relationship of samples by highlighting the high-order manifold information in a hypergraph. Specifically, we incorporate high-order manifold information by graph$p$-Laplacian into a hypergraph and propose$p$-Laplacian-based hypergraph neural networks (pLapHGNN) to significantly learn hidden layer representations that encode both the high-order structure of data and the high-order manifold geometrical information. Confronting the difficulties of obtaining optimized high-order manifolds of the data, we propose an effective approximate approach by graph$p$-Laplacian representing the relationship of hyperedges in the hypergraph. Furthermore, we study the weights of hyperedges in a hypergraph with high-order manifold information. Experiments on the ModelNet40 dataset and NTU dataset demonstrate that the proposed method is more effective than the other popular methods for 3D shape recognition. Extensive experiments on other visual classification tasks and citation networks also show the superiority of our proposed method for representation learning.
Xueqi Ma, Weifeng Liu 0001, Qi Tian 0001, Yue Gao 0002
IEEE Trans. Multim.2
2022 Domain-invariant Graph for Adaptive Semi-supervised Domain Adaptation
abstract
Domain adaptation aims to generalize a model from a source domain to tackle tasks in a related but different target domain. Traditional domain adaptation algorithms assume that enough labeled data, which are treated as the prior knowledge are available in the source domain. However, these algorithms will be infeasible when only a few labeled data exist in the source domain, thus the performance decreases significantly. To address this challenge, we propose a Domain-invariant Graph Learning (DGL) approach for domain adaptation with only a few labeled source samples. Firstly, DGL introduces the Nyström method to construct a plastic graph that shares similar geometric property with the target domain. Then, DGL flexibly employs the Nyström approximation error to measure the divergence between the plastic graph and source graph to formalize the distribution mismatch from the geometric perspective. Through minimizing the approximation error, DGL learns a domain-invariant geometric graph to bridge the source and target domains. Finally, we integrate the learned domain-invariant graph with the semi-supervised learning and further propose an adaptive semi-supervised model to handle the cross-domain problems. The results of extensive experiments on popular datasets verify the superiority of DGL, especially when only a few labeled source samples are available.
Weifeng Liu 0001, Yicong Zhou, Jun Yu 0002, Dapeng Tao, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Answer Questions with Right Image Regions: A Visual Attention Regularization Approach
abstract
Visual attention in Visual Question Answering (VQA) targets at locating the right image regions regarding the answer prediction, offering a powerful technique to promote multi-modal understanding. However, recent studies have pointed out that the highlighted image regions from the visual attention are often irrelevant to the given question and answer, leading to model confusion for correct visual reasoning. To tackle this problem, existing methods mostly resort to aligning the visual attention weights with human attentions. Nevertheless, gathering such human data is laborious and expensive, making it burdensome to adapt well-developed models across datasets. To address this issue, in this article, we devise a novel visual attention regularization approach, namely, AttReg, for better visual grounding in VQA. Specifically, AttReg first identifies the image regions that are essential for question answering yet unexpectedly ignored (i.e., assigned with low attention weights) by the backbone model. And then a mask-guided learning scheme is leveraged to regularize the visual attention to focus more on these ignored key regions. The proposed method is very flexible and model-agnostic, which can be integrated into most visual attention-based VQA models and require no human attention supervision. Extensive experiments over three benchmark datasets, i.e., VQA-CP v2, VQA-CP v1, and VQA v2, have been conducted to evaluate the effectiveness of AttReg. As a by-product, when incorporating AttReg into the strong baseline LMH, our approach can achieve a new state-of-the-art accuracy of 60.00% with an absolute performance gain of 7.01% on the VQA-CP v2 benchmark dataset. In addition to the effectiveness validation, we recognize that the faithfulness of the visual attention in VQA has not been well explored in literature. In the light of this, we propose to empirically validate such property of visual attention and compare it with the prevalent gradient-based approaches.
Yibing Liu, Jianhua Yin 0001, Xuemeng Song, Weifeng Liu 0001, Liqiang Nie, Min Zhang 0005
ACM Trans. Multim. Comput. Commun. Appl.5
2021 Bigram and Unigram Based Text Attack via Adaptive Monotonic Heuristic Search
abstract
Deep neural networks (DNNs) are known to be vulnerable to adversarial images, while their robustness in text classification are rarely studied. Several lines of text attack methods have been proposed in the literature, such as character-level, word-level, and sentence-level attacks. However, it is still a challenge to minimize the number of word distortions necessary to induce misclassification, while simultaneously ensuring the lexical correctness, syntactic correctness, and semantic similarity. In this paper, we propose the Bigram and Unigram based Monotonic Heuristic Search (BU-MHS) method to examine the vulnerability of deep models. Our method has three major merits. Firstly, we propose to attack text documents not only at the unigram word level but also at the bigram level to avoid producing meaningless outputs. Secondly, we propose a hybrid method to replace the input words with both their synonyms and sememe candidates, which greatly enriches potential substitutions compared to only using synonyms. Lastly, we design a search algorithm, i.e., Monotonic Heuristic Search (MHS), to determine the priority of word replacements, aiming to reduce the modification cost in an adversarial attack. We evaluate the effectiveness of BU-MHS on IMDB, AG's News, and Yahoo! Answers text datasets by attacking four state-of-the-art DNNs models. Experimental results show that our BU-MHS achieves the highest attack success rate by changing the smallest number of words compared with other existing models.
Xinghao Yang, Weifeng Liu 0001, James Bailey 0001, Dacheng Tao, Wei Liu 0007
AAAI2
2021 Leveraging GANs via Non-local Features
Xuyang Peng, Weifeng Liu 0001, Baodi Liu, Kai Zhang 0029, Yicong Zhou
ICANN (2)2
2021 Linked Attention-Based Dynamic Graph Convolution Module for Point Cloud Classification
abstract
With the rapid development of 3D technology, point cloud data is becoming more and more popular, which arouses researchers’ interest. But its properties – irregularity and disorder – make it difficult to analyze. In this work, we combine the attention module with the dynamic graph convolutional neural network to pay attention to the target’s critical part. Then, the modules are densely connected to guarantee that each layer is fully utilized. Finally, we carry out experiments on several benchmark datasets to verify the proposed model and achieve state-of-the-art performance.
Baodi Liu, Weifeng Liu 0001, Kai Zhang 0029
ICIP3
2021 Affine Non-Negative Collaborative Representation for Deep Metric Learning
Baodi Liu, Weifeng Liu 0001, Kai Zhang 0029
ICIP3
2021 Hazy Re-ID: An Interference Suppression Model for Domain Adaptation Person Re-Identification Under Inclement Weather Condition
abstract
In a conventional domain adaptation person Re-identification (Re-ID) task, both the training and test images in target domain are collected under the sunny weather. However, in reality, the pedestrians to be retrieved may be obtained under severe weather conditions such as hazy, dusty and snowing, etc. This paper proposes a novel Interference Suppression Model (ISM) to deal with the interference caused by the hazy weather in domain adaptation person Re-ID. A teacher-student model is used in the ISM to distill the interference information at the feature level by reducing the discrepancy between the clear and the hazy intrinsic similarity matrix. Furthermore, in the distribution level, the extra discriminator is introduced to assist the student model make the interference feature distribution more clear. The experimental results show that the proposed method achieves the superior performance on two synthetic datasets than the state-of-the-art methods. The related code will be released online https://github.com/pangjian123/ISM-ReID.
Jian Pang, Dacheng Zhang, Huafeng Li 0001, Weifeng Liu 0001, Zhengtao Yu 0001
ICME4
2021 BESA: BERT-based Simulated Annealing for Adversarial Text Attacks
abstract
Modern Natural Language Processing (NLP) models are known immensely brittle towards text adversarial examples. Recent attack algorithms usually adopt word-level substitution strategies following a pre-computed word replacement mechanism. However, their resultant adversarial examples are still imperfect in achieving grammar correctness and semantic similarities, which is largely because of their unsuitable candidate word selections and static optimization methods. In this research, we propose BESA, a BERT-based Simulated Annealing algorithm, to address these two problems. Firstly, we leverage the BERT Masked Language Model (MLM) to generate contextual-aware candidate words to produce fluent adversarial text and avoid grammar errors. Secondly, we employ Simulated Annealing (SA) to adaptively determine the word substitution order. The SA provides sufficient word replacement options via internal simulations, with an objective to obtain both a high attack success rate and a low word substitution rate. Besides, our algorithm is able to jump out of local optima with a controlled probability, making it closer to achieve the best possible attack (i.e., the global optima). Experiments on five popular datasets manifest the superiority of BESA compared with existing methods, including TextFooler, BAE, BERT-Attack, PWWS, and PSO.
Xinghao Yang, Weifeng Liu 0001, Dacheng Tao, Wei Liu 0007
IJCAI2
2021 Collaborative Representation for Deep Meta Metric Learning
abstract
Most metric learning methods utilize all training data to construct a single metric, and it is usually over-fitting on the "salient" feature. To overcome this issue, we propose a deep meta metric learning method based on collaborative representation. We construct multiple episodes from the original training data to train a general metric, where each episode consists of a query set and a support set. Then, we introduce a collaborative representation method, which fits the query sample with the support samples per class. We predict the query sample's label via the optimal fitness among the query sample and the support samples in each specific class. Besides, we adopt a hard mining strategy to learn a more discriminative metric according to increasing the training tasks' difficulty. Experiments verify that our method achieves state-of-the-art results on three re-ID benchmark datasets.
Weifeng Liu 0001, Kai Zhang 0029, Baodi Liu
ICMR2
2021 Adaptive Eigenmodes for Robust Object Tracking
abstract
Discriminative correlation filters based algorithms have attracted extensive attention due to their strong tracking capability. However, object tracking still faces many challenges due to object appearance variations, background clutter, occlusion, plane rotation, etc. In this paper, to better express the object, multiple features are integrated to make full use of the advantage of different features. Furthermore, the adaptive eigen-decomposition and reconstruction ("eigenmodes") method is applied to carry out the integrated-feature decomposition, and optimal expression of the object is established through simple eigen-relationships. It has proved experimentally that the predicted value after eigenmodes is closer to the groundtruth than before. Furthermore, to solve the tracking failure caused by interfering objects or background clutters and improve the tracking accuracy, the average peak-correlation energy (APCE) method is utilized as an optimized update strategy in this paper. A large number of experimental results on the known reference datasets indicate that our algorithm has good performance compared to the existing tracking methods.
Yujuan Qi, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001
SMC5
2021 From Objects to a Whole Painting
abstract
Style image painting is the process of using some stylized strokes to redraw a reference image purposefully and meaningfully. It is a kind of style image generation. In recent years, the application of GAN has greatly improved the quality of generated images for style image generation. However, those methods which use GAN are usually non-serialized. To solve this problem, reinforcement learning and RNN methods are applied to the generation of style images based on strokes. To speed up training, stroke-based style image painting using CNN framework is also proposed. But none of the existing style image painting methods takes into account the content distribution in the reference image, which makes the painting process lack a clear order. We extract the feature of the object contained in the image through the content acquisition module and use the optimal transmission theory to construct a compound loss function to integrate the content information into the painting process. As more advanced macro content information is added to the painting process, our painting method can draw images in a more orderly way. We have conducted a lot of experiments to prove that our method is superior to the current state-of-the-art style image painting method.
Fei Wang 0032, Baodi Liu, Weifeng Liu 0001
SMC3
2021 CNN-combined graph residual network with multilevel feature fusion for hyperspectral image classification
abstract
Abstract The application of graph convolutional networks (GCN) in hyperspectral image (HSI) classification has become a promising method, thanks to its flexible convolution operation in any irregular image region. For the classification of HSI, GCN can extract more superpixel‐level features with a topological structure, in comparison to the traditional convolutional neural networks (CNNs) using fixed square kernels distilling pixel‐level features. To fully leverage the different levels of features, this study proposes a novel deep network referred to as a CNN‐combined graph residual network (GRN), which integrates the multilevel graph residual module and spectral‐spatial features continuous learning module. During the extraction of topology information using the former module, HSI pixels are divided into superpixels and served as input nodes of the module to reduce the computational complexity and obtain the multilevel spatial relevance between adjacent superpixels. Besides, for the latter module, the spectral‐spatial features are learnt continuously, which could obtain the finer pixel‐level features. Finally, the captured spectral‐spatial features of different levels are concatenated. This strategy could not only adequately utilize the correlation and difference of adjacent spatial but also obtain the finer and more valuable spectral‐spatial information, which makes a significant boost in the HSI classification. Additionally, the experiment results demonstrate the superiority and availability of the GRN on three benchmark datasets of HSI, compared with the state‐of‐the‐art methods for the classification of HSI.
Wenhui Guo, Guixun Xu, Weifeng Liu 0001, Baodi Liu, Yanjiang Wang 0001
IET Comput. Vis.3
2021 Learn from Object Counting: Crowd Counting with Meta-learning
abstract
Abstract The objective of crowd counting is to learn a counter that can estimate the number of people in a single image. So far, most of the proposed work evaluates the crowd density by fitting the constructed density map corresponding to the sample. The performance of those algorithms depends on a large amount of carefully prepared data. However, a significant problem with crowd data sets is the difficulty of labeling. To address such a situation, utilizing object counting data in few‐shot scenes is considered and an efficient algorithm to extract the meta‐information is proposed, thus improving the accuracy and convergence rate of the crowd counting tasks. Specifically, the counting network is trained with only object counting tasks constructed on different domains during the meta‐training phase. Then, the meta‐counter is testing on crowd counting tasks in the meta‐testing stage. Experimentally, it is demonstrated that the above way improves the converge rate and accuracy of crowd counting tasks on three crowd counting datasets when meta‐training on ten‐type object counting tasks.
Changtong Zan, Baodi Liu, Weili Guan, Kai Zhang 0029, Weifeng Liu 0001
IET Image Process.5
2021 Example-feature graph convolutional networks for semi-supervised classification
Sichao Fu, Weifeng Liu 0001, Kai Zhang 0029, Yicong Zhou
Neurocomputing2
2021 Human activity recognition by manifold regularization based dynamic graph convolutional networks
Weifeng Liu 0001, Sichao Fu, Yicong Zhou, Zhengjun Zha, Liqiang Nie
Neurocomputing1
2021 Accurately modeling the human brain functional correlations with hypergraph Laplacian
Jichao Ma, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001
Neurocomputing4
2021 Targeted Attention Attack on Deep Learning Models in Road Sign Recognition
abstract
Real-world traffic sign recognition is an important step toward building autonomous vehicles, most of which highly dependent on deep neural networks (DNNs). Recent studies demonstrated that DNNs are surprisingly susceptible to adversarial examples. Many attack methods have been proposed to understand and generate adversarial examples, such as gradient-based attack, score-based attack, decision-based attack, and transfer-based attacks. However, most of these algorithms are ineffective in real-world road sign attack, because 1) iteratively learning perturbations for each frame is not realistic for a fast moving car and 2) most optimization algorithms traverse all pixels equally without considering their diverse contribution. To alleviate these problems, this article proposes the targeted attention attack (TAA) method for real-world road sign attack. Specifically, we have made the following contributions: 1) we leverage the soft attention map to highlight those important pixels and skip those zero-contributed areas-this also helps to generate natural perturbations; 2) we design an efficient universal attack that optimizes a single perturbation/noise based on a set of training images under the guidance of the pretrained attention map; 3) we design a simple objective function that can be easily optimized; and 4) we evaluate the effectiveness of TAA on real-world data sets. Experimental results validate that the TAA method improves the attack successful rate (nearly 10%) and reduces the perturbation loss (about a quarter) compared with the popular RP2 method. Additionally, our TAA also provides good properties, e.g., transferability and generalization capability. We provide code and data to ensure the reproducibility: https://github.com/AdvAttack/RoadSignAttack.
Xinghao Yang, Weifeng Liu 0001, Shengli Zhang 0001, Wei Liu 0007, Dacheng Tao
IEEE Internet Things J.2
2021 Semi-supervised classification by graph p-Laplacian convolutional networks
Sichao Fu, Weifeng Liu 0001, Kai Zhang 0029, Yicong Zhou, Dapeng Tao
Inf. Sci.2
2021 PaintNet: A shape-constrained generative framework for generating clothing from fashion model
Xuemeng Song, Tian Gan 0002, Yiyang Yao, Weifeng Liu 0001, Liqiang Nie
Multim. Tools Appl.5
2021 Unified Cross-domain Classification via Geometric and Statistical Adaptations
Weifeng Liu 0001, Baodi Liu, Weili Guan, Yicong Zhou, Changsheng Xu
Pattern Recognit.1
2021 MKEL: Multiple Kernel Ensemble Learning via Unified Ensemble Loss for Image Classification
abstract
In this article, a novel ensemble model, called Multiple Kernel Ensemble Learning (MKEL), is developed by introducing a unified ensemble loss. Different from the previous multiple kernel learning (MKL) methods, which attempt to seek a linear combination of basis kernels as a unified kernel, our MKEL model aims to find multiple solutions in corresponding Reproducing Kernel Hilbert Spaces (RKHSs) simultaneously. To achieve this goal, multiple individual kernel losses are integrated into a unified ensemble loss. Therefore, each model can co-optimize to learn its optimal parameters by minimizing a unified ensemble loss in multiple RKHSs. Furthermore, we apply our proposed ensemble loss into the deep network paradigm and take the sub-network as a kernel mapping from the original input space into a feature space, named Deep-MKEL (D-MKEL). Our D-MKEL model can utilize the diversified deep individual sub-networks into a whole unified network to improve the classification performance. With this unified loss design, our D-MKEL model can make our network much wider than other traditional deep kernel networks and more parameters are learned and optimized. Experimental results on several mediate UCI classification and computer vision datasets demonstrate that our MKEL model can achieve the best classification performance among comparative MKL methods, such as Simple MKL, GMKL, Spicy MKL, and Matrix-Regularized MKL. On the contrary, experimental results on large-scale CIFAR-10 and SVHN datasets concretely show the advantages and potentialities of the proposed D-MKEL approach compared to state-of-the-art deep kernel methods.
Xiangjun Shen, Kou Lu, Sumet Mehta, Weifeng Liu 0001, Jianping Fan 0001, Zhengjun Zha
ACM Trans. Intell. Syst. Technol.5
2021 A Survey on Canonical Correlation Analysis
abstract
In recent years, the advances in data collection and statistical analysis promotes canonical correlation analysis (CCA) available for more advanced research. CCA is the main technique for two-set data dimensionality reduction such that the correlation between the pairwise variables in the common subspace is mutually maximized. Over 80-years of developments, a number of CCA models have been proposed according to different machine learning mechanisms. However, the field lacks an insightful review for the state-of-art developments. This survey targets to provide a well-organized overview for CCA and its extensions. Specifically, we first review the CCA theory from the perspective of both model formation and model optimization. The association between two popular solution methods, i.e., eigen value decomposition (EVD) and singular value decomposition (SVD), are discussed. Following that, we present a taxonomy of current progresses and classify them into seven groups: 1) multi-view CCA, 2) probabilistic CCA, 3) deep CCA, 4) kernel CCA, 5) discriminative CCA, 6) sparse CCA and 7) locality preserving CCA. For each group, we demonstrate two or three representative mathematical models, identifying their strengths and limitations. We summarize the representative applications and numerical results of these seven groups in real-world practices, collecting the data sets and open-sources for implementation. In the end, we provide several promising future research directions that can improve the current state of the art.
Xinghao Yang, Weifeng Liu 0001, Wei Liu 0007, Dacheng Tao
IEEE Trans. Knowl. Data Eng.2
2021 Dynamic Graph Learning Convolutional Networks for Semi-supervised Classification
abstract
Over the past few years, graph representation learning (GRL) has received widespread attention on the feature representations of the non-Euclidean data. As a typical model of GRL, graph convolutional networks (GCN) fuse the graph Laplacian-based static sample structural information. GCN thus generalizes convolutional neural networks to acquire the sample representations with the variously high-order structures. However, most of existing GCN-based variants depend on the static data structural relationships. It will result in the extracted data features lacking of representativeness during the convolution process. To solve this problem, dynamic graph learning convolutional networks (DGLCN) on the application of semi-supervised classification are proposed. First, we introduce a definition of dynamic spectral graph convolution operation. It constantly optimizes the high-order structural relationships between data points according to the loss values of the loss function, and then fits the local geometry information of data exactly. After optimizing our proposed definition with the one-order Chebyshev polynomial, we can obtain a single-layer convolution rule of DGLCN. Due to the fusion of the optimized structural information in the learning process, multi-layer DGLCN can extract richer sample features to improve classification performance. Substantial experiments are conducted on citation network datasets to prove the effectiveness of DGLCN. Experiment results demonstrate that the proposed DGLCN obtains a superior classification performance compared to several existing semi-supervised classification models.
Sichao Fu, Weifeng Liu 0001, Weili Guan, Yicong Zhou, Dapeng Tao, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Urban Perception: Sensing Cities via a Deep Interactive Multi-task Learning Framework
abstract
Social scientists have shown evidence that visual perceptions of urban attributes, such as safe, wealthy, and beautiful perspectives of the given cities, are highly correlated to the residents’ behaviors and quality of life. Despite their significance, measuring visual perceptions of urban attributes is challenging due to the following facts: (1) Visual perceptions are subjectively contradistinctive rather than absolute. (2) Perception comparisons between image pairs are usually conducted region by region, and highly related to the specific urban attributes. And (3) the urban attributes have both the shared and specific information. To address these problems, in this article, we present a Deep inteRActive Multi-task leArning scheme, DRAMA for short. DRAMA comparatively quantifies the perceptions of urban attributes by jointly integrating the pairwise comparisons, regional interactions, and urban attribute correlations within a unified deep scheme. In DRAMA, each urban attribute is treated as a task, whereby the task-sharing and the task-specific information is fully explored. By conducting extensive experiments over a public large-scale benchmark dataset, it is demonstrated that our proposed DRAMA scheme outperforms several state-of-the-art baselines. Meanwhile, we applied the pairwise comparisons of our DRAMA model to further quantify the urban attributes and hence rank cities with respect to the given urban attributes. As a byproduct, we have released the codes and parameter settings to facilitate other researches.
Weili Guan, Zhaozheng Chen, Fuli Feng, Weifeng Liu 0001, Liqiang Nie
ACM Trans. Multim. Comput. Commun. Appl.4
2020 Local structure alignment guided domain adaptation with few source samples
abstract
Domain adaptation has received lots of attention for its high efficiency in dealing with cross-domain learning tasks. Most existing domain adaptation methods adopt the strategies relying on large amounts of source label information, which limits their applications in the real world where only a few label samples are available. We exploit the local geometric connections to tackle this problem and propose a Local Structure Alignment (LSA) guided domain adaptation method in this paper. LSA leverages the Nyström method to describe the distribution difference from the geometric perspective and then perform the distribution alignment between domains. Specifically, LSA constructs a domain-invariant Hessian matrix to locally connect the data of the two domains through minimizing the Nyström approximation error. And then it integrates the domain-invariant Hessian matrix with the semi-supervised learning and finally builds an adaptive semi-supervised model. Extensive experimental results validate that the proposed LSA outperforms the traditional domain adaptation methods especially when only sparse source label information is available.
Yuying Cai, Baodi Liu, Weifeng Liu 0001, Kai Zhang 0029, Changsheng Xu
MMAsia4
2020 Label embedded dictionary learning for image classification
Shuai Shao 0006, Rui Xu 0012, Weifeng Liu 0001, Baodi Liu, Yanjiang Wang 0001
Neurocomputing3
2020 HesGCN: Hessian graph convolutional networks for semi-supervised classification
Sichao Fu, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou, Liqiang Nie
Inf. Sci.2
2020 Domain Adaptation with Few Labeled Source Samples by Graph Regularization
Weifeng Liu 0001, Yicong Zhou, Dapeng Tao, Liqiang Nie
Neural Process. Lett.2
2020 Class specific or shared? A cascaded dictionary learning framework for image classification
Yanjiang Wang 0001, Shuai Shao 0006, Rui Xu 0012, Weifeng Liu 0001, Baodi Liu
Signal Process.4
2019 Two-order graph convolutional networks for semi-supervised classification
abstract
Currently, deep learning (DL) algorithms have achieved great success in many applications including computer vision and natural language processing. Many different kinds of DL models have been reported, such as DeepWalk, LINE, diffusionconvolutional neural networks, graph convolutional networks (GCN), and so on. The GCN algorithm is a variant of convolutional neural network and achieves significant superiority by using a one‐order localised spectral graph filter. However, only a one‐order polynomial in the Laplacian of GCN has been approximated and implemented, which ignores undirect neighbour structure information. The lack of rich structure information reduces the performance of the neural networks in the graph structure data. In this study, the authors deduce and simplify the formula of two‐order spectral graph convolutions to preserve rich local information. Furthermore, they build a layerwise GCN based on this two‐order approximation, i.e. two‐order GCN (TGCN) for semi‐supervised classification. With the two‐order polynomial in the Laplacian, the proposed TGCN model can assimilate abundant localised structure information of graph data and then boosts the classification significantly. To evaluate the proposed solution, extensive experiments are conducted on several popular datasets including the Citeseer, Cora, and PubMed dataset. Experimental results demonstrate that the proposed TGCN outperforms the state‐of‐art methods.
Sichao Fu, Weifeng Liu 0001, Yicong Zhou
IET Image Process.2
2019 Advances in data representation and learning for pattern analysis
C. L. Philip Chen, Xinge You, Xinbo Gao 0001, Tongliang Liu, Fionn Murtagh, Weifeng Liu 0001
Neurocomputing6
2019 HpLapGCN: Hypergraph p-Laplacian graph convolutional networks
Sichao Fu, Weifeng Liu 0001, Yicong Zhou, Liqiang Nie
Neurocomputing2
2019 Hessian-Regularized Multitask Dictionary Learning for Remote Sensing Image Recognition
abstract
Learning effective image representations is a vital issue for remote sensing (RS) image recognition tasks. Although numerous algorithms have been proposed, it is still challenging due to the limited labeled data. One representative work is the Laplacian-regularized multitask dictionary learning (LR-MTDL) that employs graph Laplacian regularization terms to fully utilize both the labeled and unlabeled information. However, it probably conduces to poor extrapolating power because Laplacian regularization biases the solution toward a constant function. In this letter, we propose a Hessian-regularized multitask dictionary learning to learn a source-data set-shared but target-data set-biased representation for RS image recognition. Particularly, Hessian can properly exploit the intrinsic local geometry of the data manifold and finally leverage the performance. Extensive experiments on four RS image data sets validate the effectiveness of the proposed method by comparing with baseline algorithms including single-task dictionary learning and LR-MTDL.
Guanhua Feng, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou
IEEE Geosci. Remote. Sens. Lett.2
2019 Effective human action recognition by combining manifold regularization and pairwise constraints
Xueqi Ma, Dapeng Tao, Weifeng Liu 0001
Multim. Tools Appl.3
2019 Hessian Regularized Distance Metric Learning for People Re-Identification
Guanhua Feng, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou
Neural Process. Lett.2
2019 $p$ -Laplacian Regularization for Scene Recognition
abstract
The explosive growth of multimedia data on the Internet makes it essential to develop innovative machine learning algorithms for practical applications especially where only a small number of labeled samples are available. Manifold regularized semi-supervised learning (MRSSL) thus received intensive attention recently because it successfully exploits the local structure of data distribution including both labeled and unlabeled samples to leverage the generalization ability of a learning model. Although there are many representative works in MRSSL, including Laplacian regularization (LapR) and Hessian regularization, how to explore and exploit the local geometry of data manifold is still a challenging problem. In this paper, we introduce a fully efficient approximation algorithm of graph p -Laplacian, which significantly saving the computing cost. And then we propose p -LapR (pLapR) to preserve the local geometry. Specifically, p -Laplacian is a natural generalization of the standard graph Laplacian and provides convincing theoretical evidence to better preserve the local structure. We apply pLapR to support vector machines and kernel least squares and conduct the implementations for scene recognition. Extensive experiments on the Scene 67 dataset, Scene 15 dataset, and UC-Merced dataset validate the effectiveness of pLapR in comparison to the conventional manifold regularization methods.
Weifeng Liu 0001, Xueqi Ma, Yicong Zhou, Dapeng Tao, Jun Cheng 0002
IEEE Trans. Cybern.1
2019 Hypergraph $p$ -Laplacian Regularization for Remotely Sensed Image Recognition
abstract
Graph-based and manifold-regularization (MR)-based semisupervised learning, including Laplacian regularization (LapR) and hypergraph LapR (HLapR), have achieved prominent performance in preserving locality and similarity information. However, it is still a great challenge to exactly explore and exploit the local structure of the data distribution. In this paper, we present an efficient and effective approximation algorithm of hypergraph${p}$-Laplacian and then propose hypergraph${p}$-LapR (HpLapR) to preserve the geometry of the probability distribution. In particular, hypergraph is a generalization of a standard graph while hypergraph${p}$-Laplacian is a nonlinear generalization of the standard graph Laplacian. The proposed HpLapR shows great potential to exploit the local structures. We integrate HpLapR with logistic regression for remote sensing image recognition. Experiments on UC-Merced data set demonstrate that the proposed HpLapR has superior performance compared with several popular MR methods including LapR and HLapR.
Xueqi Ma, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou
IEEE Trans. Geosci. Remote. Sens.2
2018 Biological modeling of human visual system for object recognition using GLoP filters and sparse coding on multi-manifolds
Limiao Deng, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001, Yujuan Qi
Mach. Vis. Appl.4
2018 Reinforcement online learning for emotion prediction by using physiological signals
Weifeng Liu 0001, Lianbo Zhang, Dapeng Tao, Jun Cheng 0002
Pattern Recognit. Lett.1
2018 Data mining in human activity analysis
Xinmei Tian 0001, Weifeng Liu 0001, Fionn Murtagh
Signal Process.2
2017 Multiple Scale Canonical Correlation Analysis Networks for Two-View Object Recognition
Xinghao Yang, Weifeng Liu 0001
ICONIP (1)2
2017 Canonical correlation analysis networks for two-view image recognition
Xinghao Yang, Weifeng Liu 0001, Dapeng Tao, Jun Cheng 0002
Inf. Sci.2
2017 Support vector machine active learning by Hessian regularization
Weifeng Liu 0001, Lianbo Zhang, Dapeng Tao, Jun Cheng 0002
J. Vis. Commun. Image Represent.1
2017 Dictionary Learning-Based Hough Transform for Road Detection in Multispectral Image
abstract
It is of great importance to determine the location and orientation of a straight road in multispectral images for remote sensing. One of the classical methods for straight line detection is the Hough transform that is widely used in binary images. Although there are many previous works for straight road detection, it is still in its infancy to extract a straight road in multispectral images for remote sensing. In this letter, we propose a multiview dictionary learning formulation to approximate the Hough transform for straight road detection in multispectral images. Our formulation can exploit the complementary among the multiple spectral channels. Furthermore, it is natural to incorporate regularizations of prior information to significantly leverage the performance. We consider L1-norm regularization as a case study and conduct extensive experiments on RSSCN7 data set to verify the proposed algorithm. The experimental results demonstrate the superiority of our method in comparison with traditional methods.
Weifeng Liu 0001, Zhenqing Zhang, Xinghua Chen, Yicong Zhou
IEEE Geosci. Remote. Sens. Lett.1
2017 Multiview Canonical Correlation Analysis Networks for Remote Sensing Image Recognition
abstract
In the past decade, deep learning (DL) algorithms have been widely used for remote sensing (RS) image recognition tasks. As the most typical DL model, convolutional neural networks (CNNs) achieves outstand performance for big RS data classification. Recently, a variant of CNN, dubbed canonical correlation analysis network (CCANet), was proposed to abstract the two-view image features. Extensive experiments conducted on several benchmark databases validate the effectiveness of CCANet. However, the CCANet structure is powerless when the observations arrive from more than two sources. To serve the multiview purpose, in this letter, we propose multiview CCANets (MCCANets). Particularly, the MCCANet model learns the stacked multiperspective filter banks by the MCCA method and builds a deep convolutional structure. In the output stage, the binarization and the blockwise histogram are employed as nonlinear processing and feature pooling, respectively. To access the effectiveness of the MCCANet, we conduct a host of experiments on the RSSCN7 RS database. Extensive experimental results demonstrate that the MCCANet outperforms the two-view CCANet.
Xinghao Yang, Weifeng Liu 0001, Dapeng Tao, Jun Cheng 0002
IEEE Geosci. Remote. Sens. Lett.2
2017 Cauchy Estimator Discriminant Learning for RGB-D Sensor-based Scene Classification
Dapeng Tao, Xipeng Yang, Weifeng Liu 0001, Shuifa Sun, Yanan Guo 0003, Jianxin Pang
Multim. Tools Appl.3
2017 LMAE: A large margin Auto-Encoders for classification
Weifeng Liu 0001, Tengzhou Ma, Qiangsheng Xie, Dapeng Tao, Jun Cheng 0002
Signal Process.1
2017 Multiview Cauchy Estimator Feature Embedding for Depth and Inertial Sensor-Based Human Action Recognition
abstract
The ever-growing popularity of Kinect and inertial sensors has prompted intensive research efforts on human action recognition. Since human actions were extracted from Kinect and inertial sensors, they can be characterized by multiple feature representations. By encoding the multiview features into a unified space, it could be optimal for human action recognition. In this paper, we propose a new unsupervised feature fusion method termed multiview Cauchy estimator feature embedding (MCEFE) for human action recognition. By minimizing empirical risk, MCEFE integrates the encoded complementary information in multiple views to find the unified data representation and the projection matrices. To enhance robustness to outliers, the Cauchy estimator is imposed on the reconstruction error. Furthermore, ensemble manifold regularization is enforced on the projection matrices to encode the correlations between different views and avoid overfitting. Experiments are conducted on the new Chinese Academy of Sciences—Yunnan University—multimodal human action database to demonstrate the effectiveness and robustness of MCEFE for human action recognition.
Yanan Guo 0003, Dapeng Tao, Weifeng Liu 0001, Jun Cheng 0002
IEEE Trans. Syst. Man Cybern. Syst.3
2016 Class specific dictionary learning based kernel collaborative representation for fine-grained image classification
abstract
Recently, dictionary learning based sparse representation algorithm has been widely adopted and achieved satisfying performance in image classification. However, sparse representation based classification (SRC) as well as collaborative representation based classification (CRC) always result in high residual error due to their basic assumption that considers training samples as dictionary directly for each category. And conventional class specific dictionary learning algorithm usually operates in the Euclidean space and fails to capture nonlinear information. To deal with these problems, we propose a classification algorithm which is called class specific dictionary learning based kernel collaborative representation (CSDL-KCRC) to enhance the classification accuracy. Extensive experimental results operated on three fine-grained image datasets, such as Oxford 102-Flowers dataset, Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset and Stanford Dogs dataset, demonstrate the effectiveness of CSDL-KCRC in image classification.
Xiaojie Feng, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001
SMC4
2016 Hessian regularization by patch alignment framework
Weifeng Liu 0001, Dapeng Tao
Neurocomputing1
2016 Manifold regularized kernel logistic regression for web image annotation
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001, Ke Lu 0002
Neurocomputing1
2016 HSAE: A Hessian regularized sparse auto-encoders
Weifeng Liu 0001, Tengzhou Ma, Dapeng Tao, Jane You
Neurocomputing1
2016 Large-scale paralleled sparse principal component analysis
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001, Ke Lu 0002
Multim. Tools Appl.1
2016 Big data meets multimedia analytics
Tat-Seng Chua, Xiangjian He, Weifeng Liu 0001, Massimo Piccardi, Yonggang Wen 0001, Dacheng Tao
Signal Process.3
2015 Hessian Regularized Sparse Coding for Human Action Recognition
Weifeng Liu 0001, Zhen Wang 0004, Dapeng Tao, Jun Yu 0002
MMM (2)1
2015 A general framework for co-training and its applications
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001
Neurocomputing1
2015 Multiview Hessian regularized logistic regression for action recognition
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001, Ke Lu 0002
Signal Process.1
2014 Class specific subspace learning for collaborative representation
abstract
Collaborative representation based classification (CRC) has been successfully used for visual recognition and showed impressive performance recently. However, it directly uses the training samples from each class as the subspaces to calculate the minimum residual error for a given testing sample. This leads to high residual error and instability, which is critical especially for a small number of training samples in each class. In this paper, we propose a class specific subspace learning algorithm for collaborative representation. By introducing the dual form of subspace learning, it presents an explicit relationship between the basis vectors and the original image features, and thus enhances the interpretability. Lagrange multipliers are then applied to optimize the corresponding objective function, i.e., learning the weights used in constructing the subspaces. Extensive experimental results demonstrate that the proposed algorithm has achieved superior performance in several visual recognition tasks.
Baodi Liu, Bin Shen 0002, Yu-Xiong Wang, Weifeng Liu 0001, Yanjiang Wang 0001
SMC4
2014 Multiview Hessian discriminative sparse coding for image annotation
Weifeng Liu 0001, Dacheng Tao, Jun Cheng 0002, Yuan Yan Tang
Comput. Vis. Image Underst.1
2013 Discriminant Multi-component Face Analysis
abstract
Sparse representation based classification (SRC) has attracted much attention in face analysis such as face recognition (FR) and face expression recognition (FER). Currently, most of SRC based methods treated face as a whole component which results in under-utilization of the complementary in different facial parts. In this paper, we present an approach which can effectively explore the complementary of different facial parts to boost the performance of face analysis. In particular, we employ multi-view sparse coding techniques to learn the factorized representation of different facial components. Furthermore, we incorporate label information into the objective function to enforce the discriminability. To evaluate the performance, we conduct face analysis experiments including FR and FER on JAFFE database. Experimental results demonstrate that the proposed method can significantly boost the performance of face analysis.
Weifeng Liu 0001, Liping Dong, Yanjiang Wang 0001
SMC2
2013 Self-Explanatory Convex Sparse Representation for Image Classification
abstract
Sparse representation technique has been widely used in various areas of computer vision over the last decades. Unfortunately, in the current formulations, there are no explicit relationship between the learned dictionary and the original data. By tracing back and connecting sparse representation with the K-means algorithm, a novel variation scheme termed as self-explanatory convex sparse representation (SCSR) has been proposed in this paper. To be specific, the basis vectors of the dictionary are refined as convex combination of the data points. The atoms now would capture a notion of centroids similar to K-means, leading to enhanced interpretability. Sparse representation and K-means are thus unified under the same framework in this sense. Besides, an appealing property also emerges that the weight and code matrices both tend to be naturally sparse without additional constraints. Compared with the standard formulations, SCSR is easier to be extended into the kernel space. To solve the corresponding sparse coding sub problem and dictionary learning sub problem, block-wise coordinate descent and Lagrange multipliers are proposed accordingly. To validate the proposed algorithm, it is implemented in image classification, a successful applications of sparse representation. Experimental results on several benchmark data sets, such as UIUC-Sports, Scene 15, and Caltech-256 demonstrate the effectiveness of our proposed algorithm.
Baodi Liu, Yu-Xiong Wang, Bin Shen 0002, Yu-Jin Zhang, Yanjiang Wang 0001, Weifeng Liu 0001
SMC6
2013 Multiview Hessian Regularization for Image Annotation
abstract
The rapid development of computer hardware and Internet technology makes large scale data dependent models computationally tractable, and opens a bright avenue for annotating images through innovative machine learning algorithms. Semisupervised learning (SSL) therefore received intensive attention in recent years and was successfully deployed in image annotation. One representative work in SSL is Laplacian regularization (LR), which smoothes the conditional distribution for classification along the manifold encoded in the graph Laplacian, however, it is observed that LR biases the classification function toward a constant function that possibly results in poor generalization. In addition, LR is developed to handle uniformly distributed data (or single-view data), although instances or objects, such as images and videos, are usually represented by multiview features, such as color, shape, and texture. In this paper, we present multiview Hessian regularization (mHR) to address the above two problems in LR-based image annotation. In particular, mHR optimally combines multiple HR, each of which is obtained from a particular view of instances, and steers the classification function that varies linearly along the data manifold. We apply mHR to kernel least squares and support vector machines as two examples for image annotation. Extensive experiments on the PASCAL VOC'07 dataset validate the effectiveness of mHR by comparing it with baseline algorithms, including LR and HR.
Weifeng Liu 0001, Dacheng Tao
IEEE Trans. Image Process.1
2013 Hessian Regularized Support Vector Machines for Mobile Image Annotation on the Cloud
abstract
With the rapid development of the cloud computing and mobile service, users expect a better experience through multimedia computing, such as automatic or semi-automatic personal image and video organization and intelligent user interface. These functions heavily depend on the success of image understanding, and thus large-scale image annotation has received intensive attention in recent years. The collaboration between mobile and cloud opens a new avenue for image annotation, because the heavy computation can be transferred to the cloud for immediately responding user actions. In this paper, we present a scheme for image annotation on the cloud, which transmits mobile images compressed by Hamming compressed sensing to the cloud and conducts semantic annotation through a novel Hessian regularized support vector machine on the cloud. We carefully explained the rationality of Hessian regularization for encoding the local geometry of the compact support of the marginal distribution and proved that Hessian regularized support vector machine in the reproducing kernel Hilbert space is equivalent to conduct Hessian regularized support vector machine in the space spanned by the principal components of the kernel principal component analysis. We conducted experiments on the PASCAL VOC'07 dataset and demonstrated the effectiveness of Hessian regularized support vector machine for large-scale image annotation.
Dapeng Tao, Weifeng Liu 0001, Xuelong Li 0001
IEEE Trans. Multim.3
2012 Facial expression recognition based on Gabor features and sparse representation
abstract
In this paper, we present a facial expression recognition method based on Gabor feature and sparse representation. Sparse Representation based Classification (SRC) has been widely used in computer vision and pattern recognition. And Gabor filter banks can be used to approximately model the signal processing in visual primary cortex. We believe that the nature of the attractive performance of SRC and Gabor feature lies in that they both followed the natures of signal perception of retina and information processing of cortex in human vision. Therefore, we combined the Gabor feature and SRC for facial expression recognition. The comparison experiments of proposed Gabor+SRC algorithm and straightforward SRC application are conducted on JAFFE database. And the experimental results showed the attractive performance of the proposed Gabor+SRC method.
Weifeng Liu 0001, Caifeng Song, Yanjiang Wang 0001
ICARCV1
2012 Subject-Independent Facial Expression Recognition with Biologically Inspired Features
abstract
Despite of much research for facial expression recognition, recognizing facial expressions across different persons is still a challenging computer vision task. However, facial expression analysis seems naturally for human visual system. Motivated by visual biology, this paper proposes an invariant feature extraction method for subject-independent facial expression recognition. In particular, we extract the biologically inspired facial features using extended visual cortex model-HMAX which consist of a template matching and a maximum pooling operation. We carefully organized the facial features and achieve subject-independent facial expression recognition using a sparse representation based classifier. The experiments on Yale database and JAFFE database demonstrate the significance of our proposed method for subject-independent facial expression recognition.
Weifeng Liu 0001, Caifeng Song, Yanjiang Wang 0001
ICMLA (1)1
2012 Cellular Differentiation Algorithm for High Dimensional Numerical Function Optimization
abstract
Inspired by the cellular differentiation mechanism of organisms, combined with the theory of artificial life and swarm intelligence, a new biomimetic optimization algorithm, cellular differentiation optimization algorithm (CDOA), is proposed in this paper. A certain number of cells are randomly distributed in the search space to find the optimal solution by activating their differential behaviors such as division, growth, migration, adhesion and apoptosis. Experimental results on several benchmark complex functions with high dimensions show that the proposed cellular differentiation optimization algorithm can rapidly converge at high quality solutions and outperform some of the state-of-art in high-dimension numerical function optimization.
Yanjiang Wang 0001, Chengna Yuan, Weifeng Liu 0001
ICMLA (1)3
2012 Facial expression recognition based on discriminative dictionary learning
Weifeng Liu 0001, Caifeng Song, Yanjiang Wang 0001
ICPR1