Fei Wu 0004

dblp:84/3254-4 · DBLP profile ↗
← Back
104ranked-venue papers
26as first author
58since 2021 · last 2026
0000-0001-5498-4947ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 16 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 8 first-author · 21 since 2021Software engineering, systems software and programming languages · 10 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 6 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Irrelevant feature filtering module for deep multi-view generative clustering
Xiaoyuan Jing, Yong-Fang Yao, Wei Liu 0200, Fei Wu 0004, Changhui Hu 0001, Ziyun Cai
Inf. Sci.5
2026 Cross-modal hashing under the federated learning framework
Qinghua Huang, Jiahuan Lu, Fei Wu 0004, Guangchuan Peng, Guangwei Gao, Xiaoyuan Jing
J. Vis. Commun. Image Represent.3
2026 Compressive self-attention transformer for low-light enhancement and zero-element pixels restoration
Changhui Hu 0001, Donghang Jing, Kerui Hu, Tiesheng Chen, Ziyun Cai, Fei Wu 0004, Xiaoyuan Jing
Pattern Recognit.7
2026 Extremely degraded face image super-resolution based on high frequency attention and noisy facial priors
Xiaoke Zhu, Jihui Hu, Xiaopan Chen, Fan Zhang 0028, Fei Wu 0004, Xiaoyuan Jing
Signal Process. Image Commun.5
2026 End-to-End Open-Set Semi-Supervised Learning for Fine-Grained Encrypted Traffic Classification
abstract
Encrypted traffic classification is crucial for enhancing network management, service quality, and security. However, real-world network environments are inherently open-world scenarios in which traffic not only consists of known classes but also includes the continuous emergence of unknown classes. Existing deep learning methods typically rely on the closed-world assumption, which significantly limits their classification performance when dealing with unknown traffic types. This limitation makes it challenging to accurately classify known traffic classes and effectively identify unknown ones. Although few studies have focused on open-world scenarios, these methods often use staged strategies and struggle to reliably detect unknown traffic or to estimate novel classes. To address these challenges, we propose an end-to-end Fine-grained Encrypted traffic Classification method based on Open-set Semi-supervised Learning, called FEC-OSL. This method comprises three mutually reinforcing core components. First, we design a dual-branch flow feature extraction module to capture detailed and discriminative flow features. Second, we introduce a novel energy-based perspective that leverages energy-boundary learning to distinguish known traffic from unknown traffic, enabling precise detection of known classes. Finally, an adaptive deep clustering approach integrates feature learning with clustering to achieve fine-grained classification of unknown flows. We conduct extensive experiments on three real-world datasets, and the results validate that our proposed method exhibits outstanding performance in handling both known and unknown encrypted traffic in open-world scenarios.
Hongyu Du, Fei Wu 0004, Shangdong Liu, Yimu Ji 0001, Kui Ren 0001
IEEE Trans. Inf. Forensics Secur.6
2025 Efficient Cross-modal Prompt Learning with Semantic Enhancement for Domain-robust Fake News Detection
abstract
With the development of multimedia technology, online social media has become a major medium for people to access news, but meanwhile, it has also exacerbated the dissemination of multi-modal fake news. An automatic and efficient multi-modal fake news detection (MFND) method is urgently needed. Existing MFND methods usually conduct cross-modal information interaction at later stage, resulting in insufficient exploration of complementary information between modalities. Another challenge lies in the differences among news data from different domains, leading to the weak generalization ability in detecting news from various domains. In this work, we propose an efficient Cross-modal Prompt Learning with Semantic enhancement method for Domain-robust fake news detection (CPLSD). Specifically, we design an efficient cross-modal prompt interaction module, which utilizes prompt as medium to realize lightweight cross-modal information interaction in the early stage of feature extraction, enabling to exploit rich modality complementary information. We design a domain-general prompt generation module that can adaptively blend domain-specific news features to generate domain-general prompts, for improving the domain generalization ability of the model. Furthermore, an image semantic enhancement module is designed to achieve image-to-text translation, fully exploring the semantic discriminative information of the image modality. Extensive experiments conducted on three MFND benchmarks demonstrate the superiority of our proposed approach over existing state-of-the-art MFND methods.
Fei Wu 0004, Changhui Hu 0001, Yimu Ji 0001, Xiaoyuan Jing, Guoping Jiang
COLING1
2025 LightGR-Transformer: Light Grouped Residual Transformer for Multispectral Object Detection
Fei Wu 0004, Yinjie Wang
CVM (3)2
2025 Complementary Graph Learning and Prompt-based Cross-modal Generation for Missing-modality Fake News Detection
abstract
Multi-modal fake news detection (MFND) has attracted increasing attention. However, due to information loading failure or access restriction, incomplete modality makes joint multi-modal information extraction be challenging. Existing MFND methods with missing-modality focus on specific missing-modality cases, and how to flexibly deal with various missing-modality cases of news in the real world has not been well studied. In this paper, we propose a novel fake news detection approach named Complementary Graph learning and Prompt-based cross-modal Generation network (CG-PG), which contains two main modules: a complementary graph learning module and a prompt-based cross-modal generation module. The complementary graph learning module explores structural complementary information in image and text graphs to implement cross-modal information propagation. To further recover the information loss caused by missing modalities, the prompt-based cross-modal generation module generates representations of the missing modality from available modalities and imposes task-related constraints on the representations with the missing-aware prompts. Experimental results on the public Weibo and Fakeddit datasets under various missing-modality cases show that CG-PG outperforms state-of-the-art related works.
Fei Wu 0004, Ruixuan Zhou, Changhui Hu 0001, Qinghua Huang, Xiaoyuan Jing
ICASSP1
2025 Joint Decision Network with Modality-Specific and Dual Interactive Features for Fake News Detection
Fei Wu 0004, Ruixuan Zhou, Yimu Ji 0001, Xiaoyuan Jing
MMM (2)1
2025 Multi-view learning based on product and process metrics for software defect prediction
Ying Sun 0023, Fei Wu 0004, Di Wu 0014, Xiaoyuan Jing, Yanfei Sun
Appl. Intell.2
2025 AF-MCDC: Active Feedback-Based Malicious Client Dynamic Detection
Hongyu Du, Shouhui Zhang, Xi Xv, Yimu Ji 0001, Fei Wu 0004, Shangdong Liu
Comput. Networks6
2025 Sample-pair learning network for extremely imbalanced classification
abstract
In data classification, class-balanced data is ideal, but real datasets are often imbalanced, necessitating rebalancing through methods like resampling. In recent years, some new generative model-based resampling methods have been proposed. However, when facing extreme class imbalance, where the minority class is strongly underrepresented and on its own does not contain enough information to conduct the generative process. Some deep learning methods have been proposed to solve extremely imbalanced classification problems, but some of them are only used for specific datasets. Therefore, we proposed a novel deep learning method that combines a generative strategy with multi-task joint learning, termed sample-pair learning network (SPLN), for extremely imbalanced classification. The network consists of data preprocessing and multi-task joint learning modules. During data preprocessing, the training set is expanded by constructing positive and negative sample-pairs, then rebalanced using a strategy combining attention and resampling, termed undersampling based on attention power values (APVUS). The multi-task joint learning module employs a Siamese convolutional subnetwork to measure the similarity between sample-pairs and a multi-layer perceptron to recognize the category of single samples. The module can reduce the risk of overfitting caused by excessive noise in the training set. Finally, we designed a voting model based on the Siamese convolutional subnetwork to infer the categories of test samples. Experimental results demonstrate that our approach outperforms state-of-the-art generative model-based methods and is effective and general for extremely imbalanced classification.
Linjun Chen, Xiaoyuan Jing, Runhang Chen, Fei Wu 0004, Yongchang Ding, Changhui Hu 0001, Ziyun Cai
Neurocomputing4
2025 A Mimic Honeypot Construction Method Based on Incomplete Information Zero-Sum Stochastic Games and Q-Learning
abstract
Honeypots based on deception technology offer a promising solution to address the asymmetry of attack and defense in the Internet of Things (IoTs). However, as the IoT security situation continues to evolve, attackers can identify honeypots by analyzing system characteristics and network behaviors, launching targeted virtual escape attacks that may exploit the honeypot as a stepping stone to compromise other systems. Once the IoT honeypot itself is successfully identified and attacked, the current defense measures typically rely on post-attack remediation. To address this challenge, we propose a mimic honeypot construction method based on incomplete information zero-sum stochastic game and Q-Learning. This method enhances the IoT honeypot’s deceptive capabilities while ensuring the security of the honeypot itself. Firstly, inspired by the concept of mimic defense, we design a dynamic heterogeneous redundancy honeypot (mimic honeypot), which contains multiple business executors composed of both business and virtualization layers. Secondly, we establish an incomplete information zero-sum stochastic game model to represent the honeypot attack-defense scenario. The Q-Learning algorithm is employed to solve for the Bayesian Nash equilibrium, enabling the mimic honeypot to adaptively adjust its deployment strategies based on the attacker’s observed actions. Finally, the experimental results demonstrate that the proposed mimic honeypot outperforms existing methods in terms of deceptive effectiveness and honeypot self-protection capabilities, significantly reducing the likelihood of honeypot compromise and ensuring robust network defense.
Zongkai Ji, Xukun Qian, Fei Wu 0004, Shangdong Liu, Yimu Ji 0001
IEEE Internet Things J.4
2025 Saliency Feature Learning for Multimodality Person Retrieval in Visual Internet of Things
abstract
In the visual Internet of Things (VIoT), smart surveillance is an important component of the multimodality person retrieval (MPR) task. Capturing discriminative pedestrian information from images aids in identifying target individuals through sketch or text descriptions, improving the reliability of VIoT systems. Current granularity-level matching methods face challenges with information redundancy and modality differences in MPR tasks. In this article, we propose a novel approach for intra and intermodality saliency feature learning by multigranularity feature selection and relation (MFSR). First, to reduce redundant information within each modality, the intramodality feature selection module (IFSM) employs the adaptive weighting mechanism to enhance salient pedestrian features while suppressing irrelevant features. Second, the multigranularity feature relation module (MFRM) aligns cross-modality salient person features by increasing the similarity scores of local and global features across modalities, to reduce differences between multimodality visual-text pairs. Finally, the cross-modality similarity matching (CSM) loss is designed to enhance consistency in visual-text pairs of the same identity, ensuring compact intraclass features by minimizing discrepancies between cross-modality similarity distributions and identity-matching distributions. Experimental results show that our approach achieves state-of-the-art performance on benchmark datasets.
Shuai You, Yujian Feng, Shitao Wang, Fei Wu 0004, Yuchen Sha, Yimu Ji 0001, Xiaoyuan Jing
IEEE Internet Things J.4
2025 Semi-supervised cross-modality person re-identification based on pseudo label learning
Fei Wu 0004, Ruixuan Zhou, Yang Gao 0001, Yujian Feng, Qinghua Huang, Xiaoyuan Jing
Image Vis. Comput.1
2025 Output difference feedback and system benefit control based dynamic heterogeneous redundancy architecture
abstract
Mimic active defense technology effectively disrupts attack routes and reduces the probability of successful attacks by using a dynamic heterogeneous redundancy (DHR) architecture. However, current approaches often overlook the adaptability of the adjudication mechanism in complex and variable network environments, focusing primarily on system security while neglecting performance considerations. To address these limitations, we propose an output difference feedback and system benefit control based DHR architecture. This architecture introduces an adjudication mechanism based on output difference feedback, which enhances adaptability by considering the impact of each executor’s output deviation on the global decision. Additionally, the architecture incorporates a scheduling strategy based on system benefit, which models the quality of service and switching overhead as a bi-objective optimization problem, balancing security with reduced computational costs and system overhead. Simulation results demonstrate that our architecture improves adaptability towards different network environments and effectively reduces both the attack success rate and average failure rate.
Zhibo He, Shangdong Liu, Weili Zhang, Fei Wu 0004, Fukang Zeng, Jun Zuo, Longfei Zhou, Yukun Niu, Yimu Ji 0001
Frontiers Inf. Technol. Electron. Eng.5
2025 Cluster-graph convolution networks for robust multi-view clustering
Xiaoyuan Jing, Wei Liu 0200, Fei Wu 0004, Changhui Hu 0001, Bo Du 0001
Knowl. Based Syst.4
2025 Learning multi-granularity representation with transformer for visible-infrared person re-identification
Yujian Feng, Feng Chen 0047, Guozi Sun, Fei Wu 0004, Yimu Ji 0001, Tianliang Liu, Shangdong Liu, Xiaoyuan Jing, Jiebo Luo 0001
Pattern Recognit.4
2025 Homogeneous and heterogeneous relational graph for visible-infrared person re-identification
Yujian Feng, Feng Chen 0047, Jian Yu 0007, Yimu Ji 0001, Fei Wu 0004, Shangdong Liu, Xiaoyuan Jing
Pattern Recognit.5
2025 Balanced Multi-modal Learning with Hierarchical Fusion for Fake News Detection
Fei Wu 0004, Guangwei Gao, Yimu Ji 0001, Xiaoyuan Jing
Pattern Recognit.1
2025 PFBL: Prototype-Based Fully Balanced Learning for Multimodal Fake News Detection
abstract
Multimodal fake news detection (MFND) has received widespread attention. As a typical multimodal task, MFND is troubled by modality imbalance, where the dominant modality suppresses other modalities during the optimization process. Meanwhile, the social credibility maintains the amount of fake news less than that of real news, which causes class imbalance. We call MFND with intertwined influence of modality imbalance and class imbalance as dual imbalanced MFND, i.e., DI-MFND, which has not been well studied. In this article, we propose an approach called prototype-based fully balanced learning (PFBL) for DI-MFND. Specifically, our model contains three main parts. 1) The prototype-based modality-balanced learning (PMB) part, which constructs modality prototypes for each modality. It designs the prototype-based unimodal loss to enhance the intraclass compactness of the poorly performing modality and a constrained gradient optimization strategy to suppress the dominant modality for optimization. 2) The prototype-based class-balanced learning (PCB) part constructs class prototypes. A prototype-oriented discriminative enhancement loss is designed to effectively align samples with corresponding prototypes and increase the inter-class distance, thus enhancing the separability of different categories. 3) The multimodal fusion and classification (MFC) part employs a cross-transformer block to perform interaction between modalities and further integrates features of different modalities by attention mechanism. We propose a joint updating strategy for modality prototypes and class prototypes. Extensive experiments on three widely-used news datasets demonstrate that our approach outperforms state-of-the-art approaches.
Fei Wu 0004, Zhe-Ying Deng, Di Wu 0014, Chao Lan, Yimu Ji 0001, Xiaoyuan Jing
IEEE Trans. Comput. Soc. Syst.1
2025 An Active Defense Adjudication Method Based on Adaptive Anomaly Sensing for Mimic IoT
abstract
Security issues in the Internet of Things (IoT) are inevitable. Uncertain threats, such as known vulnerabilities and backdoors exist within IoT, and traditional passive network security technologies are ineffective against uncertain threats. To address the above issues, we propose an active defense adjudication method based on adaptive anomaly sensing for mimic IoT. The method constructs a mimic IoT active defense architecture, improving system security and reliability despite prevailing security threats. In addition, an intelligent anomaly sensing algorithm is integrated into the adjudication module of the mimic IoT active defense architecture to support arbitration. An adaptive anomaly sensing model based on multi-feature selection is used to determine the anomaly score of the IoT device outputs, and this model fully considers the reliability of the adjudication data and improves the accuracy of the adjudication. Finally, we conduct a comparative analysis of the proposed adjudication algorithm against three others via a mimic power communication IoT system as an application scenario. The experimental results show that our algorithm can improve security and reduce the failure rate of the mimic IoT system.
Tiansheng Gu, Yijun Nie, Zongkai Ji, Fei Wu 0004, Zhongjie Ba, Yimu Ji 0001, Kui Ren 0001, Guozi Sun
IEEE Trans. Serv. Comput.5
2024 Joint Multi-modal Graph Structure and Representation Learning for Fake News Detection
abstract
With the development of the Internet and multi-media technologies, fake news spread on the Internet has become a problem that cannot be ignored by the public and the government. Multi-modal fake news detection methods have tried to make use of the powerful representation ability of graph convolutional network in the fake news detection task. However, how to learn optimal graph structure for subsequent graph learning and effectively bridge the inter-modality gap is still a challenge. In this paper, we propose a novel approach called Joint Multi-modal Graph Structure and Representation Learning (JMGSRL) for multi-modal fake news detection. It consists of two main modules, i.e., a multi-modal graph structure learning module and a modality discriminator module. The multi-modal graph structure learning module jointly optimizes the graph structures of multiple modalities based on graph convolutional network by adding the potential edges or reweighting the unreasonable edges. The optimized graph structures are further used for subsequent graph learning. To effectively deal with the modality difference issue, the modality discriminator is designed to learn modality-invariant features, and the network training is guided by the adversarial scheme. Experiments on two public real datasets demonstrate that JMGSRL can outperform the state-of-the-art related multi-modal fake news detection methods.
Jiahuan Lu, Fei Wu 0004, Yimu Ji 0001, Xiaoyuan Jing
HPCC3
2024 Semantic Distillation and Structural Alignment Network for Fake News Detection
abstract
In recent years, the rapid proliferation of multi-modal fake news has posed potential harm across various sectors of society, making the detection of multi-modal fake news crucial. Most existing methods can not effectively reduce the redundant information and preserve both semantic and structural information. To address these problems, this paper proposes a semantic distillation and structural alignment (SDSA) network. We design an semantic distillation module for modality-specific features to preserve task-relevant semantic information and eliminate redundant information. Then, we propose a triple similarity alignment module to preserve structural information. Specifically, intra-modal similarity alignment mines intra-modal consistency by preserving the neighborhood structure within each modality, inter-modal similarity alignment explores cross-modality consistency by bringing the cross-modality feature neighborhood structures, and joint similarity alignment aims to preserve the structural information of fused features. Experiments conducted on two widely used fake news datasets demonstrate that the SDSA method outperforms state-of-the-art approaches.
Shangdong Liu, Xiaofan Yue, Fei Wu 0004, Yujian Feng, Yimu Ji 0001
ICASSP3
2024 An empirical study of data sampling techniques for just-in-time software defect prediction
Zhiqiang Li 0003, Qiannan Du, Hongyu Zhang 0002, Xiaoyuan Jing, Fei Wu 0004
Autom. Softw. Eng.5
2024 Resan: A Residual Dual-Attention Network for Abnormal Cardiac Activity Detection
abstract
ABSTRACT Cardiovascular disease is one of the leading causes of death worldwide. Early and accurate detection of abnormal cardiac activity can be an effective way to prevent serious cardiovascular events. Electrocardiogram (ECG) and phonocardiogram (PCG) signals provide an objective evaluation of the heart's electrical and acoustic functions, enabling medical professionals to make an accurate diagnosis. Therefore, the cardiologists often use them to make a preliminary diagnosis of abnormal cardiac activity in clinical practice. For this reason, many diagnostic models have been proposed. However, these models fail to utilize the interaction information within and between the signals to aid the diagnosis of disease. To address this issue, we designed a residual dual‐attention network (ResAN) for the detection of abnormal cardiac activity using synchronized ECG and PCG signals. First, ResAN uses a feature learning module with two parallel residual networks, for example, ECG‐ResNet and PCG‐ResNet to automatically learn the deep modal‐specific features from the ECG and PCG sequences, respectively. Second, to fully utilize the available information of different modal signals, ResAN uses a dual‐attention fusion module to capture the salient features of the integrated ECG and PCG features learned by the feature learning module, as well as the alternating features between them based on the attention mechanisms. Finally, these fused features are merged and fed to the classification module to detect abnormal cardiac activity. Our model achieves an accuracy of 96.1%, surpassing the performances of comparison models by 1.0% to 9.9% when using synchronized ECG and PCG signals. Furthermore, the ablation study confirmed the efficacy of the components in ResAN and also showed that ResAN performs better with synchronized ECG and PCG signals compared to using single‐modal signals. Overall, ResAN provides a valid solution for the early detection of abnormal cardiac activity using ECG and PCG signals.
Fei Wu 0004, Datun Qi, Xiaoyuan Jing
Comput. Intell.3
2024 Mimic turbo compiled code structure for wireless communication systems
abstract
Abstract Turbo codes play a crucial role in wireless communication systems, and their compiled code structures are key factors affecting the performance of the entire communication system. As a result, the study of turbo compiled code structures has been a focal point for researchers. The iterative decoding of turbo code structures has multiple limitations and large storage resource consumption, leading to poor system anti‐interference ability and a rapid increase in BER. To address these issues, this paper proposes the mimic turbo compiled code structure (MTCCS) for wireless communication systems. MTCCS is based on the DHR idea, incorporating dynamic, heterogeneous, and redundancy characteristics. Dynamicity is achieved through a dynamic scheduling algorithm based on abnormal feedback information. Heterogeneity is achieved through a codec component collection design method based on intrinsic and extrinsic heterogeneity. Redundancy is achieved through a majority voting algorithm. At the beginning of information transmission, MTCCS randomly selects heterogeneous codecs from the heterogeneous codec collection to enter the runtime pool. After the information transmission is complete, the majority voting algorithm is used to adjudicate the multi‐mode output of the codecs, resulting in a relatively accurate decoding outcome. Meanwhile, the dynamic scheduling module calculates the abnormal feedback information of each codec and accordingly dynamically schedules the mimic turbo codecs to replace the abnormal ones. Through the above process, MTCCS realizes the adaptive compilation code and improves the anti‐interference ability of turbo code. Simulation experiments are conducted on MTCCS in both non‐interference and interference scenarios. Simulation experiments show that MTCCS introducing the DHR idea achieves a balance between anti‐interference and decoding performance. It effectively addresses the issue of poor anti‐interference ability in turbo codes, and the decoding performance of MTCCS is superior to that of the previous single conventional turbo codes.
Shangdong Liu, Yimu Ji 0001, Fei Wu 0004, Tiansheng Gu, Yulu Zheng, Yijun Nie, Zongkai Ji, Cailing Sun, Zeng Chen, Yawei Sun
IET Commun.5
2024 Sharpness-aware gradient guidance for few-shot class-incremental learning
Runhang Chen, Xiaoyuan Jing, Fei Wu 0004
Knowl. Based Syst.3
2024 Cycle mapping with adversarial event classification network for fake news detection
Fei Wu 0004, Yujian Feng, Guangwei Gao, Yimu Ji 0001, Xiaoyuan Jing
Multim. Tools Appl.1
2024 Swin transformer and ResNet based deep networks for low-light image enhancement
Lintao Xu, Changhui Hu 0001, Fei Wu 0004, Ziyun Cai
Multim. Tools Appl.4
2024 MFECLIP: CLIP With Mapping-Fusion Embedding for Text-Guided Image Editing
abstract
Recently, generative adversarial networks (GAN) have made remarkable progress, particularly with the advent of Contrastive Language-Image Pretraining (CLIP), which take image and text into a joint latent space, bridging the gap between these two modalities. Several impressive text-guided image editing methods based on GANs and CLIP have emerged. However, in these studies, most of them simply minimize the distance between the target image embedding and text embedding in the CLIP space, and take this objective as network's optimization goal, overlooking the real distance between them may be large. This may result in inability to accurately guide the editing process according to the text prompts and the changes in text-irrelevant attributes. To mitigate this issue, we propose a novel approach named CLIP with Mapping-Fusion Embedding (MFECLIP) for text-guided image editing, which comprises two components: the MFE Block and MFE Loss. Through the MFE Block, we obtain Mapping-Fusion Embedding (MFE), which can further eliminate the modality gap, and it can serve as a superior guide for editing process instead of the original text embedding. Based on contrastive learning, the MFE Loss is designed to achieve accurate alignment between the target image and text prompt. We have conducted extensive experiments on real datasets, CUB and Oxford, demonstrating the favorable performance of the proposed method.
Fei Wu 0004, Yongheng Ma, Xiaoyuan Jing, Guoping Jiang
IEEE Signal Process. Lett.1
2024 $\text{Offset}^{3}\text{Net}$: Simple Joint 3-D Detection and Tracking With Three-Step Offset Learning
abstract
Light-detection-and-ranging-based multiobject detection and tracking play fundamental roles in autonomous driving systems. Most existing detection and tracking methods inevitably require complex pairing permutations for object association across frames, making the framework slow. Moreover, the occlusion and viewpoint changes lead to missed and false detection. To solve the abovementioned issues, this article proposes a simple joint 3-D detection and tracking approach with three-step offset learning ($\text{Offset}^{3}\text{Net}$). Specifically,$\text{Offset}^{3}\text{Net}$incorporates three task-specific output subnetworks to learn three offsets: 1) center offset, 2) motion offset, and 3) association offset. The learning of abovementioned offsets eliminates the complex bipartite matching processing. Specifically, the center offset guides the model to generate precise detections, whereas the motion offset transforms the track from the previous frame to the current frame, and the association offset minimizes the distance between detection and motion-updated track of the same object. Then, a simple read-off operation is conducted for data association on a hybrid-time centerness map, which represents the detections and offset-updated tracks. In addition, we design a detection-feature-enhanced module that captures the temporal coherence of the object motion and appearance information, avoiding the missed and false detection. Experiments on nuScenes have demonstrated the effectiveness of our$\text{Offset}^{3}\text{Net}$in terms of accuracy and speed compared with most 3-D detection and tracking methods.
Yimu Ji 0001, Jing He 0004, Fei Wu 0004, Yanfei Sun
IEEE Trans. Ind. Informatics4
2024 Cross-Modality Spatial-Temporal Transformer for Video-Based Visible-Infrared Person Re-Identification
abstract
Video-based visible-infrared person re-identification (VVI-ReID) aims to match the identity of a person captured in video sequences from both visible and infrared cameras. The VVI-ReID task requires considering both the spatial relationship between body parts within each frame and the temporal change of appearance between successive frames. Existing VVI Re-ID methods employ Convolutional Neural Networks to extract local spatial features and Long Short-Term Memory to form temporal associations. However, these methods can not effectively capture the global spatial feature and the long-range temporal dependencies in ultra-long sequences. In this paper, we propose a Cross-modality Spatial-temporal Transformer (CST) including a Cross-frame Tube Transformer Module (CTTM) and a Multi-frame Transformer Fusion Module (MTFM) to address these challenges. Firstly, CTTM tokenizes a video clip into multiple 3D tubes, each encapsulating local spatial-temporal information of pedestrians, and then obtains global spatial-temporal representations by establishing the relationship between tubes. Secondly, we design MTFM to exchange information between multiple frames using message tokens, thus modeling the long-range temporal dependencies of features of pedestrians. In addition, to prevent the potential representation collapse caused by triplet-based loss functions, we propose a diversity-consistency (DC) loss function to preserve the diversity and consistency of cross-modality feature representations by imposing variance, invariance, and covariance constraints in feature representations. Extensive benchmark experiments demonstrate that our approach outperforms the state-of-the-art methods with large margins.
Yujian Feng, Feng Chen 0047, Jian Yu 0007, Yimu Ji 0001, Fei Wu 0004, Tianliang Liu, Shangdong Liu, Xiaoyuan Jing, Jiebo Luo 0001
IEEE Trans. Multim.5
2023 DE-net: Dynamic Text-Guided Image Editing Adversarial Networks
abstract
Text-guided image editing models have shown remarkable results. However, there remain two problems. First, they employ fixed manipulation modules for various editing requirements (e.g., color changing, texture changing, content adding and removing), which results in over-editing or insufficient editing. Second, they do not clearly distinguish between text-required and text-irrelevant parts, which leads to inaccurate editing. To solve these limitations, we propose: (i) a Dynamic Editing Block (DEBlock) that composes different editing modules dynamically for various editing requirements. (ii) a Composition Predictor (Comp-Pred), which predicts the composition weights for DEBlock according to the inference on target texts and source images. (iii) a Dynamic text-adaptive Convolution Block (DCBlock) that queries source image features to distinguish text-required parts and text-irrelevant parts. Extensive experiments demonstrate that our DE-Net achieves excellent performance and manipulates source images more correctly and accurately.
Ming Tao 0002, Bing-Kun Bao, Hao Tang 0005, Fei Wu 0004, Longhui Wei, Qi Tian 0001
AAAI4
2023 Inter-modal Fusion Network with Graph Structure Preserving for Fake News Detection
Fei Wu 0004, Xiaoke Zhu, Xiaoyuan Jing
ICONIP (6)2
2023 A DHR executor selection algorithm based on historical credibility and dissimilarity clustering
Yimu Ji 0001, Weili Zhang, Shangdong Liu, Fei Wu 0004, Fukang Zeng, Jun Zuo, Longfei Zhou
Sci. China Inf. Sci.7
2023 Task-specific parameter decoupling for class incremental learning
Runhang Chen, Xiaoyuan Jing, Fei Wu 0004, Yaru Hao
Inf. Sci.3
2023 Semi-supervised cross-modal hashing via modality-specific and cross-modal graph convolutional networks
Fei Wu 0004, Guangwei Gao, Yimu Ji 0001, Xiaoyuan Jing, Zhiguo Wan
Pattern Recognit.1
2023 Occluded Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to match person images between the visible and near-infrared modalities. Previous VI-ReID methods are based on holistic pedestrian images and achieve excellent performance. However, in real-world scenarios, images captured by visible and near-infrared cameras usually contain occlusions. The performance of these methods degrades significantly due to the loss of information of discriminative features from the occlusion of the images. We define visible-infrared person re-identification in this occlusion scene as Occluded VI-ReID, where only partial content information of pedestrian images can be used to match images of different modalities from different cameras. In this paper, we propose a matching framework for occlusion scenes, which contains a local feature enhance module (LFEM) and a modality information fusion module (MIFM). LFEM adopts Transformer to learn features of each modality, and adjusts the importance of patches to enhance the representation ability of local features of the non-occluded areas. MIFM utilizes a co-attention mechanism to infer the correlation between each image for reducing the difference between modalities. We construct two occluded VI-ReID datasets, namely Occluded-SYSU-MM01 and Occluded-RegDB datasets. Our approach outperforms existing state-of-the-art methods on two occlusion datasets, while remains top performance on two holistic datasets.
Yujian Feng, Yimu Ji 0001, Fei Wu 0004, Guangwei Gao, Yang Gao 0001, Tianliang Liu, Shangdong Liu, Xiaoyuan Jing, Jiebo Luo 0001
IEEE Trans. Multim.3
2023 Visible-Infrared Person Re-Identification via Cross-Modality Interaction Transformer
abstract
Visible-infrared person re-identification (VI Re-ID) is designed to match person images of the same identity from visible and infrared cameras. Transformer structures have been successfully applied in the field of VI Re-ID. However, previous Transformer-based methods were mainly designed to capture global content information in a single modality, and could not simultaneously perceive semantic information between two modalities from a global perspective. To solve this problem, we propose a novel framework named the cross-modality interaction Transformer (CMIT). It has strong abilities in modeling spatial and sequential features that can capture dependencies between long-range features, and explicitly improves the discriminativeness of features by exchanging information across modalities, thus contributing to obtaining modality-invariant representations. Specifically, CMIT utilizes a cross-modality attention mechanism to enrich the feature representations of each patch token by interacting with the patch tokens of the other modality, and aggregates local features of the CNN structure and global information of the Transformer structure to mine feature saliency representation. Furthermore, the modality-discriminative (MD) loss function is proposed to learn potential consistency between modalities to encourage intra-modality compactness within class and inter-modality separation between classes. Extensive experiments on two benchmarks demonstrate that our approach outperforms state-of-the-art methods.
Yujian Feng, Jian Yu 0007, Feng Chen 0047, Yimu Ji 0001, Fei Wu 0004, Shangdong Liu, Xiaoyuan Jing
IEEE Trans. Multim.5
2023 JDSR-GAN: Constructing an Efficient Joint Learning Network for Masked Face Super-Resolution
abstract
With the growing importance of preventing the COVID-19 virus in cyber-manufacturing security, face images obtained in most video surveillance scenarios are usually low resolution together with mask occlusion. However, most of the previous face super-resolution solutions can not efficiently handle both tasks in one model. In this work, we consider both tasks simultaneously and construct an efficient joint learning network, called JDSR-GAN, for masked face super-resolution tasks. Given a low-quality face image with mask as input, the role of the generator composed of a denoising module and super-resolution module is to acquire a high-quality high-resolution face image. The discriminator utilizes some carefully designed loss functions to ensure the quality of the recovered face images. Moreover, we incorporate the identity information and attention mechanism into our network for feasible correlated feature expression and informative feature learning. By jointly performing denoising and face super-resolution, the two tasks can complement each other and attain promising performance. Extensive qualitative and quantitative results show the superiority of our proposed JDSR-GAN over some competitive methods.
Guangwei Gao, Fei Wu 0004, Huimin Lu 0001, Jian Yang 0003
IEEE Trans. Multim.3
2022 Feature Distillation Interaction Weighting Network for Lightweight Image Super-resolution
abstract
Convolutional neural networks based single-image superresolution (SISR) has made great progress in recent years. However, it is difficult to apply these methods to real-world scenarios due to the computational and memory cost. Meanwhile, how to take full advantage of the intermediate features under the constraints of limited parameters and calculations is also a huge challenge. To alleviate these issues, we propose a lightweight yet efficient Feature Distillation Interaction Weighted Network (FDIWN). Specifically, FDIWN utilizes a series of specially designed Feature Shuffle Weighted Groups (FSWG) as the backbone, and several novel mutual Wide-residual Distillation Interaction Blocks (WDIB) form an FSWG. In addition, Wide Identical Residual Weighting (WIRW) units and Wide Convolutional Residual Weighting (WCRW) units are introduced into WDIB for better feature distillation. Moreover, a Wide-Residual Distillation Connection (WRDC) framework and a Self-Calibration Fusion (SCF) unit are proposed to interact features with different scales more flexibly and efficiently. Extensive experiments show that our FDIWN is superior to other models to strike a good balance between model performance and efficiency. The code is available at https://github.com/IVIPLab/FDIWN.
Guangwei Gao, Juncheng Li 0003, Fei Wu 0004, Huimin Lu 0001, Yi Yu 0001
AAAI4
2022 DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis
abstract
Synthesizing high-quality realistic images from text descriptions is a challenging task. Existing text-to-image Generative Adversarial Networks generally employ a stacked architecture as the backbone yet still remain three flaws. First, the stacked architecture introduces the entanglements between generators of different image scales. Second, existing studies prefer to apply and fix extra networks in adversarial learning for text-image semantic consistency, which limits the supervision capability of these networks. Third, the cross-modal attention-based text-image fusion that widely adopted by previous works is limited on several special image scales because of the computational cost. To these ends, we propose a simpler but more effective Deep Fusion Generative Adversarial Networks (DF-GAN). To be specific, we propose: (i) a novel one-stage text-to-image backbone that directly synthesizes high-resolution images without entanglements between different generators, (ii) a novel Target-Aware Discriminator composed of Matching-Aware Gradient Penalty and One-Way Output, which enhances the text-image semantic consistency without introducing extra networks, (iii) a novel deep text-image fusion block, which deepens the fusion process to make a full fusion between text and visual features. Compared with current state-of-the-art methods, our proposed DF-GAN is simpler but more efficient to synthesize realistic and text-matching images and achieves better performance on widely used datasets. Code is available at https://github.com/tobran/DF-GAN.
Ming Tao 0002, Hao Tang 0005, Fei Wu 0004, Xiaoyuan Jing, Bing-Kun Bao, Changsheng Xu
CVPR3
2022 Simulator Attack+ for Black-Box Adversarial Attack
abstract
Numerous researches on adversarial black-box attacks have proved that deep neural networks have certain insecurity. However, the current black-box attack methods still have shortages in incomplete utilization of query information. The newly proposed Simulator Attack based on meta-learning shows good performance in query-efficiency but still misses some hidden information. For this disadvantage, our research finds the usability of the feature layer output information in a simulator model for the first time. Then we propose an optimized Simulator Attack+ framework based on this discovery. By conducting experiments on the CIFAR-10 and CIFAR-100 datasets, results legibly show that Simulator Attack+ can further reduce the number of consuming queries to improve query-efficiency meanwhile maintaining attack effect. Our code is available at https://github.com/Rain117E/SimulatorAttackplus.
Yimu Ji 0001, Jianyu Ding, Zhiyu Chen 0004, Fei Wu 0004, Shangdong Liu
ICIP4
2022 DVO + LCLMF: A web service recommendation mechanism with QoS privacy preservation
abstract
Abstract QoS‐aware based web service recommendation is one of the crucial solutions to help users find high‐quality web services. To accurately predict the QoS values of candidate services, it is usually required to collect historical QoS data of users (QoS data for short). If these collected QoS data are improperly processed, QoS data privacy may be threatened. However, how to accurately predict the QoS values of candidate services while protecting QoS data privacy has not been well studied. In response to the situation, we propose a hybrid web service recommendation mechanism, which is divided into three parts. In the first part, the QoS data privacy preservation algorithm, which called DVO, is proposed based on keeping the cosine similarity of QoS data unchanged, that is, to realize the confusion of QoS data while ensuring the availability of QoS data remains unchanged. In the second part, a hybrid matrix factorization model based on location information and service features, which called LCLMF, is proposed to improve the accuracy of QoS values prediction. According to DVO and LCLMF, the DVO + LCLMF is designed in the third part, which can accurately predict QoS values while protecting QoS data privacy. The experimental results show that DVO + LCLMF can accurately predict the QoS values of candidate services on the basis of attaining QoS data privacy protection.
Yimu Ji 0001, Shangdong Liu, Fei Wu 0004, Haichang Yao, Jing He 0004, Yanlan Liu, Shuai You
Concurr. Comput. Pract. Exp.4
2022 Semi-supervised multi-view graph convolutional networks with application to webpage classification
Fei Wu 0004, Xiaoyuan Jing, Pengfei Wei 0001, Chao Lan, Yimu Ji 0001, Guoping Jiang, Qinghua Huang
Inf. Sci.1
2022 Dual-aligned unsupervised domain adaptation with graph convolutional networks
Fei Wu 0004, Pengfei Wei 0001, Guangwei Gao, Changhui Hu 0001, Qi Ge, Xiaoyuan Jing
Multim. Tools Appl.1
2022 JSPNet: Learning joint semantic & instance segmentation of point clouds via feature self-similarity and cross-task probability
Feng Chen 0047, Fei Wu 0004, Guangwei Gao, Yimu Ji 0001, Guoping Jiang, Xiaoyuan Jing
Pattern Recognit.2
2022 Modality and Event Adversarial Networks for Multi-Modal Fake News Detection
abstract
With the popularity of news on social media, fake news has become an important issue for the public and government. There exist some fake news detection methods that focus on information exploration and utilization from multiple modalities, e.g., text and image. However, how to effectively learn both modality-invariant and event-invariant discriminant features is still a challenge. In this paper, we propose a novel approach named Modality and Event Adversarial Networks (MEAN) for fake news detection. It contains two parts: a multi-modal generator and a dual discriminator. The multi-modal generator extracts latent discriminant feature representations of text and image modalities. A decoder is adopted to reduce information loss in the generation process for each modality. The dual discriminator includes a modality discriminator and an event discriminator. The discriminator learns to classify the event or the modality of features, and network training is guided by the adversarial scheme. Experiments on two widely used datasets show that MEAN can perform better than state-of-the-art related multi-modal fake news detection methods.
Pengfei Wei 0001, Fei Wu 0004, Ying Sun 0023, Xiaoyuan Jing
IEEE Signal Process. Lett.2
2022 Leaning compact and representative features for cross-modality person re-identification
Guangwei Gao, Hao Shao, Fei Wu 0004, Meng Yang 0001, Yi Yu 0001
World Wide Web3
2021 Semantic Preserving Generative Adversarial Network For Cross-Modal Hashing
abstract
Cross-modal hashing has achieved significant progress in recent years. However, how to effectively learn more discriminative hash codes of each modality and simultaneous alleviate the loss of modality information is still a challenging problem. Focusing on this problem, in this paper, we propose a novel cross-modal hashing approach named Semantic Preserving Generative Adversarial Network (SPGAN). The overall network architecture consists of two sub-networks, i.e., a semantic preserving generative adversarial network module and a discriminative hashing module. The generator maps text features into the image feature space. And the discriminator judges whether the feature representations are real image features or generated image features. The adversarial learning process can effectively reduce modality difference and preserve information of the image modality as much as possible. The discriminative hashing module projects the real and generated image features into a Hamming space to obtain hash codes, and explores semantic similarities for enhancing the discriminant ability of hash codes. Experiments on two widely used datasets demonstrate that SPGAN can outperform state-of-the-art related works.
Fei Wu 0004, Xiaokai Luo, Qinghua Huang, Pengfei Wei 0001, Ying Sun 0023, Xiwei Dong, Zhiyong Wu 0006
ICIP1
2021 Adaptive deformable convolutional network
Feng Chen 0047, Fei Wu 0004, Guangwei Gao, Qi Ge, Xiaoyuan Jing
Neurocomputing2
2021 Semi-supervised Heterogeneous Defect Prediction with Open-source Projects on GitHub
abstract
The heterogeneous defect prediction (HDP) technique can predict defects in a target company using heterogeneous metric data from external company, which has received substantial research attention. However, existing HDP methods assume that source data is labeled but labeling data is expensive. Semi-supervised defect prediction technique can perform defect prediction with few labeled data. In this paper, we investigate a new problem — semi-supervised HDP (SHDP). To solve this problem, we propose a new approach named cost-sensitive kernel semi-supervised correlation analysis (CKSCA) as a solution of SHDP problem. It introduces unified metric representation and canonical correlation analysis to make the data distributions of different company projects more similar. CKSCA also designs a cost-sensitive kernel semi-supervised discriminant analysis mechanism to utilize the limited labeled data and sufficient real-life unlabeled data from different companies. Besides we collect lots of open-source projects from GitHub website to construct a new large-scale unlabeled dataset called GITHUB dataset. It contains 26,407 modules and is greater than each public project dataset. It has been public online and can be extended continuously. Experiments on the GITHUB dataset and other public datasets indicate that unlabeled GITHUB data can help prediction model improve prediction performance, and CKSCA is effective and efficient for solving SHDP problem.
Ying Sun 0023, Xiaoyuan Jing, Fei Wu 0004, Xiwei Dong, Yanfei Sun, Ruchuan Wang 0001
Int. J. Softw. Eng. Knowl. Eng.3
2021 Multiset Feature Learning for Highly Imbalanced Data Classification
abstract
With the expansion of data, increasing imbalanced data has emerged. When the imbalance ratio (IR) of data is high, most existing imbalanced learning methods decline seriously in classification performance. In this paper, we systematically investigate the highly imbalanced data classification problem, and propose an uncorrelated cost-sensitive multiset learning (UCML) approach for it. Specifically, UCML first constructs multiple balanced subsets through random partition, and then employs the multiset feature learning (MFL) to learn discriminant features from the constructed multiset. To enhance the usability of each subset and deal with the non-linearity issue existed in each subset, we further propose a deep metric based UCML (DM-UCML) approach. DM-UCML introduces the generative adversarial network technique into the multiset constructing process, such that each subset can own similar distribution with the original dataset. To cope with the non-linearity issue, DM-UCML integrates deep metric learning with MFL, such that more favorable performance can be achieved. In addition, DM-UCML designs a new discriminant term to enhance the discriminability of learned metrics. Experiments on eight traditional highly class-imbalanced datasets and two large-scale datasets indicate that: the proposed approaches outperform state-of-the-art highly imbalanced learning methods and are more robust to high IR.
Xiaoyuan Jing, Xinyu Zhang 0012, Xiaoke Zhu, Fei Wu 0004, Xinge You, Yang Gao 0001, Shiguang Shan, Jing-Yu Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Face illumination recovery for the deep learning feature under severe illumination variations
Changhui Hu 0001, Jian Yu 0007, Fei Wu 0004, Yang Zhang 0067, Xiaoyuan Jing, Xiaobo Lu, Pan Liu 0013
Pattern Recognit.3
2021 Spectrum-aware discriminative deep feature learning for multi-spectral face recognition
Fei Wu 0004, Xiaoyuan Jing, Yujian Feng, Yimu Ji 0001, Ruchuan Wang 0001
Pattern Recognit.1
2021 Efficient Cross-Modality Graph Reasoning for RGB-Infrared Person Re-Identification
abstract
The modality and pose variance between RGB and infrared (IR) images are two key challenges for RGB-IR person re-identification. Existing methods mainly focus on leveraging pixel or feature alignment to handle the intra-class variations and cross-modality discrepancy. However, these methods are hard to keep semantic identity consistency between global and local representation, which the consistency is important for the cross-modality pedestrian re-identification task. In this work, we propose a novel cross-modality graph reasoning method (CGRNet) to globally model and reason over relations between modalities and context, and to keep semantic identity consistency between global and local representation. Specifically, we propose a local modality-similarity module to put the distribution of modality-specific features into a common subspace without losing identity information. Besides, we squeeze the input feature of RGB and IR images into a channel-wise global vector, and through graph reasoning, the identity relationship and modality relationship in each vector are inferred. Extensive experiments on two datasets demonstrate the superior performance of our approach over the existing state-of-the-art. The code is available athttps://github.com/fegnyujian/CGRNet.
Yujian Feng, Feng Chen 0047, Yimu Ji 0001, Fei Wu 0004
IEEE Signal Process. Lett.4
2021 IBE-BCIOT: an IBE based cross-chain communication mechanism of blockchain in IoT
Xiaoying Xiao, Weiheng Gu, Yicheng Lu, Shangdong Liu, Fei Wu 0004, Jing He 0004, Yimu Ji 0001, Fen Mei
World Wide Web9
2020 SEBF: A Single-Chain based Extension Model of Blockchain for Fintech
abstract
The traditional blockchain has the shortcoming that a single-chain can only deal with one or a few specific data types. The research question of how to make blockchain be able to deal with various data types has not been well studied. In this paper, we propose a single-chain based extension model of blockchain for fintech (SEBF). In the financial environment, we design a four-layer architecture for this model. By employing the external trusted or-acle group and a financial regulator agency, a variety types of data can be effectively stored in the blockchain, such that the data type extension based on a single-chain is realized. The experimental results indicate that the proposed model can improve the efficiency of simplified payment verifi-cation.
Yimu Ji 0001, Weiheng Gu, Xiaoying Xiao, Shangdong Liu, Jing He 0004, Yunyao Li 0002, Fen Mei, Fei Wu 0004
IJCAI11
2020 Cross-View Image Synthesis with Deformable Convolution and Attention Mechanism
Songsong Wu, Hao Tang 0005, Fei Wu 0004, Guangwei Gao, Xiaoyuan Jing
PRCV (1)4
2020 Low-rank tensor completion for visual data recovery via the tensor train rank-1 decomposition
abstract
In this study, the authors study the problem of tensor completion, in particular for three‐dimensional arrays such as visual data. Previous works have shown that the low‐rank constraint can produce impressive performances for tensor completion. These works are often solved by means of Tucker rank. However, Tucker rank does not capture the intrinsic correlation of the tensor entries. Therefore, the authors propose a new proximal operator for the approximation of tensor nuclear norms based on tensor‐train rank‐1 decomposition via the singular value decomposition. The proximal operator will perform a soft‐thresholding operation on tensor singular values. In addition, the low‐rank constraint can capture the global structure of data well, but it does not exploit local smooth of visual data. Therefore, they integrate total variation as a regularisation term into low‐rank tensor completion. Finally, they use a primal–dual splitting to achieve optimisation. Experimental results have shown that the proposed method, can preserve the multi‐dimensional nature inherent in the data, and thus provide superior results over many state‐of‐the‐art tensor completion techniques.
Xiaoyuan Jing, Guijin Tang, Fei Wu 0004, Xiwei Dong
IET Image Process.4
2020 Manifold embedded distribution adaptation for cross-project defect prediction
abstract
Cross‐project defect prediction (CPDP) technology refers to the constructing prediction model to predict the instance label of the target project by utilising labelled data from an external project. The challenge of CPDP methods is the distribution difference between the data from different projects. Transfer learning can transfer the knowledge from the source domain to the target domain with the aim to minimise the domain difference between different domains. However, most existing methods reduce the distribution discrepancy in the original feature space, where the features are high‐dimensional and non‐linear, which makes it hard to reduce the distribution distance between different projects. Moreover, previous works mainly consider marginal distribution or conditional distribution difference. In this study, the authors proposed a manifold embedded distribution adaptation (MDA) approach to narrow the distribution gap in manifold feature subspace. MDA maps source and target project data to manifold subspace and then joint distribution adaptation of conditional and marginal distributions is performed on manifold subspace. To evaluate the effectiveness of MDA, the authors perform extensive experiments on 20 public projects with three indicators. The experiment results show that MDA improves the average performance, but the improvement is not statistically significant in comparison to HYDRA (one of the baselines).
Ying Sun 0023, Xiaoyuan Jing, Fei Wu 0004, Yanfei Sun
IET Softw.3
2020 Dynamic attention network for semantic segmentation
Fei Wu 0004, Feng Chen 0047, Xiaoyuan Jing, Changhui Hu 0001, Qi Ge, Yimu Ji 0001
Neurocomputing1
2020 Multi-view semantic learning network for point cloud based 3D object detection
Yongguang Yang, Feng Chen 0047, Fei Wu 0004, Deliang Zeng, Yimu Ji 0001, Xiaoyuan Jing
Neurocomputing3
2020 Unsupervised visual domain adaptation via discriminative dictionary evolution
Songsong Wu, Guangwei Gao, Fei Wu 0004, Xiaoyuan Jing
Pattern Anal. Appl.4
2020 Modality-specific and shared generative adversarial network for cross-modal retrieval
Fei Wu 0004, Xiaoyuan Jing, Zhiyong Wu 0006, Yimu Ji 0001, Xiwei Dong, Xiaokai Luo, Qinghua Huang, Ruchuan Wang 0001
Pattern Recognit.1
2020 Intraspectrum Discrimination and Interspectrum Correlation Analysis Deep Network for Multispectral Face Recognition
abstract
Multispectral images contain rich recognition information since the multispectral camera can reveal information that is not visible to the human eye or to the conventional RGB camera. Due to this characteristic of multispectral images, multispectral face recognition has attracted lots of research interest. Although some multispectral face recognition methods have been presented in the last decade, how to fully and effectively explore the intraspectrum discriminant information and the useful interspectrum correlation information in multispectral face images for recognition has not been well studied. To boost the performance of multispectral face recognition, we propose an intraspectrum discrimination and interspectrum correlation analysis deep network (IDICN) approach. Multiple spectra are divided into several spectrum-sets, with each containing a group of spectra within a small spectral range. The IDICN network contains a set of spectrum-set-specific deep convolutional neural networks attempting to extract spectrum-set-specific features, followed by a spectrum pooling layer, whose target is to select a group of spectra with favorable discriminative abilities adaptively. IDICN jointly learns the nonlinear representations of the selected spectra, such that the intraspectrum Fisher loss and the interspectrum discriminant correlation are minimized. Experiments on the well-known Hong Kong Polytechnic University, Carnegie Mellon University, and the University of Western Australia multispectral face datasets demonstrate the superior performance of the proposed approach over several state-of-the-art methods.
Fei Wu 0004, Xiaoyuan Jing, Xiwei Dong, Ruimin Hu, Dong Yue 0001, Lina Wang 0001, Yimu Ji 0001, Ruchuan Wang 0001, Guoliang Chen 0008
IEEE Trans. Cybern.1
2020 Toward Driver Face Recognition in the Intelligent Traffic Monitoring Systems
abstract
This paper models the driver face recognition problem under the intelligent traffic monitoring systems as severe illumination variation face recognition with single sample problem. Firstly, in the point of view of numerical value sign, the current illumination invariant unit is derived from the subtraction of two pixels in the face local region, which may be positive or negative, we propose a generalized illumination robust (GIR) model based on positive and negative illumination invariant units to tackle severe illumination variations. Then, the GIR model can be used to generate several GIR images based on the local edge-region or the local block-region, which results in the edge-region based GIR (EGIR) image or the block-region based GIR (BGIR) image. For single GIR image based classification, the GIR image utilizes the saturation function and the nearest neighbor classifier, which can develop EGIR-face and BGIR-face. For multi GIR images based classification, the GIR images employ the extended sparse representation classification (ESRC) as the classifier that can form the EGIR image based classification (GIRC) and the BGIR image based classification (BGIRC). Further, the GIR model is integrated with the pre-trained deep learning (PDL) model to construct the GIR-PDL model. Finally, the performances of the proposed methods are verified on the Extended Yale B, CMU PIE, AR, self-built Driver and VGGFace2 face databases. The experimental results indicate that the proposed methods are efficient to tackle severe illumination variations.
Changhui Hu 0001, Yang Zhang 0067, Fei Wu 0004, Xiaobo Lu, Pan Liu 0013, Xiaoyuan Jing
IEEE Trans. Intell. Transp. Syst.3
2019 A Cost-Sensitive Shared Hidden Layer Autoencoder for Cross-Project Defect Prediction
Juanjuan Li, Xiaoyuan Jing, Fei Wu 0004, Ying Sun 0023, Yongguang Yang
PRCV (3)3
2019 Modality Consistent Generative Adversarial Network for Cross-Modal Retrieval
Zhiyong Wu 0006, Fei Wu 0004, Xiaokai Luo, Xiwei Dong, Cailing Wang, Xiaoyuan Jing
PRCV (3)2
2019 Adversarial Domain Alignment Feature Similarity Enhancement Learning for Unsupervised Domain Adaptation
Fei Wu 0004, Ying Sun 0023, Songsong Wu, Xiaoyuan Jing
PRCV (3)2
2019 Semi-supervised Multi-view Individual and Sharable Feature Learning for Webpage Classification
abstract
Semi-supervised multi-view feature learning (SMFL) is a feasible solution for webpage classification. However, how to fully extract the complementarity and correlation information effectively under semi-supervised setting has not been well studied. In this paper, we propose a semi-supervised multi-view individual and sharable feature learning (SMISFL) approach, which jointly learns multiple view-individual transformations and one sharable transformation to explore the view-specific property for each view and the common property across views. We design a semi-supervised multi-view similarity preserving term, which fully utilizes the label information of labeled samples and similarity information of unlabeled samples from both intra-view and inter-view aspects. To promote learning of diversity, we impose a constraint on view-individual transformation to make the learned view-specific features to be statistically uncorrelated. Furthermore, we train a linear classifier, such that view-specific and shared features can be effectively combined for classification. Experiments on widely used webpage datasets demonstrate that SMISFL can significantly outperform state-of-the-art SMFL and webpage classification methods.
Fei Wu 0004, Xiaoyuan Jing, Yimu Ji 0001, Chao Lan, Qinghua Huang, Ruchuan Wang 0001
WWW1
2019 General logarithm difference model for severe illumination variation face recognition
Changhui Hu 0001, Xiaobo Lu, Fei Wu 0004, Songsong Wu, Xiaoyuan Jing
Multim. Tools Appl.3
2019 Semi-supervised multiple kernel intact discriminant space learning for image recognition
Xiwei Dong, Fei Wu 0004, Xiaoyuan Jing
Neural Comput. Appl.2
2019 Multi-view Intact Discriminant Space Learning for Image Classification
Xiwei Dong, Fei Wu 0004, Xiaoyuan Jing, Songsong Wu
Neural Process. Lett.2
2018 Cost-sensitive transfer kernel canonical correlation analysis for heterogeneous defect prediction
Zhiqiang Li 0003, Xiaoyuan Jing, Fei Wu 0004, Xiaoke Zhu, Baowen Xu
Autom. Softw. Eng.3
2018 Low-rank representation for semi-supervised software defect prediction
abstract
Software defect prediction based on machine learning is an active research topic in the field of software engineering. The historical defect data in software repositories may contain noises because automatic defect collection is based on modified logs and defect reports. When the previous defect labels of modules are limited, predicting the defect‐prone modules becomes a challenging problem. In this study, the authors propose a graph‐based semi‐supervised defect prediction approach to solve the problems of insufficient labelled data and noisy data. Graph‐based semi‐supervised learning methods used the labelled and unlabelled data simultaneously and consider them as the nodes of the graph at the training phase. Therefore, they solve the problem of insufficient labelled samples. To improve the stability of noisy defect data, a powerful clustering method, low‐rank representation (LRR), and neighbourhood distance are used to construct the relationship graph of samples. Therefore, they propose a new semi‐supervised defect prediction approach, named low‐rank representation‐based semi‐supervised software defect prediction (LRRSSDP). The widely used datasets from NASA projects and noisy datasets are employed as test data to evaluate the performance. Experimental results show that (i) LRRSSDP outperforms several representative state‐of‐the‐art semi‐supervised defect prediction methods; and (ii) LRRSSDP can maintain robustness in noisy environments.
Zhiwu Zhang 0002, Xiaoyuan Jing, Fei Wu 0004
IET Softw.3
2018 Multi-view local discrimination and canonical correlation analysis for image classification
Xiaoyuan Jing, Fei Wu 0004
Neurocomputing3
2018 "Like charges repulsion and opposite charges attraction" law based multilinear subspace analysis for face recognition
Fei Wu 0004, Xiaoyuan Jing, Songsong Wu, Guangwei Gao, Qi Ge, Ruchuan Wang 0001
Knowl. Based Syst.1
2018 Cross-Project and Within-Project Semisupervised Software Defect Prediction: A Unified Approach
abstract
When there exist not enough historical defect data for building an accurate prediction model, semisupervised defect prediction (SSDP) and cross-project defect prediction (CPDP) are two feasible solutions. Existing CPDP methods assume that the available source data are well labeled. However, due to expensive human efforts for labeling a large amount of defect data, usually, we can only utilize the suitable unlabeled source data. We call CPDP in this scenario as cross-project semisupervised defect prediction (CSDP). Although some within-project semisupervised defect prediction (WSDP) methods have been developed in recent years, there still exists much room for improvement on prediction performance. In this paper, we aim to provide a unified and effective solution for both CSDP and WSDP problems. We introduce the semisupervised dictionary learning technique and propose a cost-sensitive kernelized semisupervised dictionary learning (CKSDL) approach. CKSDL can make full use of the limited labeled defect data and a large amount of unlabeled data in the kernel space. In addition, CKSDL considers the misclassification costs in the dictionary learning process. Extensive experiments on 16 projects indicate that CKSDL outperforms state-of-the-art WSDP methods, using unlabeled cross-project defect data can help improve the WSDP performance, and CKSDL generally obtains significantly better prediction performance than related SSDP methods in the CSDP scenario.
Fei Wu 0004, Xiaoyuan Jing, Ying Sun 0023, Fangyi Cui, Yanfei Sun
IEEE Trans. Reliab.1
2017 Semi-Supervised Multi-View Correlation Feature Learning with Application to Webpage Classification
abstract
Webpage classification has attracted a lot of research interest. Webpage data is often multi-view and high-dimensional, and the webpage classification application is usually semi-supervised. Due to these characteristics, using semi-supervised multi-view feature learning (SMFL) technique to deal with the webpage classification problem has recently received much attention. However, there still exists room for improvement for this kind of feature learning technique. How to effectively utilize the correlation information among multi-view of webpage data is an important research topic. Correlation analysis on multi-view data can facilitate extraction of the complementary information. In this paper, we propose a novel SMFL approach, named semi-supervised multi-view correlation feature learning (SMCFL), for webpage classification. SMCFL seeks for a discriminant common space by learning a multi-view shared transformation in a semi-supervised manner. In the discriminant space, the correlation between intra-class samples is maximized, and the correlation between inter-class samples and the global correlation among both labeled and unlabeled samples are minimized simultaneously. We transform the matrix-variable based nonconvex objective function of SMCFL into a convex quadratic programming problem with one real variable, and can achieve a global optimal solution. Experiments on widely used datasets demonstrate the effectiveness and efficiency of the proposed approach.
Xiaoyuan Jing, Fei Wu 0004, Xiwei Dong, Shiguang Shan, Songcan Chen
AAAI2
2017 Multiset Feature Learning for Highly Imbalanced Data Classification
abstract
With the expansion of data, increasing imbalanced data has emerged. When the imbalance ratio of data is high, most existing imbalanced learning methods decline in classification performance. To address this problem, a few highly imbalanced learning methods have been presented. However, most of them are still sensitive to the high imbalance ratio. This work aims to provide an effective solution for the highly imbalanced data classification problem. We conduct highly imbalanced learning from the perspective of feature learning. We partition the majority class into multiple blocks with each being balanced to the minority class and combine each block with the minority class to construct a balanced sample set. Multiset feature learning (MFL) is performed on these sets to learn discriminant features. We thus propose an uncorrelated cost-sensitive multiset learning (UCML) approach. UCML provides a multiple sets construction strategy, incorporates the cost-sensitive factor into MFL, and designs a weighted uncorrelated constraint to remove the correlation among multiset features. Experiments on five highly imbalanced datasets indicate that: UCML outperforms state-of-the-art imbalanced learning methods.
Fei Wu 0004, Xiaoyuan Jing, Shiguang Shan, Wangmeng Zuo, Jing-Yu Yang 0001
AAAI1
2017 Multi-Kernel Low-Rank Dictionary Pair Learning for Multiple Features Based Image Classification
abstract
Dictionary learning (DL) is an effective feature learning technique, and has led to interesting results in many classification tasks. Recently, by combining DL with multiple kernel learning (which is a crucial and effective technique for combining different feature representation information), a few multi-kernel DL methods have been presented to solve the multiple feature representations based classification problem. However, how to improve the representation capability and discriminability of multi-kernel dictionary has not been well studied. In this paper, we propose a novel multi-kernel DL approach, named multi-kernel low-rank dictionary pair learning (MKLDPL). Specifically, MKLDPL jointly learns a kernel synthesis dictionary and a kernel analysis dictionary by exploiting the class label information. The learned synthesis and analysis dictionaries work together to implement the coding and reconstruction of samples in the kernel space. To enhance the discriminability of the learned multi-kernel dictionaries, MKLDPL imposes the low-rank regularization on the analysis dictionary, which can make samples from the same class have similar representations. We apply MKLDPL for multiple features based image classification task. Experimental results demonstrate the effectiveness of the proposed approach.
Xiaoke Zhu, Xiaoyuan Jing, Fei Wu 0004, Di Wu 0014, Li Cheng 0006, Ruimin Hu
AAAI3
2017 Learning Heterogeneous Dictionary Pair with Feature Projection Matrix for Pedestrian Video Retrieval via Single Query Image
abstract
Person re-identification (re-id) plays an important role in video surveillance and forensics applications. In many cases, person re-id needs to be conducted between image and video clip, e.g., re-identifying a suspect from large quantities of pedestrian videos given a single image of him. We call re-id in this scenario as image to video person re-id (IVPR). In practice, image and video are usually represented with different features, and there usually exist large variations between frames within each video. These factors make matching between image and video become a very challenging task. In this paper, we propose a joint feature projection matrix and heterogeneous dictionary pair learning (PHDL) approach for IVPR. Specifically, PHDL jointly learns an intra-video projection matrix and a pair of heterogeneous image and video dictionaries. With the learned projection matrix, the influence of variations within each video to the matching can be reduced. With the learned dictionary pair, the heterogeneous image and video features can be transformed into coding coefficients with the same dimension, such that the matching can be conducted using coding coefficients. Furthermore, to ensure that the obtained coding coefficients have favorable discriminability, PHDL designs a point-to-set coefficient discriminant term. Experiments on the public iLIDS-VID and PRID 2011 datasets demonstrate the effectiveness of the proposed approach.
Xiaoke Zhu, Xiaoyuan Jing, Fei Wu 0004, Yunhong Wang 0001, Wangmeng Zuo, Wei-Shi Zheng 0001
AAAI3
2017 Discriminant Tensor Dictionary Learning with Neighbor Uncorrelation for Image Set Based Classification
abstract
Image set based classification (ISC) has attracted lots of research interest in recent years. Several ISC methods have been developed, and dictionary learning technique based methods obtain state-of-the-art performance. However, existing ISC methods usually transform the image sample of a set into a vector for subsequent processing, which breaks the inherent spatial structure of image sample and the set. In this paper, we utilize tensor to model an image set with two spatial modes and one set mode, which can fully explore the intrinsic structure of image set. We propose a novel ISC approach, named discriminant tensor dictionary learning with neighbor uncorrelation (DTDLNU), which jointly learns two spatial dictionaries and one set dictionary. The spatial and set dictionaries are composed by set-specific sub-dictionaries corresponding to the class labels, such that the reconstruction error is discriminative. To obtain dictionaries with favorable discriminative power, DTDLNU designs a neighbor-uncorrelated discriminant tensor dictionary term, which minimizes the within-class scatter of the training sets in the projected tensor space and reduces tensor dictionary correlation among set-specific sub-dictionaries corresponding to neighbor sets from different classes. Experiments on three challenging datasets demonstrate the effectiveness of DTDLNU.
Fei Wu 0004, Xiaoyuan Jing, Wangmeng Zuo, Xiaoke Zhu
IJCAI1
2017 Large-scale image recognition based on parallel kernel supervised and semi-supervised subspace learning
Fei Wu 0004, Xiaoyuan Jing, Qian Liu 0010, Songsong Wu
Neural Comput. Appl.1
2017 Multi-view Discriminant Dictionary Learning via Learning View-specific and Shared Structured Dictionaries for Image Classification
Fei Wu 0004, Xiaoyuan Jing, Dong Yue 0001
Neural Process. Lett.1
2017 Image denoising using weighted nuclear norm minimization with multiple strategies
Xiaoyuan Jing, Guijin Tang, Fei Wu 0004, Qi Ge
Signal Process.4
2017 Structure-Based Low-Rank Model With Graph Nuclear Norm Regularization for Noise Removal
abstract
Nonlocal image representation methods, including group-based sparse coding and block-matching 3-D filtering, have shown their great performance in application to low-level tasks. The nonlocal prior is extracted from each group consisting of patches with similar intensities. Grouping patches based on intensity similarity, however, gives rise to disturbance and inaccuracy in estimation of the true images. To address this problem, we propose a structure-based low-rank model with graph nuclear norm regularization. We exploit the local manifold structure inside a patch and group the patches by the distance metric of manifold structure. With the manifold structure information, a graph nuclear norm regularization is established and incorporated into a low-rank approximation model. We then prove that the graph-based regularization is equivalent to a weighted nuclear norm and the proposed model can be solved by a weighted singular-value thresholding algorithm. Extensive experiments on additive white Gaussian noise removal and mixed noise removal demonstrate that the proposed method achieves a better performance than several state-of-the-art algorithms.
Qi Ge, Xiaoyuan Jing, Fei Wu 0004, Zhihui Wei, Liang Xiao 0001, Wenze Shao, Dong Yue 0001, Haibo Li 0001
IEEE Trans. Image Process.3
2017 Super-Resolution Person Re-Identification With Semi-Coupled Low-Rank Discriminant Dictionary Learning
abstract
Person re-identification has been widely studied due to its importance in surveillance and forensics applications. In practice, gallery images are high resolution (HR), while probe images are usually low resolution (LR) in the identification scenarios with large variation of illumination, weather, or quality of cameras. Person re-identification in this kind of scenarios, which we call super-resolution (SR) person re-identification, has not been well studied. In this paper, we propose a semi-coupled low-rank discriminant dictionary learning (SLD2L) approach for SR person re-identification task. With the HR and LR dictionary pair and mapping matrices learned from the features of HR and LR training images, SLD2L can convert the features of the LR probe images into HR features. To ensure that the converted features have favorable discriminative capability and the learned dictionaries can well characterize intrinsic feature spaces of the HR and LR images, we design a discriminant term and a low-rank regularization term for SLD2L. Moreover, considering that low resolution results in different degrees of loss for different types of visual appearance features, we propose a multi-view SLD2L (MVSLD2L) approach, which can learn the type-specific dictionary pair and mappings for each type of feature. Experimental results on multiple publicly available data sets demonstrate the effectiveness of our proposed approaches for the SR person re-identification task.
Xiaoyuan Jing, Xiaoke Zhu, Fei Wu 0004, Ruimin Hu, Xinge You, Yunhong Wang 0001, Jing-Yu Yang 0001
IEEE Trans. Image Process.3
2017 An Improved SDA Based Defect Prediction Framework for Both Within-Project and Cross-Project Class-Imbalance Problems
abstract
Background.Solving the class-imbalance problem of within-project software defect prediction (SDP) is an important research topic. Although some class-imbalance learning methods have been presented, there exists room for improvement. For cross-project SDP, we found that the class-imbalanced source usually leads to misclassification of defective instances. However, only one work has paid attention to this cross-project class-imbalance problem.Objective.We aim to provide effective solutions for both within-project and cross-project class-imbalance problems.Method.Subclass discriminant analysis (SDA), an effective feature learning method, is introduced to solve the problems. It can learn features with more powerful classification ability from original metrics. For within-project prediction, we improve SDA for achieving balanced subclasses and propose the improved SDA (ISDA) approach. For cross-project prediction, we employ the semi-supervised transfer component analysis (SSTCA) method to make the distributions of source and target data consistent, and propose the SSTCA+ISDA prediction approach.Results. Extensive experiments on four widely used datasets indicate that: 1) ISDA-based solution performs better than other state-of-the-art methods for within-project class-imbalance problem; 2) SSTCA+ISDA proposed for cross-project class-imbalance problem significantly outperforms related methods.Conclusion. Within-project and cross-project class-imbalance problems greatly affect prediction performance, and we provide a unified and effective prediction framework for both problems.
Xiaoyuan Jing, Fei Wu 0004, Xiwei Dong, Baowen Xu
IEEE Trans. Software Eng.2
2016 Distance learning by treating negative samples differently and exploiting impostors with symmetric triplet constraint for person re-identification
abstract
Distance learning (DL) is an effective technique for person reidentification (PR-ID). DL based methods learn the distance metric by exploiting the discriminative information contained in samples. In PR-ID, different types of negative samples own different amounts of discriminative information, and impostor samples usually own more than other well separable negative samples (WSN-samples). Therefore, how to make full use of the different discriminative information conveyed by all negative samples in the DL process is a critical issue to be investigated. In this paper, we propose a novel DL approach for PR-ID. Specifically, for each target sample, we divide its negative samples into impostors and WSN-samples. Then we learn the distance metric by utilizing impostors and WSN-samples differently. For impostors, we design a symmetric triplet constraint, which requires the impostor to be far away from both samples of its corresponding positive sample pair simultaneously; for WSN-samples, we require them to keep their favorable separability. Experimental results on three benchmark datasets demonstrate the effectiveness and efficiency of our approach.
Xiaoke Zhu, Xiaoyuan Jing, Fei Wu 0004, Wei-Shi Zheng 0001, Ruimin Hu, Chunxia Xiao, Chao Liang 0001
ICME3
2016 Missing data imputation based on low-rank recovery and semi-supervised regression for software effort estimation
abstract
Software effort estimation (SEE) is a crucial step in software development. Effort data missing usually occurs in real-world data collection. Focusing on the missing data problem, existing SEE methods employ the deletion, ignoring, or imputation strategy to address the problem, where the imputation strategy was found to be more helpful for improving the estimation performance. Current imputation methods in SEE use classical imputation techniques for missing data imputation, yet these imputation techniques have their respective disadvantages and might not be appropriate for effort data. In this paper, we aim to provide an effective solution for the effort data missing problem. Incompletion includes the drive factor missing case and effort label missing case. We introduce the low-rank recovery technique for addressing the drive factor missing case. And we employ the semi-supervised regression technique to perform imputation in the case of effort label missing. We then propose a novel effort data imputation approach, named low-rank recovery and semi-supervised regression imputation (LRSRI). Experiments on 7 widely used software effort datasets indicate that: (1) the proposed approach can obtain better effort data imputation effects than other methods; (2) the imputed data using our approach can apply to multiple estimators well.
Xiaoyuan Jing, Fumin Qi, Fei Wu 0004, Baowen Xu
ICSE3
2016 Video-Based Person Re-Identification by Simultaneously Learning Intra-Video and Inter-Video Distance Metrics
Xiaoke Zhu, Xiaoyuan Jing, Fei Wu 0004
IJCAI3
2016 Privacy preserving via interval covering based subclass division and manifold learning based bi-directional obfuscation for effort estimation
abstract
When a company lacks local data in hand, engineers can build an effort model for the effort estimation of a new project by utilizing the training data shared by other companies. However, one of the most important obstacles for data sharing is the privacy concerns of software development organizations. In software engineering, most of existing privacy-preserving works mainly focus on the defect prediction, or debugging and testing, yet the privacy-preserving data sharing problem has not been well studied in effort estimation. In this paper, we aim to provide data owners with an effective approach of privatizing their data before release. We firstly design an Interval Covering based Subclass Division (ICSD) strategy. ICSD can divide the target data into several subclasses by digging a new attribute (i.e., class label) from the effort data. And the obtained class label is beneficial to maintaining the distribution of the target data after obfuscation. Then, we propose a manifold learning based bi-directional data obfuscation (MLBDO) algorithm, which uses two nearest neighbors, which are selected respectively from the previous and next subclasses by utilizing the manifold learning based nearest neighbor selector, as the disturbances to obfuscate the target sample. We call the entire approach as ICSD&MLBDO. Experimental results on seven public effort datasets show that: 1) ICSD&MLBDO can guarantee the privacy and maintain the utility of obfuscated data. 2) ICSD&MLBDO can achieve better privacy and utility than the compared privacy-preserving methods.
Fumin Qi, Xiaoyuan Jing, Xiaoke Zhu, Fei Wu 0004, Li Cheng 0006
ASE4
2016 Group recursive discriminant subspace learning with image set decomposition
Fei Wu 0004, Xiaoyuan Jing, Yong-Fang Yao, Dong Yue 0001, Jun Chen 0001
Neural Comput. Appl.1
2016 Multi-spectral low-rank structured dictionary learning for face recognition
Xiaoyuan Jing, Fei Wu 0004, Xiaoke Zhu, Xiwei Dong, Fei Ma 0004, Zhiqiang Li 0003
Pattern Recognit.2
2016 Uncorrelated multi-set feature learning for color face recognition
Fei Wu 0004, Xiaoyuan Jing, Xiwei Dong, Qi Ge, Songsong Wu, Qian Liu 0010, Dong Yue 0001, Jing-Yu Yang 0001
Pattern Recognit.1
2016 Multi-view low-rank dictionary learning for image classification
Fei Wu 0004, Xiaoyuan Jing, Xinge You, Dong Yue 0001, Ruimin Hu, Jing-Yu Yang 0001
Pattern Recognit.1
2016 Multi-Label Dictionary Learning for Image Annotation
abstract
Image annotation has attracted a lot of research interest, and multi-label learning is an effective technique for image annotation. How to effectively exploit the underlying correlation among labels is a crucial task for multi-label learning. Most existing multi-label learning methods exploit the label correlation only in the output label space, leaving the connection between the label and the features of images untouched. Although, recently some methods attempt toward exploiting the label correlation in the input feature space by using the label information, they cannot effectively conduct the learning process in both the spaces simultaneously, and there still exists much room for improvement. In this paper, we propose a novel multi-label learning approach, named multi-label dictionary learning (MLDL) with label consistency regularization and partial-identical label embedding MLDL, which conducts MLDL and partial-identical label embedding simultaneously. In the input feature space, we incorporate the dictionary learning technique into multi-label learning and design the label consistency regularization term to learn the better representation of features. In the output label space, we design the partial-identical label embedding, in which the samples with exactly same label set can cluster together, and the samples with partial-identical label sets can collaboratively represent each other. Experimental results on the three widely used image datasets, including Corel 5K, IAPR TC12, and ESP Game, demonstrate the effectiveness of the proposed approach.
Xiaoyuan Jing, Fei Wu 0004, Zhiqiang Li 0003, Ruimin Hu, David Zhang 0001
IEEE Trans. Image Process.2
2015 Super-resolution Person re-identification with semi-coupled low-rank discriminant dictionary learning
abstract
Person re-identification has been widely studied due to its importance in surveillance and forensics applications. In practice, gallery images are high-resolution (HR) while probe images are usually low-resolution (LR) in the identification scenarios with large variation of illumination, weather or quality of cameras. Person re-identification in this kind of scenarios, which we call super-resolution (SR) person re-identification, has not been well studied. In this paper, we propose a semi-coupled low-rank discriminant dictionary learning (SLD2L) approach for SR person re-identification. For the given training image set which consists of HR gallery and LR probe images, we aim to convert the features of LR images into discriminating HR features. Specifically, our approach learns a pair of HR and LR dictionaries and a mapping from the features of HR gallery images and LR probe images. To ensure that the converted features using the learned dictionaries and mapping have favorable discriminative capability, we design a discriminant term which requires the converted HR features of LR probe images should be close to the features of HR gallery images from the same person, but far away from the features of HR gallery images from different persons. In addition, we apply low-rank regularization in dictionary learning procedure such that the learned dictionaries can well characterize intrinsic feature space of HR and LR images. Experimental results on public datasets demonstrate the effectiveness of SLD2L.
Xiaoyuan Jing, Xiaoke Zhu, Fei Wu 0004, Xinge You, Qinglong Liu, Dong Yue 0001, Ruimin Hu, Baowen Xu
CVPR3
2015 Web Page Classification Based on Uncorrelated Semi-Supervised Intra-View and Inter-View Manifold Discriminant Feature Extraction
Xiaoyuan Jing, Qian Liu 0010, Fei Wu 0004, Baowen Xu, Yang-Ping Zhu, Songcan Chen
IJCAI3
2015 Heterogeneous cross-company defect prediction by unified metric representation and CCA-based transfer learning
abstract
Cross-company defect prediction (CCDP) learns a prediction model by using training data from one or multiple projects of a source company and then applies the model to the target company data. Existing CCDP methods are based on the assumption that the data of source and target companies should have the same software metrics. However, for CCDP, the source and target company data is usually heterogeneous, namely the metrics used and the size of metric set are different in the data of two companies. We call CCDP in this scenario as heterogeneous CCDP (HCCDP) task. In this paper, we aim to provide an effective solution for HCCDP. We propose a unified metric representation (UMR) for the data of source and target companies. The UMR consists of three types of metrics, i.e., the common metrics of the source and target companies, source-company specific metrics and target-company specific metrics. To construct UMR for source company data, the target-company specific metrics are set as zeros, while for UMR of the target company data, the source-company specific metrics are set as zeros. Based on the unified metric representation, we for the first time introduce canonical correlation analysis (CCA), an effective transfer learning method, into CCDP to make the data distributions of source and target companies similar. Experiments on 14 public heterogeneous datasets from four companies indicate that: 1) for HCCDP with partially different metrics, our approach significantly outperforms state-of-the-art CCDP methods; 2) for HCCDP with totally different metrics, our approach obtains comparable prediction performances in contrast with within-project prediction results. The proposed approach is effective for HCCDP.
Xiaoyuan Jing, Fei Wu 0004, Xiwei Dong, Fumin Qi, Baowen Xu
ESEC/SIGSOFT FSE2
2014 Uncorrelated Multi-View Discrimination Dictionary Learning for Recognition
abstract
Dictionary learning (DL) has now become an important feature learning technique that owns state-of-the-art recognition performance. Due to sparse characteristic of data in real-world applications, DL uses a set of learned dictionary bases to represent the linear decomposition of a data point. Fisher discrimination DL (FDDL) is a representative supervised DL method, which constructs a structured dictionary whose atoms correspond to the class labels. Recent years have witnessed a growing interest in multi-view (more than two views) feature learning techniques. Although some multi-view (or multi-modal) DL methods have been presented, there still exists much room for improvement. How to enhance the total discriminability of dictionaries and reduce their redundancy is a crucial research topic. To boost the performance of multi-view DL technique, we propose an uncorrelated multi-view discrimination DL (UMDDL) approach for recognition. By making dictionary atoms correspond to the class labels such that the obtained reconstruction error is discriminative, UMDDL aims to jointly learn multiple dictionaries with totally favorable discriminative power. Furthermore, we design the uncorrelated constraint for multi-view DL, so as to reduce the redundancy among dictionaries learned from different views. Experiments on several public datasets demonstrate the effectiveness of the proposed approach.
Xiaoyuan Jing, Ruimin Hu, Fei Wu 0004, Xilin Chen 0001, Qian Liu 0010, Yong-Fang Yao
AAAI3