Dapeng Tao

dblp:55/10400 · DBLP profile ↗
← Back
181ranked-venue papers
21as first author
84since 2021 · last 2026
0000-0003-0783-5273ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 89 · 10 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 72 · 7 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 14 · 3 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ProGMLP: A Progressive Framework for GNN-to-MLP Knowledge Distillation with Efficient Trade-offs
abstract
GNN-to-MLP (G2M) methods have emerged as a promising approach to accelerate Graph Neural Networks (GNNs) by distilling their knowledge into simpler Multi-Layer Perceptrons (MLPs). These methods bridge the gap between the expressive power of GNNs and the computational efficiency of MLPs, making them well-suited for resource-constrained environments. However, existing G2M methods are limited by their inability to flexibly adjust inference cost and accuracy dynamically, a critical requirement for real-world applications where computational resources and time constraints can vary significantly. To address this, we introduce a Progressive framework designed to offer flexible and on-demand trade-offs between inference cost and accuracy for GNN-to-MLP knowledge distillation (ProGMLP). ProGMLP employs a Progressive Training Structure (PTS), where multiple MLP students are trained in sequence, each building on the previous one. Furthermore, ProGMLP incorporates Progressive Knowledge Distillation (PKD) to iteratively refine the distillation process from GNNs to MLPs, and Progressive Mixup Augmentation (PMA) to enhance generalization by progressively generating harder mixed samples. Our approach is validated through comprehensive experiments on eight real-world graph datasets, demonstrating that ProGMLP maintains high accuracy while dynamically adapting to varying runtime scenarios, making it highly effective for deployment in diverse application settings.
Weigang Lu 0001, Ziyu Guan, Wei Zhao 0019, Yaming Yang 0002, Yibing Zhan, Dapeng Tao
AAAI8
2026 Cross-modal attention fusion and label co-occurrence feature enhancement for multi-label postoperative adverse reaction prediction
Jierong Li, Yibing Zhan, Dapeng Tao, Chongchong Qi
Expert Syst. Appl.4
2026 Causal Mask in Transformer via Transfer Entropy Estimation from Vector Autoregressive Learning for Multivariate Time Series Forecasting
abstract
Time series forecasting remains challenging in domains such as finance and climate science, where complex interactions among variables often induce spurious correlations. We propose ARCausal, a forecasting framework that integrates transfer entropy (TE)-based causal discovery with Transformer attention modeling. ARCausal introduces a sparse causal masking mechanism derived from TE and refined via vector autoregression (VAR) estimation to capture dynamic causal interactions. The mask suppresses noninformative dependencies and distinguishes autocorrelation from cross-variable causal effects, improving both predictive performance and interpretability. Experiments on nine benchmark datasets demonstrate consistent improvements over strong baselines, achieving up to [Formula: see text] reduction in MSE while maintaining computational efficiency. Visualization results further illustrate the interpretability of the learned causal structures. The code is publicly available at https://github.com/jancely/ARCausal/.
Chengli Zhou, Yaqun Huang, Dapeng Tao, Chunna Zhao
Int. J. Neural Syst.4
2026 Evaluating large language models for real-world perioperative clinical consultation
Yibing Zhan, Baosheng Yu, Pingbo Xu, Lijing Chen, Chong Zhang 0013, Chengli Zhou, Xiongbin Wang, Dapeng Tao
Neurocomputing10
2026 Evolving classifiers with background suppression transformer for open-set long-tailed class-incremental remote sensing scene classification
Sichao Fu, Hongquan Xin, Wuli Wang, Peng Ren 0001, Baodi Liu, Weihua Ou, Dapeng Tao
Neural Networks8
2026 Causal-guided strength differential independence sample weighting for out-of-distribution generalization
Haoran Yu 0005, Weifeng Liu 0001, Yingjie Wang 0007, Baodi Liu, Dapeng Tao, Honglong Chen
Pattern Recognit.5
2026 Joint subgraph independence for graph out-of-distribution generalization
Weifeng Liu 0001, Baodi Liu, Dapeng Tao, Honglong Chen
Pattern Recognit.5
2026 Mind the data: Evaluating data quality sensitivity in medical LLMs
Xiaodong Han, Yibing Zhan, Baosheng Yu, Dapeng Tao
Pattern Recognit. Lett.5
2025 AGMixup: Adaptive Graph Mixup for Semi-supervised Node Classification
abstract
Mixup is a data augmentation technique that enhances model generalization by interpolating between data points using a mixing ratio lambda in the image domain. Recently, the concept of mixup has been adapted to the graph domain through node-centric interpolations. However, these approaches often fail to address the complexity of interconnected relationships, potentially damaging the graph's natural topology and undermining node interactions. Furthermore, current graph mixup methods employ a one-size-fits-all strategy with a randomly sampled lambda for all mixup pairs, ignoring the diverse needs of different pairs. This paper proposes an Adaptive Graph Mixup (AGMixup) framework for semi-supervised node classification. AGMixup introduces a subgraph-centric approach, which treats each subgraph similarly to how images are handled in Euclidean domains, thus facilitating a more natural integration of mixup into graph-based learning. We also propose an adaptive mechanism to tune the mixing ratio lambda for diverse mixup pairs, guided by the contextual similarity and uncertainty of the involved subgraphs. Extensive experiments across seven datasets on semi-supervised node classification benchmarks demonstrate AGMixup's superiority over state-of-the-art graph mixup methods.
Weigang Lu 0001, Ziyu Guan, Wei Zhao 0019, Yaming Yang 0002, Yibing Zhan, Yiheng Lu, Dapeng Tao
AAAI7
2025 Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated Reasoning
abstract
Recently, increasing attention has been focused on improving the ability of Large Language Models (LLMs) to perform complex reasoning. Advanced methods, such as Chain-of-Thought (CoT) and its variants, are found to enhance their reasoning skills by designing suitable prompts or breaking down complex problems into more manageable sub-problems. However, little concentration has been put on exploring the reasoning process, i.e., we discovered that most methods resort to Direct Reasoning (DR) and disregard Indirect Reasoning (IR). This can make LLMs difficult to solve IR tasks, which are often encountered in the real world. To address this issue, we propose a Direct-Indirect Reasoning (DIR) method, which considers DR and IR as multiple parallel reasoning paths that are merged to derive the final answer. We stimulate LLMs to implement IR by crafting prompt templates incorporating the principles of contrapositive and contradiction. These templates trigger LLMs to assume the negation of the conclusion as true, combine it with the premises to deduce a conclusion, and utilize the logical equivalence of the contrapositive to enhance their comprehension of the rules used in the reasoning process. Our DIR method is simple yet effective and can be straightforwardly integrated with existing variants of CoT methods. Experimental results on four datasets related to logical reasoning and mathematic proof demonstrate that our DIR method, when combined with various baseline methods, significantly outperforms all the original methods.
Yanfang Zhang 0001, Yiliu Sun, Yibing Zhan, Dapeng Tao, Dacheng Tao, Chen Gong 0002
COLING4
2025 NT-FAN: A simple yet effective noise-tolerant few-shot adaptation network
Wenjing Yang 0002, Haoang Chi, Yibing Zhan, Xiaoguang Ren, Dapeng Tao, Long Lan
Artif. Intell.6
2025 One Unsupervised Feature Selection Method for the Classical Linear Classifier in Land Coverage Classification With PolSAR Imagery
abstract
ABSTRACT Land coverage mapping and classification is one of the critical information‐based tools for sustainable agricultural development, enabling relevant departments to carry out agricultural resource adjustments, yield predictions, and other tasks in advance. As a vital means of acquiring land cover and usage information, SAR sensors have become an important research direction due to their all‐weather and all‐day working capabilities. Nevertheless, traditional classification methods in PolSAR image classification often input a combination of various scattering features, i.e., high‐dimensional feature combination, into classifiers, leading to mutual interference among different features and consequently degrading classification performance, especially for linear classifiers such as NRS and SVM. To mitigate this interference, this paper proposed an unsupervised feature selection based on spectral clustering (FSSC) that constructs a targeted approach by leveraging the linear expression capabilities of high‐dimensional features. In this method, the linear relationships between different features are first analyzed, and the linear similarity between features can be quantitatively expressed using Pearson correlation coefficients, forming a feature similarity matrix. Subsequently, the similarity matrix undergoes unsupervised similarity partitioning through spectral clustering, dividing the features into distinct combinations. Features within clustering subsets can be considered as combinations with high linear similarity. Therefore, KL divergence is applied to select the most representative features within each cluster, and the resulting representative feature combinations from different clustering subsets are combined to form an optimal feature set, achieving the purpose of feature selection. This method maps high‐dimensional feature combinations into low‐dimensional ones while preserving the essential attributes of the original data, thereby retaining the valuable feature information and enhancing classification performance. Experimental outcomes conclusively show that the proposed method enhances the overall accuracy (OA) of SVM by 4.51% and the OA of NRS by 2.34% in the Flevoland Dataset, underscoring its efficacy in PolSAR image classification, especially for linear classifiers.
Xichao Liu, Dapeng Tao
Comput. Intell.3
2025 Noise-Robust Few-Shot Classification via Variational Adversarial Data Augmentation
abstract
Few-shot classification models trained with clean samples poorly classify samples from the real world with various scales of noise. To enhance the model for recognizing noisy samples, researchers usually utilize data augmentation or use noisy samples generated by adversarial training for model training. However, existing methods still have problems: (i) The effects of data augmentation on the robustness of the model are limited. (ii) The noise generated by adversarial training usually causes overfitting and reduces the generalization ability of the model, which is very significant for few-shot classification. (iii) Most existing methods cannot adaptively generate appropriate noise. Given the above three points, this paper proposes a noise-robust few-shot classification algorithm, VADA—Variational Adversarial Data Augmentation. Unlike existing methods, VADA utilizes a variational noise generator to generate an adaptive noise distribution according to different samples based on adversarial learning, and optimizes the generator by minimizing the expectation of the empirical risk. Applying VADA during training can make few-shot classification more robust against noisy data, while retaining generalization ability. In this paper, we utilize FEAT and ProtoNet as baseline models, and accuracy is verified on several common few-shot classification datasets, including MiniImageNet, TieredImageNet, and CUB. After training with VADA, the classification accuracy of the models increases for samples with various scales of noise.
Baodi Liu, Kai Zhang 0029, Honglong Chen, Dapeng Tao, Weifeng Liu 0001
Comput. Vis. Media5
2025 End-to-End Cascaded Image Restoration and Object Detection for Rain and Fog Conditions
abstract
ABSTRACT Adverse weather conditions in real‐world scenarios can degrade the performance of deep learning‐based object detection models. A commonly used approach is to apply image restoration before object detection to improve degraded images. However, there is no direct correlation between the visual quality of image restoration and the object detection accuracy. Furthermore, image restoration and object detection have potential conflicting objectives, making joint optimisation difficult. To address this, we propose an end‐to‐end object detection network specifically designed for rainy and foggy conditions. Our approach cascades an image restoration subnetwork with a detection subnetwork and optimises them jointly through a shared objective. Specifically, we introduce an expanded dilated convolution block and a weather attention block to enhance the effectiveness and robustness of the restoration network under various weather degradations. Additionally, we incorporate an auxiliary alignment branch with feature alignment loss to align the features of restored and clean images within the detection backbone, enabling joint optimisation of both subnetworks. A novel training strategy is also proposed to further improve object detection performance under rainy and foggy conditions. Extensive experiments on the vehicle‐rain‐fog, VOC‐fog and real‐world fog datasets demonstrate that our method outperforms recent state‐of‐the‐art approaches in image restoration quality and detection accuracy. The code is available at https://github.com/HappyPessimism/RainFog‐Restoration‐Detection .
Peng Li 0039, Dapeng Tao
IET Comput. Vis.3
2025 Sample-Cohesive Pose-Aware Contrastive Facial Representation Learning
abstract
Abstract Self-supervised facial representation learning (SFRL) methods, especially contrastive learning (CL) methods, have been increasingly popular due to their ability to perform face understanding without heavily relying on large-scale well-annotated datasets. However, analytically, current CL-based SFRL methods still perform unsatisfactorily in learning facial representations due to their tendency to learn pose-insensitive features, resulting in the loss of some useful pose details. This could be due to the inappropriate positive/negative pair selection within CL. To conquer this challenge, we propose a Pose-disentangled Contrastive Facial Representation Learning (PCFRL) framework to enhance pose awareness for SFRL. We achieve this by explicitly disentangling the pose-aware features from non-pose face-aware features and introducing appropriate sample calibration schemes for better CL with the disentangled features. In PCFRL, we first devise a pose-disentangled decoder with a delicately designed orthogonalizing regulation to perform the disentanglement; therefore, the learning on the pose-aware and non-pose face-aware features would not affect each other. Then, we introduce a false-negative pair calibration module to overcome the issue that the two types of disentangled features may not share the same negative pairs for CL. Our calibration employs a novel neighborhood-cohesive pair alignment method to identify pose and face false-negative pairs, respectively, and further help calibrate them to appropriate positive pairs. Lastly, we devise two calibrated CL losses, namely calibrated pose-aware and face-aware CL losses, for adaptively learning the calibrated pairs more effectively, ultimately enhancing the learning with the disentangled features and providing robust facial representations for various downstream tasks. In the experiments, we perform linear evaluations on four challenging downstream facial tasks with SFRL using our method, including facial expression recognition, face recognition, facial action unit detection, and head pose estimation. Experimental results show that PCFRL outperforms existing state-of-the-art methods by a substantial margin, demonstrating the importance of improving pose awareness for SFRL. Our evaluation code and model will be available at https://github.com/fulaoze/CV/tree/main .
Yuanyuan Liu 0004, Shaoze Feng, Yibing Zhan, Dapeng Tao, Zijing Chen, Zhe Chen 0013
Int. J. Comput. Vis.5
2025 G-NodeMixup: Enhancing graph neural networks reachability under extremely limited labels
Ziyu Guan, Beilei Ling, Weigang Lu 0001, Meng Yan 0013, Yaming Yang 0002, Wei Zhao 0019, Yibing Zhan, Dapeng Tao
Neurocomputing8
2025 Fractional Position With Predictive Attention for Multivariate Time Series Forecasting
abstract
With the proliferation of the Internet of Things (IoT), a wealth of multivariate time series data is being generated across various domains, creating new demands for accurate and efficient forecasting models. Despite the success of attention-based models in capturing dependencies within time series, they often fail to address two critical challenges: (1) the lag effect between output and input, which can significantly distort predictions, and (2) the limitations of classic trigonometric positional embeddings, which lack scalability and adaptability to diverse temporal patterns. To address these challenges, we propose FPPformer, a novel forecasting model that introduces two key innovations: (i) a Fractional Positional Embedding (FPE), which leverages fractional calculus to enable scalable and adaptive positional representations, and (ii) a Predictive Attention Mechanism (PAM), which explicitly models the lag effect, aligning output and input more effectively. The FPPformer architecture consists of encoder-only structure, with the core of encoder module utilizing the PAM. Experimental results demonstrate that FPPformer significantly improves forecasting perfromance, reducing the mean squared error (MSE) by 28% and the mean absolute error (MAE) by 17% across six datasets spanning four domains -electricity, weather, economy, and transportation -especially on large-scale datasets such as Traffic and Electricity. These results highlight FPPformer’s ability to address fundamental challenges in time series forecasting, providing a new perspective on leveraging positional representations and lag-aware attention mechanisms. The code for this project is available at https://github.com/jancely/FPPformer.
Chengli Zhou, Junjie Ye 0003, Yanli Zhou, Xiaojun Zhou 0004, Yaqun Huang, Dapeng Tao, Chunna Zhao
IEEE Internet Things J.7
2025 WMANet:Weighted multiple adaptive feature attention for self-supervised single remote-sensing image denoising
Weifeng Liu 0001, Dapeng Tao, Baodi Liu, Yanjiang Wang 0001
Knowl. Based Syst.5
2025 PPBU: Progressive Pixel Bank Updating Strategy for Single Remote Sensing Image Denoising
Baodi Liu, Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001
IEEE Geosci. Remote. Sens. Lett.4
2025 Multi-scale region selection network in deep features for full-field mammogram classification
abstract
Early diagnosis and treatment of breast cancer can effectively reduce mortality. Since mammogram is one of the most commonly used methods in the early diagnosis of breast cancer, the classification of mammogram images is an important work of computer-aided diagnosis (CAD) systems. With the development of deep learning in CAD, deep convolutional neural networks have been shown to have the ability to complete the classification of breast cancer tumor patches with high quality, which makes most previous CNN-based full-field mammography classification methods rely on region of interest (ROI) or segmentation annotation to enable the model to locate and focus on small tumor regions. However, the dependence on ROI greatly limits the development of CAD, because obtaining a large number of reliable ROI annotations is expensive and difficult. Some full-field mammography image classification algorithms use multi-stage training or multi-feature extractors to get rid of the dependence on ROI, which increases the computational amount of the model and feature redundancy. In order to reduce the cost of model training and make full use of the feature extraction capability of CNN, we propose a deep multi-scale region selection network (MRSN) in deep features for end-to-end training to classify full-field mammography without ROI or segmentation annotation. Inspired by the idea of multi-example learning and the patch classifier, MRSN filters the feature information and saves only the feature information of the tumor region to make the performance of the full-field image classifier closer to the patch classifier. MRSN first scores different regions under different dimensions to obtain the location information of tumor regions. Then, a few high-scoring regions are selected by location information as feature representations of the entire image, allowing the model to focus on the tumor region. Experiments on two public datasets and one private dataset prove that the proposed MRSN achieves the most advanced performance.
Luhao Sun, Bowen Han 0001, Wenzong Jiang, Weifeng Liu 0001, Baodi Liu, Dapeng Tao, Chao Li 0075
Medical Image Anal.6
2025 Parentheses insertion based sentence-level text adversarial attack
Xinghao Yang, Baodi Liu, Honglong Chen, Dapeng Tao, Weifeng Liu 0001
Multim. Syst.5
2025 Correction: Parentheses insertion based sentence-level text adversarial attack
Xinghao Yang, Baodi Liu, Honglong Chen, Dapeng Tao, Weifeng Liu 0001
Multim. Syst.5
2025 IW-ViT: Independence-Driven Weighting Vision Transformer for out-of-distribution generalization
Weifeng Liu 0001, Haoran Yu 0005, Yingjie Wang 0007, Baodi Liu, Dapeng Tao, Honglong Chen
Pattern Recognit.5
2025 Distilling interaction knowledge for semi-supervised egocentric action recognition
Haoran Wang 0001, Baosheng Yu, Yibing Zhan, Dapeng Tao, Haibin Ling
Pattern Recognit.5
2025 SCAWaveNet: A Spatial-Channel Attention-Based Network for Global Significant Wave Height Retrieval
abstract
Recent advancements in spaceborne GNSS missions have produced extensive global datasets, providing a robust basis for deep learning-based significant wave height (SWH) retrieval. While existing deep learning models predominantly utilize CYGNSS data with four-channel information, they often adopt single-channel inputs or simple channel concatenation without leveraging the benefits of cross-channel information interaction during training. To address this limitation, a novel spatial–channel attention-based network, namely SCAWaveNet, is proposed for SWH retrieval. Specifically, features from each channel of the DDMs are modeled as independent attention heads, enabling the fusion of spatial and channel-wise information. For auxiliary parameters, a lightweight attention mechanism is designed to assign weights along the spatial and channel dimensions. The final feature integrates both spatial and channel-level characteristics. Model performance is evaluated using four-channel CYGNSS data. Quantitative and qualitative experiments were conducted on CYGNSS-ERA5 test set, SCAWaveNet achieves an average RMSE of 0.438 m. Compared to state-of-the-art models, SCAWaveNet reduces RMSE by at least 3.52%. Furthermore, evaluations on WW3, Jason-3, and NDBC buoy data, as well as in wind speed, rainstorm, typhoon and noisy scenarios, further confirm the superiority of SCAWaveNet. The code is available at https://github.com/Clifx9908/SCAWaveNet.
Chong Zhang 0013, Xichao Liu, Jinwei Bu, Yibing Zhan, Dapeng Tao
IEEE Trans. Geosci. Remote. Sens.6
2025 MSC-GAN: A Multistream Complementary Generative Adversarial Network With Grouping Learning for Multitemporal Cloud Removal
abstract
Optical remote sensing images have extensive application value, but cloud contamination greatly limits their potential use in the field of geographic information. Cloud removal aims to restore clear, unobstructed images from cloud-covered ones for subsequent in-depth analysis. Due to severe cloud cover problems such as thick clouds in some areas of remote sensing images, cloud removal tasks have become challenging. Recently, many methods have attempted to incrementally fill in obscured regions by fusing cloud-free information from multitemporal data. However, most of these methods fail to effectively utilize the interaction among different temporal data, and some information of data is easily lost in the process of deep transmission, this causes problems such as inadequate cloud removal and blurred recovery of ground under the clouds. Therefore, we propose a multistream complementary generative adversarial network (MSC-GAN) for cloud removal using multitemporal data. First, it employs a multistream complementary (MSC) architecture in the down-sampling feature encoding stage to effectively promote the interaction of feature information across multitemporal data, alleviating information loss as network depth increases. Second, to reduce the feature blur, we design a group feature reweighting (GFR) module as a complementary connection of long-distance information, in which the grouping learning and multidimensional parallel architecture can cost-effectively enhance semantic fusion between low-level and high-level features. Moreover, a channel enhancement method is introduced to assist in processing the underlying transition information, minimizing the interference of invalid information. Experimental results on multiple benchmark datasets under a series of image quality assessment metrics demonstrate the effectiveness of the proposed method.
Yanjiang Wang 0001, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Trans. Geosci. Remote. Sens.4
2025 See Degraded Objects: A Physics-Guided Approach for Object Detection in Adverse Environments
abstract
In adverse environments, the detector often fails to detect degraded objects because they are almost invisible and their features are weakened by the environment. Common approaches involve image enhancement to support detection, but they inevitably introduce human-invisible noise that negatively impacts the detector. In this work, we propose a physics-guided approach for object detection in adverse environments, which gives a straightforward solution that injects the physical priors into the detector, enabling it to detect poorly visible objects. The physical priors, derived from the imaging mechanism and image property, include environment prior and frequency prior. The environment prior is generated from the physical model, e.g., the atmospheric model, which reflects the density of environmental noise. The frequency prior is explored based on an observation that the amplitude spectrum could highlight object regions from the background. The proposed two priors are complementary in principle. Furthermore, we present a physics-guided loss that incorporates a novel weight item, which is estimated by applying the membership function on physical priors and could capture the extent of degradation. By backpropagating the physics-guided loss, physics knowledge is injected into the detector to aid in locating degraded objects. We conduct experiments in synthetic foggy environment, real foggy environment, and real underwater scenario. The results demonstrate that our method is effective and achieves state-of-the-art performance. The code is available at https://github.com/PangJian123/See-Degraded-Objects.
Weifeng Liu 0001, Jian Pang, Bingfeng Zhang, Baodi Liu, Dapeng Tao
IEEE Trans. Image Process.6
2025 Disentangling Inter- and Intra-Video Relations for Multi-Event Video-Text Retrieval and Grounding
abstract
Video-text retrieval aims to precisely search for videos most relevant to text queries within a video corpus. However, existing methods are largely limited to single-text (single-event) queries and are not effective at handling multi-text (multi-event) queries. Furthermore, these methods typically focus solely on retrieval and do not attempt to locate multiple events within the retrieved videos. To address these limitations, our paper proposes a novel method named Disentangling Inter- and Intra-Video Relations, which jointly addresses multi-event video-text retrieval and grounding. This method leverages both inter-video and intra-video event relationships to enhance retrieval and grounding performance. At the retrieval level, we devise a Relational Event-Centric Video-Text Retrieval module based on the principle that comprehensive textual information leads to precise correspondence between text and video. It incorporates event relationship features at different hierarchical levels and exploits the hierarchical structure of video relationships to achieve multi-level contrastive learning between events and videos. This approach enhances the richness, accuracy, and comprehensiveness of event descriptions, improving alignment precision between text and video and enabling effective differentiation among videos. For event grounding, we propose Event Contrast-Driven Video Grounding, which accounts for positional differences among events on the 2D temporal score map and achieves precise grounding of multiple events through divergence learning for their locations. Our solution not only provides efficient text-to-video retrieval but also accurately grounds events within the retrieved videos, addressing the shortcomings of existing methods. Extensive experimental results on the ActivityNet Captions and Charades-STA benchmark datasets demonstrate the superior performance of our method, validating its effectiveness. The innovation of this research lies in introducing a new joint framework for video-text retrieval and multi-event grounding while offering new ideas for further research and applications in related fields. The code is available at https://github.com/X7J92/MVT-RG.
Mengzhao Wang 0002, Huafeng Li 0001, Jinxing Li 0003, Dapeng Tao, Zhengtao Yu 0001
IEEE Trans. Image Process.5
2025 Cps-STS: Bridging the Gap Between Content and Position for Coarse-Point-Supervised Scene Text Spotter
abstract
Recently, weakly supervised methods for scene text spotter are increasingly popular with researchers due to their potential to significantly reduce dataset annotation efforts. The latest progress in this field is text spotter based on single or multi-point annotations. However, this method struggles with the sensitivity of text recognition to the precise annotation location and fails to capture the relative positions and shapes of characters, leading to impaired recognition of texts with extensive rotations and flips. To address these challenges, this paper develops a novel method named Coarse-point-supervised Scene Text Spotter (Cps-STS). Cps-STS first utilizes a few approximate points as text location labels and introduces a learnable position modulation mechanism, easing the accuracy requirements for annotations and enhancing model robustness. Additionally, we incorporate a Spatial Compatibility Attention (SCA) module for text decoding to effectively utilize spatial data such as position and shape. This module fuses compound queries and global feature maps, serving as a bias in the SCA module to express text spatial morphology. In order to accurately locate and decode text content, we introduce features containing spatial morphology information and text content into the input features of the text decoder. By introducing features with spatial morphology information as bias terms into the text decoder, ablation experiments demonstrate that this operation enables the model to effectively identify and utilize the relationship between text content and position to enhance the recognition performance of our model. One significant advantage of Cps-STS is its ability to achieve full supervision-level performance with just a few imprecise coarse points at a low cost. Extensive experiments validate the effectiveness and superiority of Cps-STS over existing approaches.
Weida Chen, Jie Jiang 0015, Linfei Wang, Huafeng Li 0001, Yibing Zhan, Dapeng Tao
IEEE Trans. Multim.6
2025 Progressive Feature Mining and External Knowledge-Assisted Text-Pedestrian Image Retrieval
abstract
Text-Pedestrian Image Retrieval employs textual description of pedestrian's appearance to identify the corresponding pedestrian image. This task involves modality discrepancy and the challenges posed by textual diversity of pedestrians with the same identity. Although advancements have been made in text-pedestrian image retrieval, current methods do not comprehensively address these challenges. Thus, this paper proposes a progressive feature mining and external knowledge- assisted feature purification method. Specifically, we implement a progressive mining mode, enabling the model to extract discriminative features from overlooked information. This enhances the model's feature representation capabilities and prevents the loss of discriminative information. To further mitigate the challenges posed by modality discrepancy and text diversity in cross-modal matching, we propose to use external knowledge of other samples from the same modality. This approach accentuates identity-consistent features and diminishes identity-inconsistent ones, refining feature representation and reducing interference from textual diversity and negative sample correlation features of the same modality. Extensive experiments on three challenging datasets demonstrate the effectiveness and superiority of the proposed method, with its retrieval performance outstripping that of large-scale model-based methods on large-scale datasets.
Huafeng Li 0001, Shedan Yang, Dapeng Tao, Zhengtao Yu 0001
IEEE Trans. Multim.4
2025 Facial Expression Recognition With Heatmap Neighbor Contrastive Learning
abstract
Many supervised learning-based facial expression recognition (FER) methods achieve good performance with the assistance of expression labels and a complex framework. However, there are inconsistent annotations in different expression datasets, making the above methods disadvantageous for new expression datasets or datasets with limited training data. The objective of this paper is to learn self-supervised facial expression features that enable the FER model not to rely on the annotation consistency of the different datasets. Most current self-supervised learning algorithms based on contrastive learning learn the representation by forcing different augmented views of the same image close in the embedding space, but they cannot cover all variances within a semantic class. We propose a heatmap neighbor contrastive learning (HNCL) method for FER. It treats the images corresponding to the heatmap nearest neighbors of expressions as other positives, providing more semantic variations than pre-defined augmented transformations. Therefore, our HNCL can learn better expression features covering more intra-class variances, improving the performance of the FER model based on self-supervised learning. After fine-tuning, HNCL with a simple framework achieves top-three performance on the in-the-lab datasets and even matches the performance of state-of-the-art supervised learning methods on the in-the-wild datasets.
Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Yibing Zhan, Dapeng Tao, Jun Wan 0005
IEEE Trans. Multim.6
2025 Dual-Task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding
abstract
Video Paragraph Grounding (VPG) aims to precisely locate the most appropriate moments within a video that are relevant to a given textual paragraph query. However, existing methods typically rely on large-scale annotated temporal labels and assume that the correspondence between videos and paragraphs is known. This is impractical in real-world applications, as constructing temporal labels requires significant labor costs, and the correspondence is often unknown. To address this issue, we propose a Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding method (DMR-JRG). In this method, retrieval and grounding tasks are mutually reinforced rather than being treated as separate issues. DMR-JRG mainly consists of two branches: a retrieval branch and a grounding branch. The retrieval branch uses inter-video contrastive learning to roughly align the global features of paragraphs and videos, reducing modality differences and constructing a coarse-grained feature space to break free from the need for correspondence between paragraphs and videos. Additionally, this coarse-grained feature space further facilitates the grounding branch in extracting fine-grained contextual representations. In the grounding branch, we achieve precise cross-modal matching and grounding by exploring the consistency between local, global, and temporal dimensions of video segments and textual paragraphs. By synergizing these dimensions, we construct a fine-grained feature space for video and textual features, greatly reducing the need for large-scale annotated temporal labels. Meanwhile, we design a grounding reinforcement retrieval module (GRRM) that brings the coarse-grained feature space of the retrieval branch closer to the fine-grained feature space of the grounding branch, thereby reinforcing retrieval branch through grounding branch, and finally achieving mutual reinforcement between tasks. Extensive experiments on three challenging datasets demonstrate the effectiveness of our proposed method. The code is available athttps://github.com/X7J92/DMR-JRG.
Mengzhao Wang 0002, Huafeng Li 0001, Jinxing Li 0003, Minghong Xie, Dapeng Tao
IEEE Trans. Multim.6
2025 SpliceMix: A Cross-Scale and Semantic Blending Augmentation Strategy for Multi-Label Image Classification
abstract
Recently, Mix-style data augmentation methods (e.g., Mixup and CutMix) have shown promising performance in various visual tasks. However, these methods are primarily designed for single-label images, ignoring the considerable discrepancies between single- and multi-label images,i.e., a multi-label image involves multiple co-occurred categories and fickle object scales. On the other hand, previous multi-label image classification (MLIC) methods tend to design elaborate models, bringing expensive computation. In this article, we introduce a simple but effective augmentation strategy for multi-label image classification, namely SpliceMix. The “splice” in our method is two-fold:1)Each mixed image is a splice of several downsampled images in the form of a grid, where the semantics of images attending to mixing are blended without object deficiencies for alleviating co-occurred bias;2)We splice mixed images and the original mini-batch to form a new SpliceMixed mini-batch, which allows an image with different scales to contribute to training together. Furthermore, such splice in our SpliceMixed mini-batch enables interactions between mixed images and original regular images. We also provide a simple and non-parametric extension based on consistency learning (SpliceMix-CL) to show the potential of extending our SpliceMix. Extensive experiments on various tasks demonstrate that only using SpliceMix with a baseline model (e.g., ResNet) achieves better performance than state-of-the-art methods. Moreover, the generalizability of our SpliceMix is further validated by the improvements in current MLIC methods when married with our SpliceMix.
Lei Wang 0095, Yibing Zhan, Leilei Ma 0002, Dapeng Tao, Liang Ding 0006, Chen Gong 0002
IEEE Trans. Multim.4
2025 Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction for Visual Grounding
abstract
Visual grounding has attracted wide attention thanks to its broad application in various visual language tasks. Although visual grounding has made significant research progress, existing methods ignore the promotion effect of the association between text and image features at different hierarchies on cross-modal matching. This paper proposes a Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction Visual Grounding method. It first generates a mask through decoupled sentence phrases, and a text and image hierarchical matching mechanism is constructed, highlighting the role of association between different hierarchies in cross-modal matching. In addition, a corresponding target object position progressive correction strategy is defined based on the hierarchical matching mechanism to achieve accurate positioning for the target object described in the text. This method can continuously optimize and adjust the bounding box position of the target object as the certainty of the text description of the target object improves. This design explores the association between features at different hierarchies and highlights the role of features related to the target object and its position in target positioning. The proposed method is validated on different datasets through experiments, and its superiority is verified by the performance comparison with the state-of-the-art methods.
Minghong Xie, Mengzhao Wang 0002, Huafeng Li 0001, Dapeng Tao, Zhengtao Yu 0001
IEEE Trans. Multim.5
2025 CLIP-Driven Semantic Discovery Network for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VIReID) primarily deals with matching identities across person images from different modalities. Due to the modality gap between visible and infrared images, cross-modality identity matching poses significant challenges. Recognizing that high-level semantics of pedestrian appearance, such as gender, shape, and clothing style, remain consistent across modalities, this paper intends to bridge the modality gap by infusing visual features with high-level semantics. Given the capability of Contrastive Language-Image Pre-training (CLIP) to sense high-level semantic information corresponding to visual representations, we explore the application of CLIP within the domain of VIReID. Consequently, we propose a CLIP-Driven Semantic Discovery Network (CSDN) that consists of Modality-specific Prompt Learner, Semantic Information Integration (SII), and High-level Semantic Embedding (HSE). Specifically, considering the diversity stemming from modality discrepancies in language descriptions, we devise bimodal learnable text tokens to capture modality-private semantic information for visible and infrared images, respectively. Additionally, acknowledging the complementary nature of semantic details across different modalities, we integrate text features from the bimodal language descriptions to achieve comprehensive semantics. Finally, we establish a connection between the integrated text features and the visual features across modalities. This process embed rich high-level semantic information into visual representations, thereby promoting the modality invariance of visual representations. The effectiveness and superiority of our proposed CSDN over existing methods have been substantiated through experimental evaluations on multiple widely used benchmarks.
Neng Dong, Liehuang Zhu, Hao Peng 0001, Dapeng Tao
IEEE Trans. Multim.5
2024 Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID
abstract
Text-to-image person re-identification (ReID) retrieves pedestrian images according to textual descriptions. Manually annotating textual descriptions is time-consuming, restricting the scale of existing datasets and therefore the generalization ability of ReID models. As a result, we study the transferable text-to-image ReID problem, where we train a model on our proposed large-scale database and directly deploy it to various datasets for evaluation. We obtain substantial training data via Multi-modal Large Language Models (MLLMs). Moreover, we identify and address two key challenges in utilizing the obtained textual descriptions. First, an MLLM tends to generate descriptions with similar structures, causing the model to overfit specific sentence patterns. Thus, we propose a novel method that uses MLLMs to caption images according to various templates. These templates are obtained using a multi-turn dialogue with a Large Language Model (LLM). Therefore, we can build a large-scale dataset with diverse textual descriptions. Second, an MLLM may produce incorrect descriptions. Hence, we introduce a novel method that automatically identifies words in a description that do not correspond with the image. This method is based on the similarity between one text and all patch token embeddings in the image. Then, we mask these words with a larger probability in the subsequent training epoch, alleviating the impact of noisy textual descriptions. The experimental results demonstrate that our methods significantly boost the direct transfer text-to-image ReID performance. Benefiting from the pre-trained model weights, we also achieve state-of-the-art performance in the traditional evaluation settings.https://github.com/WentaoTan/MLLM4Text-ReID
Wentao Tan, Changxing Ding, Jiayu Jiang, Fei Wang 0032, Yibing Zhan, Dapeng Tao
CVPR6
2024 Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld
abstract
While large language models (LLMs) excel in a simulated world of texts, they struggle to interact with the more realistic world without perceptions of other modalities such as visual or audio signals. Although vision-language models (VLMs) integrate LLM modules (1) aligned with static image features, and (2) may possess prior knowledge of world dynamics (as demonstrated in the text world), they have not been trained in an embodied visual world and thus cannot align with its dynamics. On the other hand, training an embodied agent in a noisy visual world without expert guidance is often chal-lenging and inefficient. In this paper, we train a VLM agent living in a visual world using an LLM agent excelling in a parallel text world. Specifically, we distill LLM's reflection outcomes (improved actions by analyzing mistakes) in a text world's tasks to finetune the VLM on the same tasks of the visual world, resulting in an Embodied Multi-Modal Agent (EMMA) quickly adapting to the visual world dy-namics. Such cross-modality imitation learning between the two parallel worlds is achieved by a novel DAgger-DPO algorithm, enabling EMMA to generalize to a broad scope of new tasks without any further guidance from the LLM expert. Extensive evaluations on the ALFWorld benchmark's diverse tasks highlight EMMA's superior performance to SOTA VLM-based agents, e.g., 20%-70% improvement in the success rate.
Tianyi Zhou 0001, Kanxue Li, Dapeng Tao, Lusong Li, Li Shen 0008, Xiaodong He 0001, Jing Jiang 0002, Yuhui Shi 0001
CVPR4
2024 Triple Temporal Vision Transformer for the Coverage Classification with Multi-Temporal Polsar Images
abstract
Multi-temporal SAR and Polarimetric SAR (PolSAR) images can provide the scattering change characteristics caused by vegetation growth to help the classifier capture phenological characteristics. To effectively utilize multi-temporal PolSAR data, a triple temporal vision transformer (TriTempoViT) model is proposed to capture correlation from multi-dimensional features. The method uses a three-branch network architecture to extract spatial-temporal, spatial-polarimetric, and temporal-polarization features respectively, and then the features from the three branches will be integrated into the vision transformer (ViT) for information interaction. Additionally, a 3D channel-spatial attention module (3DCSAM) is tailored to automatically weight the importance of the multi-dimensional feature maps. Moreover, a temporal interaction feature extraction module (TIFEM) is designed to comprehensively consider correlations between different temporal sequences. Compared with the recently developed state-of-the-art approach, the proposed method can improve OA of the classification accuracy in a Radarsat-2 dataset from Flevoland by about 0.84%, which proves the effectiveness of the proposed method.
Jiahui Peng, Dapeng Tao, Carlos López-Martínez
IGARSS2
2024 Where to Mask: Structure-Guided Masking for Graph Masked Autoencoders
Chuang Liu 0008, Yibing Zhan, Xueqi Ma, Dapeng Tao, Jia Wu 0001, Wenbin Hu 0001
IJCAI5
2024 MuEP: A Multimodal Benchmark for Embodied Planning with Foundation Models
Kanxue Li, Baosheng Yu, Yibing Zhan, Qiong Cao, Li Shen 0008, Lusong Li, Dapeng Tao, Xiaodong He 0001
IJCAI13
2024 An Efficient Multi-prior Hybrid Approach for Consistent 3D Generation from Single Images
Yichen Ouyang, Jiayi Ye, Wenhao Chai, Dapeng Tao, Yibing Zhan, Gaoang Wang
MMAsia4
2024 ASFusion: Adaptive visual enhancement and structural patch decomposition for infrared and visible image fusion
Yiqiao Zhou, Kangjian He, Dan Xu 0001, Dapeng Tao, Chengzhou Li
Eng. Appl. Artif. Intell.4
2024 HAG-Former: A Temporal-Polarimetric Relationship Inference Network From Local to Global
abstract
Multitemporal polarimetric SAR (PolSAR) data can provide a unique insight into the temporal scattering characteristics of targets and highlight their dynamic changes over time, therefore supporting improved classification performance. Constrained by the complexities of satellite orbit control technology and the challenges associated with time-series PolSAR data acquisition, most prevailing methodologies rely solely on a single PolSAR image to tackle land coverage classification, inherently limiting their ability to generalize across diverse scenarios. To address this limitation, this work introduces a novel Hybrid Attention-GRU Transformer (HAG-Former) model, which harnesses the power of pixel-level temporal-polarimetric change analysis and captures the dynamic variations in polarization scattering properties, thereby enhancing classification robustness and versatility. In this approach, we seamlessly integrate a self-attention mechanism, a Gated Recurrent Unit (GRU), and a transformer encoder to delve into pixel-level changes in polarimetric features. Initially, the self-attention mechanism pinpoints crucial classification-aiding features, bolstering their significance. The weighted features are then fed into the GRU model, enhancing local temporal-polarimetric relationship insights. These relationships, coupled with significant features from the self-attention mechanism, are subsequently processed by the transformer encoder, unraveling global information. Furthermore, we employ a label smoothing loss function during training, mitigating the impact of sample imbalance on classification accuracy. To validate the effectiveness of our proposed methodology, we evaluated it on two benchmark datasets. The results demonstrate a notable enhancement in classification performance, achieving an overall accuracy improvement of 2.21% and 1.79% over the state-of-the-art. The code is available athttps://github.com/Thomasakun/HAGFormer.
Carlos López-Martínez, Yibing Zhan, Dapeng Tao
IEEE Geosci. Remote. Sens. Lett.6
2024 Exploring sparsity in graph transformers
Chuang Liu 0008, Yibing Zhan, Xueqi Ma, Liang Ding 0006, Dapeng Tao, Jia Wu 0001, Wenbin Hu 0001, Bo Du 0001
Neural Networks5
2024 MCNet: Magnitude consistency network for domain adaptive object detection under inclement environments
Jian Pang, Weifeng Liu 0001, Bingfeng Zhang, Xinghao Yang, Baodi Liu, Dapeng Tao
Pattern Recognit.6
2024 Free-Form Composition Networks for Egocentric Action Recognition
abstract
Egocentric action recognition is gaining significant attention in the field of human action recognition. In this paper, we address data scarcity issue in egocentric action recognition from a compositional generalization perspective. To tackle this problem, we propose a free-form composition network (FFCN) that can simultaneously learn disentangled verb, preposition, and noun representations, and then use them to compose new samples in the feature space for rare classes of action videos. First, we use a graph to capture the spatial-temporal relations among different hand/object instances in each action video. We thus decompose each action into a set of verb and preposition spatial-temporal representations using the edge features in the graph. The temporal decomposition extracts verb and preposition representations from different video frames, while the spatial decomposition adaptively learns verb and preposition representations from action-related instances in each frame. With these spatial-temporal representations of verbs and prepositions, we can compose new samples for those rare classes in a free-form manner, which is not restricted to a rigid form of a verb and a noun. The proposed FFCN can directly generate new training data samples for rare classes, hence significantly improve action recognition performance. We evaluated our method on three popular egocentric action recognition datasets, Something-Something V2, H2O, and EPIC-KITCHENS-100, and the experimental results demonstrate the effectiveness of the proposed method for handling data scarcity problems, including long-tailed and few-shot egocentric action recognition.
Haoran Wang 0001, Qinghua Cheng, Baosheng Yu, Yibing Zhan, Dapeng Tao, Liang Ding 0006, Haibin Ling
IEEE Trans. Circuits Syst. Video Technol.5
2024 DeIoU: Toward Distinguishable Box Prediction in Densely Packed Object Detection
abstract
The Intersection over Union (IoU) has been widely employed in various stages of object detection owing to its ability to quantify the similarity between boxes objectively. However, in densely packed scenes full of crowded and small-sized objects, adjacent positive boxes often exhibit high levels of overlap. This overlap interference compromises the consistency between quality evaluation and confidence, leading to ambiguous box prediction within the previous IoU-based models. To address this issue, we design a novel learning paradigm tailored for Dense scenes based on IoU, called DeIoU. This approach effectively suppresses unnecessary overlap between predicted boxes and thereby enhances representation learning for non-salient objects. Specifically, it consists of a dense box regression loss${\mathcal {L}}_{DeIoU}$and a one-to-many (O2M) label matching strategy guided by DeIoU. These components focus on calibrating the position and shape prediction quality during the model training, learning distinguishable object features by penalizing overlap interference between neighboring boxes. Extensive experiments on four object detection datasets including SKU-110K, CrowdHuman, MS COCO 2017, and DIOR, demonstrate that our DeIoU-based learning strategy outperforms other state-of-the-art methods. Notably, the proposed method delivers a substantial improvement (average$1.3~{AP}$and$1.8~MR^{-2}$) across popular detectors on SKU-110K and CrowdHuman while exhibiting distinct competitiveness on small objects within natural scenes.
Linfei Wang, Yibing Zhan, Long Lan, Dapeng Tao, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 EME: Energy-Based Multiexpert Model for Long-Tailed Remote Sensing Image Classification
abstract
The distribution of remote sensing scene images often follows a long-tailed pattern, where there is an abundance of samples in a few dominant classes and a scarcity of samples in most other classes. This presents two major challenges when it comes to identifying this type of data: Head-Dominance: Models trained on such data tend to prioritize the dominant classes, overlooking the tail classes and resulting in poor performance when it comes to recognizing them. Tail-Interference: The presence of tail classes disrupts the learned representations for the head classes, acting as noise that negatively impacts the recognition accuracy of the head data. To address these challenges, we propose an innovative solution called the energy-based multiexpert (EME) model. The core concept behind this approach is to utilize energy-based discriminators (EDors) to separate the data into head and tail categories. Subsequently, we design multiple experts to classify the head and tail data separately, ensuring that the significant differences in data volume between these categories do not interfere with each other. Experimental results obtained by applying the EME model to three remote sensing datasets demonstrate its efficiency, outperforming current state-of-the-art (SOTA) methods. These findings underscore the effectiveness of our proposed approach in addressing the challenges posed by the long-tailed distribution in remote sensing scene images.
Shuai Shao 0006, Shiyuan Zhao, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Trans. Geosci. Remote. Sens.5
2024 Deep Location Soft-Embedding-Based Network With Regional Scoring for Mammogram Classification
abstract
Early detection and treatment of breast cancer can significantly reduce patient mortality, and mammogram is an effective method for early screening. Computer-aided diagnosis (CAD) of mammography based on deep learning can assist radiologists in making more objective and accurate judgments. However, existing methods often depend on datasets with manual segmentation annotations. In addition, due to the large image sizes and small lesion proportions, many methods that do not use region of interest (ROI) mostly rely on multi-scale and multi-feature fusion models. These shortcomings increase the labor, money, and computational overhead of applying the model. Therefore, a deep location soft-embedding-based network with regional scoring (DLSEN-RS) is proposed. DLSEN-RS is an end-to-end mammography image classification method containing only one feature extractor and relies on positional embedding (PE) and aggregation pooling (AP) modules to locate lesion areas without bounding boxes, transfer learning, or multi-stage training. In particular, the introduced PE and AP modules exhibit versatility across various CNN models and improve the model's tumor localization and diagnostic accuracy for mammography images. Experiments are conducted on published INbreast and CBIS-DDSM datasets, and compared to previous state-of-the-art mammographic image classification methods, DLSEN-RS performed satisfactorily.
Bowen Han 0001, Luhao Sun, Chao Li 0075, Wenzong Jiang, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Trans. Medical Imaging7
2024 Bounding Box Vectorization for Oriented Object Detection With Tanimoto Coefficient Regression
abstract
Current oriented object detection methods mainly utilize a vanilla coordinate-angle representation for bounding box regression, which usually suffers from inconsistency between the bounding box regression losses and prediction errors induced with respect to different rotation angles, aspect ratios, and scales. Therefore, although the existing oriented object detectors have achieved very good performances under coarse evaluation metrics such as AP50, their performance significantly degrades when using stricter evaluation metric such as AP75. To address the abovementioned issues, we propose a new regression method with bounding box vectorization that implicitly represents the shape and orientation of an object with a set of orthogonal vectors. By doing this, the proposed method delicately avoids the inconsistency issues encountered in oriented bounding box regression. During training, we introduce the Tanimoto coefficient to evaluate the similarity of the bounding box vector in a shape- and orientation-aware manner, and we refer to the proposed box-to-vector loss as the B2V loss. In addition to 2D object detection, the proposed method can be easily generalized to 3D scenarios involving orientation estimation, such as autonomous driving. We evaluate the proposed method through extensive experiments conducted on four popular oriented object detection datasets, including both 2D and 3D datasets, where the proposed method significantly outperforms the recently developed state-of-the-art methods when using a more accurate evaluation metric.
Linfei Wang, Yibing Zhan, Wei Liu 0005, Baosheng Yu, Dapeng Tao
IEEE Trans. Multim.5
2024 Logical Relation Inference and Multiview Information Interaction for Domain Adaptation Person Re-Identification
abstract
Domain adaptation person re-identification (Re-ID) is a challenging task, which aims to transfer the knowledge learned from the labeled source domain to the unlabeled target domain. Recently, some clustering-based domain adaptation Re-ID methods have achieved great success. However, these methods ignore the inferior influence on pseudo-label prediction due to the different camera styles. The reliability of the pseudo-label plays a key role in domain adaptation Re-ID, while the different camera styles bring great challenges for pseudo-label prediction. To this end, a novel method is proposed, which bridges the gap of different cameras and extracts more discriminative features from an image. Specifically, an intra-to-intermechanism is introduced, in which samples from their own cameras are first grouped and then aligned at the class level across different cameras followed by our logical relation inference (LRI). Thanks to these strategies, the logical relationship between simple classes and hard classes is justified, preventing sample loss caused by discarding the hard samples. Furthermore, we also present a multiview information interaction (MvII) module that takes features of different images from the same pedestrian as patch tokens, obtaining the global consistency of a pedestrian that contributes to the discriminative feature extraction. Unlike the existing clustering-based methods, our method employs a two-stage framework that generates reliable pseudo-labels from the views of the intracamera and intercamera, respectively, to differentiate the camera styles, subsequently increasing its robustness. Extensive experiments on several benchmark datasets show that the proposed method outperforms a wide range of state-of-the-art methods. The source code has been released at https://github.com/lhf12278/LRIMV.
Fan Li 0006, Jinxing Li 0003, Huafeng Li 0001, Bob Zhang 0001, Dapeng Tao, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Comprehensive Graph Gradual Pruning for Sparse Training in Graph Neural Networks
abstract
Graph neural networks (GNNs) tend to suffer from high computation costs due to the exponentially increasing scale of graph data and a large number of model parameters, which restricts their utility in practical applications. To this end, some recent works focus on sparsifying GNNs (including graph structures and model parameters) with the lottery ticket hypothesis (LTH) to reduce inference costs while maintaining performance levels. However, the LTH-based methods suffer from two major drawbacks: 1) they require exhaustive and iterative training of dense models, resulting in an extremely large training computation cost, and 2) they only trim graph structures and model parameters but ignore the node feature dimension, where vast redundancy exists. To overcome the above limitations, we propose a comprehensive graph gradual pruning framework termed CGP. This is achieved by designing a during-training graph pruning paradigm to dynamically prune GNNs within one training process. Unlike LTH-based methods, the proposed CGP approach requires no retraining, which significantly reduces the computation costs. Furthermore, we design a cosparsifying strategy to comprehensively trim all the three core elements of GNNs: graph structures, node features, and model parameters. Next, to refine the pruning operation, we introduce a regrowth process into our CGP framework, to reestablish the pruned but important connections. The proposed CGP is evaluated over a node classification task across six GNN architectures, including shallow models [graph convolutional network (GCN) and graph attention network (GAT)], shallow-but-deep-propagation models [simple graph convolution (SGC) and approximate personalized propagation of neural predictions (APPNP)], and deep models [GCN via initial residual and identity mapping (GCNII) and residual GCN (ResGCN)], on a total of 14 real-world graph datasets, including large-scale graph datasets from the challenging Open Graph Benchmark (OGB). Experiments reveal that the proposed strategy greatly improves both training and inference efficiency while matching or even exceeding the accuracy of the existing methods.
Chuang Liu 0008, Xueqi Ma, Yibing Zhan, Liang Ding 0006, Dapeng Tao, Bo Du 0001, Wenbin Hu 0001, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.5
2023 Gapformer: Graph Transformer with Graph Pooling for Node Classification
abstract
Graph Transformers (GTs) have proved their advantage in graph-level tasks. However, existing GTs still perform unsatisfactorily on the node classification task due to 1) the overwhelming unrelated information obtained from a vast number of irrelevant distant nodes and 2) the quadratic complexity regarding the number of nodes via the fully connected attention mechanism. In this paper, we present Gapformer, a method for node classification that deeply incorporates Graph Transformer with Graph Pooling. More specifically, Gapformer coarsens the large-scale nodes of a graph into a smaller number of pooling nodes via local or global graph pooling methods, and then computes the attention solely with the pooling nodes rather than all other nodes. In such a manner, the negative influence of the overwhelming unrelated nodes is mitigated while maintaining the long-range information, and the quadratic complexity is reduced to linear complexity with respect to the fixed number of pooling nodes. Extensive experiments on 13 node classification datasets, including homophilic and heterophilic graph datasets, demonstrate the competitive performance of Gapformer over existing Graph Neural Networks and GTs.
Chuang Liu 0008, Yibing Zhan, Xueqi Ma, Liang Ding 0006, Dapeng Tao, Jia Wu 0001, Wenbin Hu 0001
IJCAI5
2023 A Stable Vision Transformer for Out-of-Distribution Generalization
Haoran Yu 0005, Baodi Liu, Yingjie Wang 0007, Kai Zhang 0029, Dapeng Tao, Weifeng Liu 0001
PRCV (8)5
2023 Holistic ARDS Prognosis Evaluation Framework Utilizing Data Governance and Ensemble Feature Selection
abstract
Acute respiratory distress syndrome (ARDS) prognosis has become integral to modern critical care models aimed at determining expected patient outcomes, optimizing clinical pathways, and improving resource allocation. However, existing clinical studies of ARDS prognosis face severe challenges due to the redundancy of multiple sources of heterogeneous data in current healthcare datasets, the high dimensionality of patient features, and the wide variation in the importance of each feature. In this paper, we propose a holistic ARDS prognostic framework for assessing ARDS prognosis through a standardized data governance process and an integrated feature selection approach based on embedded algorithms. Specifically, we extract ARDS patient data from medical data sources by various criteria and normalize the features. Then we apply multiple embedded feature selection methods to obtain a decision matrix based on the data-governed dataset. We conducted extensive experiments to demonstrate the efficiency and superiority of our proposed framework. The experimental results show that our proposed framework performs well in both data governance and feature selection and has wide clinical application potential.
Xiaodong Han, Guanghui Xiu, Dapeng Tao
SMC4
2023 Deep Positional-Representation-Based Local Information Retention Networks for Mammography Classification
abstract
Early diagnosis of breast cancer is challenging because in the most common mammogram images, the tumor usually occupies only a very small part of the entire image, which often makes deep learning models lose attention to the tumor area. In previous work, most models solved this problem by using ROI labeling to train models, which was expensive and difficult to widely apply. Some recent ROI-free methods use multi-scale features or multi-stage training, which gets rid of the model's dependence on ROI but greatly increases the computational complexity and deployment difficulty, limiting the potential of deep neural networks. Therefore, a deep positional-representation-based local information retention networks (PR-LIR) was proposed. PR-LIR is a lightweight, end-to-end mammogram classification model, which uses positional representation (PR) and multi-scale regional pooling (MRP) modules to locate tumor regions and retain regional semantic information of small target tumors at different scales, without ROI labeling and multi-stage training, and almost no increasement in parameters and computational complexity. In particular, the proposed PR and MRP modules have good generalization performance, which can be applied to most CNN models and improve the classification accuracy of mammography images. Experimental results on two publicly available datasets show that PR-LIR achieves the best AUC and satisfactory accuracy compared to the previous state-of-the-art mammogram classification method.
Bowen Han 0001, Luhao Sun, Chao Li 0075, Wenzong Jiang, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
SMC7
2023 Low-light image enhancement for infrared and visible image fusion
abstract
Abstract Infrared and visible image fusion (IVIF) is an essential branch of image fusion, and enhancing the visible image of IVIF can significantly improve the fusion performance. However, many existing low‐light enhancement methods are unsuitable for the visible image enhancement of IVIF. In order to solve this problem, this paper proposes a new visible image enhancement method for IVIF. Firstly, the colour balance and contrast enhancement‐based self‐calibrated illumination estimation (CCSCE) is proposed to improve the input image's brightness, contrast, and colour information. Then, the method based on Mutually Guided Image Filtering (muGIF) is adopted to design a strategy to extract details adaptively from the original visible image, which can keep details without introducing additional noise effectively. Finally, the proposed visible image enhancement technique is used for IVIF tasks. In addition, the proposed method can be used for the visible image enhancement of IVIF and other low‐light images. Experiment results on different public datasets and IVIF demonstrate the authors’ method's superiority from both qualitative and quantitative comparisons. The authors’ code will be publicly available at https://github.com/yiqiao666/low‐light‐enhancement‐for‐IVIF/tree/master .
Yiqiao Zhou, Lisiqi Xie, Kangjian He, Dan Xu 0001, Dapeng Tao
IET Image Process.5
2023 Deformable Convolutional Network Constrained by Contrastive Learning for Underwater Image Enhancement
abstract
Autonomous underwater vehicles (AUVs) based on remote sensing technology have been widely applied in various underwater tasks. However, the complex underwater environment leads to challenges such as color distortion, blurred details, and fog effects in the underwater image directly acquired by AUVs. Although numerous existing methods aim to remove the color cast and restore image details, their effectiveness is still limited. This paper proposes a new method based on a deformable convolutional network and constrained by contrastive learning for underwater image enhancement. First, we propose a deformable convolutional residual block (DCRB) to achieve a more precise restoration of texture details by adaptively adjusting the convolution kernel shape. At the same time, we utilize the long-skip connection method of the U-Net architecture to preserve information that is prone to lose in shallow networks. Second, we propose a color contrastive loss function to compare the color difference between distorted images and the ground truth, resulting in a more realistic enhanced image. Finally, experimental results demonstrate that the proposed method outperforms the state-of-the-art methods regarding image quality and visual appeal.
Xinran Guo, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.4
2023 Superpixel-based adaptive salient region analysis for infrared and visible image fusion
Chengzhou Li, Kangjian He, Dan Xu 0001, Dapeng Tao, Hongzhen Shi, Wenxia Yin
Neural Comput. Appl.4
2023 Transferring fashion to surveillance with weak labels
Zheng He 0001, Chao Liang 0001, Jun Chen 0001, Chia-Wen Lin, Dapeng Tao
Neural Comput. Appl.6
2023 GSA4FDA: Deep Geometric and Statistic Alignment for Fewer Labeled Domain Adaptation
Yuying Cai, Baodi Liu, Xinghao Yang, Dapeng Tao, Weifeng Liu 0001
Neural Process. Lett.5
2023 Cross-Domain Few-Shot classification via class-shared and class-specific dictionaries
Lei Xing 0005, Baodi Liu, Dapeng Tao, Weijia Cao, Weifeng Liu 0001
Pattern Recognit.4
2023 Self-Paced Hard Task-Example Mining for Few-Shot Classification
abstract
In recent years, researchers have commonly employed assistant tasks to enhance the training phase of the few-shot classification models. Several methods have been proposed to exploit and optimize the training tasks, such as Curriculum Learning (CL) and Hard Example Mining (HEM). However, most of the existing strategies can not elaborately leverage the training tasks and share some common drawbacks, including 1) the ignorance of the target tasks’ properties, and 2) the neglect of sample relationships. In this work, we propose a Self-Paced Hard tAsk-Example Mining (SP-HAEM) method to solve these problems. Specifically, the SP-HAEM automatically chooses hard examples via the similarity between training and target tasks to optimize the support set. To represent the property of target tasks, SP-HAEM obtains a representation of the dataset, called “meta-task”. No need to apply an additional model to measure difficulty and choose hard examples like other HEM methods, SP-HAEM selects the tasks with large optimal transport distance to the meta-task as hard tasks. Thus, training with such hard tasks can not only enhances the generalization ability of the model but also eliminate the negative effect of redundancy tasks. To evaluate the effectiveness of SP-HAEM, we conduct extensive experiments on a variety of datasets, including MiniImageNet, TieredImageNet, and FC100. The results of the experiments show that SP-HAEM can achieve higher accuracy compared with the typical few-shot classification models, e.g., Prototypical Network, MAML, FEAT, and MTL.
Xinghao Yang, Xingxing Yao, Dapeng Tao, Weijia Cao, Weifeng Liu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Dynamic Adaptive Attention-Guided Self-Supervised Single Remote-Sensing Image Denoising
abstract
Optical remote sensing images are widely used in many fields, and local complex texture details in images usually play a critical role in downstream tasks. However, noise interference will destroy the complex texture in the image, thus reducing the accuracy of downstream tasks. The current attention mechanism usually focuses on the global high-level features in the image, so it cannot effectively focus on the high-frequency information in the local complex texture in the remote sensing image, and obtaining clean remote sensing images to train neural networks is difficult. Therefore, applying the current depth learning based natural image denoising methods directly to optical remote sensing images is challenging. To solve these problems, we propose a dynamic adaptive attention guided self-supervised single remote sensing image denoising network (DAA-SSID). We construct a dynamic adaptive attention module (DAAM) by dynamically calculating the activation intensity of each neuron and combining the spatial feature information extracted from remote sensing images. It can effectively extract complex texture features from remote sensing images when only a single remote sensing image participates in training. And we use independent random Bernoulli sampling in the training and inference stages respectively to prevent over-fitting caused by single-image training. Therefore, compared with other self-supervised denoising methods, our proposed model can denoise remote sensing images with more complex textures when only a single image destroyed by noise is used as the training input. Experiments on synthetic additive gaussian noise data and authentic noise data have shown that the proposed model achieves satisfactory results.
Minghao Liu 0016, Wenzong Jiang, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Trans. Geosci. Remote. Sens.4
2023 RAN: Region-Aware Network for Remote Sensing Image Super-Resolution
abstract
The remote sensing (RS) image super-resolution (SR) algorithm aims to reconstruct a high-resolution (HR) image with rich texture details from a given low-resolution (LR) image, improving the spatial resolution. It has been widely concerned in remote sensing image processing and application. Most current deep learning-based methods rely on paired training datasets. However, most datasets are often based on bicubic degradation. This single construction way limits the performance of the pre-trained network. Moreover, SR is an ill-posed problem in that multiple SR images are constructed from a single LR input. This paper proposes a Region-Aware Network (RAN) for remote sensing image super-resolution to alleviate the above issues. First, we introduce the contrastive learning strategy to mine the latent degraded representation of the image and serve as the prior knowledge of the network. Considering the RS images are acquired in specific scenes that have apparent self-similarity. Then, we propose a Region-Aware Module (RAM) based on attention mechanisms and the graph neural network to explore region information and cross-patch self-similarity. Extensive experiments have demonstrated that the proposed RAN adapts to RS image super-resolution tasks with various degradations and performs better in constructing texture information.
Baodi Liu, Lifei Zhao, Shuai Shao 0006, Weifeng Liu 0001, Dapeng Tao, Weijia Cao, Yicong Zhou
IEEE Trans. Geosci. Remote. Sens.5
2022 Masked Graph Auto-Encoder Constrained Graph Pooling
Chuang Liu 0008, Yibing Zhan, Xueqi Ma, Dapeng Tao, Bo Du 0001, Wenbin Hu 0001
ECML/PKDD (2)4
2022 Multi-view learning for hyperspectral image classification: An overview
Baodi Liu, Kai Zhang 0029, Honglong Chen, Weijia Cao, Weifeng Liu 0001, Dapeng Tao
Neurocomputing7
2022 Key point-aware occlusion suppression and semantic alignment for occluded person re-identification
Shujuan Wang, Bochun Huang, Huafeng Li 0001, Guanqiu Qi, Dapeng Tao, Zhengtao Yu 0001
Inf. Sci.5
2022 Multiorder Interaction Information Embedding-Based Multiview Fusion-Aided Hyperspectral Image Classification
abstract
Hyperspectral images (HSI) are obtained from hyperspectral imaging sensors, which capture information in hundreds of spectral bands of objects. However, how to take full advantage of spatial and spectral information from many spectral bands to improve the performance of HSI classification remains an open question. Many HSI classification works have recently been reported by employing multi-view learning (MVL) algorithms that can fully use complementary information between different view features and thus have received widespread attention. This paper proposes a multi-view fusion network based on multi-order interaction information embedding for HSI classification. Firstly, the correlation matrix between spectral bands is used to divide the original data into multiple subsets as local views. The subset after the Segmented-PCA process is used as the global view. Secondly, the features of different views are extracted separately using a feature extraction network and mapped to the same dimension. Pre-fusion is achieved by multi-order interaction of various view features. Finally, loss-weighted fusion is applied to each view according to its contribution to the classification task. To evaluate the effectiveness of the proposed method, complete experiments were conducted on three commonly used HSI datasets, namely Pavia University, Houston 2013, and Houston 2018. The experimental results demonstrate that the proposed method improves the classification performance of existing feature extraction networks and is more competitive with other methods in the field.
Weijia Cao, Kai Zhang 0029, Baodi Liu, Dapeng Tao, Weifeng Liu 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Covered Style Mining via Generative Adversarial Networks for Face Anti-spoofing
Yiqiang Wu, Dapeng Tao, Yong Luo 0002, Jun Cheng 0002, Xuelong Li 0001
Pattern Recognit.2
2022 RiFeGAN2: Rich Feature Generation for Text-to-Image Synthesis From Constrained Prior Knowledge
abstract
Text-to-image synthesis is a challenging task that generates realistic images from a textual description. The description contains limited information compared with the corresponding image and is ambiguous and abstract, which will complicate the generation and lead to low-quality images. To address this problem, we propose a novel generation text-to-image synthesis method, called RiFeGAN2, to enrich the given description. To improve the enrichment quality while accelerating the enrichment process, RiFeGAN2 exploits a domain-specific constrained model to limit the search scope and then uses an attention-based caption matching model to refine the compatible candidate captions based on constrained prior knowledge. To improve the semantic consistency between the given description and the synthesized results, RiFeGAN2 employs improved SAEMs, SAEM2s, to compact better features of the retrieved captions and effectively emphasize the descriptions via incorporating centre-attention layers. Finally, multi-caption attentional GANs are exploited to synthesize images from those features. Experiments performed on widely-used datasets show that the models can generate vivid images from enriched captions and effectually improve the semantic consistency.
Jun Cheng 0002, Fuxiang Wu, Yanling Tian, Lei Wang 0018, Dapeng Tao
IEEE Trans. Circuits Syst. Video Technol.5
2022 Triple Adversarial Learning and Multi-View Imaginative Reasoning for Unsupervised Domain Adaptation Person Re-Identification
abstract
Due to the importance of practical applications, unsupervised domain adaptation (UDA) person re-identification (re-ID) has attracted increasing attention. However, most of existing methods often lack the multi-view information reasoning and ignore the domain discrepancy of the pedestrian images with the same identity, which constrain the further improvement of recognition performance. So, this paper proposes a triple adversarial learning and multi-view imaginative reasoning network (TAL-MIRN) for UDA person re-ID, which consists of a multi-view imaginative reasoning module (IRM) and a triple adversarial learning module (TALM). IRM makes the classified pedestrian identity features from a single-view image extracted by a feature encoder consistent with the classification results of the aggregated multi-view pedestrian identity features, so the strong multi-view imaginative reasoning ability of the feature encoder is obtained. TALM is composed by the adversarial learning between the camera classifier and feature encoder, adversarial learning of joint distribution alignment, and adversarial learning of the difference between two classifiers used in classification. In particular, the domain-invariant features at camera level are guaranteed by the adversarial learning between the feature extractor and camera classifier. The joint alignment of identity and domain is achieved by the competition between the feature extractor and classifier integrated with identity and domain. The discriminability and robustness of the learned features are enhanced by playing a MinMax game between two different identity classifiers. Furthermore, a simple normalization operation named as cross normalization (CN) is proposed to increase both modeling and generalization capability of the proposed TAL-MIRN across multiple domains. The proposed TAL-MIRN is applied to five benchmark datasets, and the comparative experimental results confirm its superiority over the state-of-the-art methods. The related source codes is available athttps://github.com/lhf12278/TALM-IRM.
Huafeng Li 0001, Neng Dong, Zhengtao Yu 0001, Dapeng Tao, Guanqiu Qi
IEEE Trans. Circuits Syst. Video Technol.4
2022 Adversarial UV-Transformation Texture Estimation for 3D Face Aging
abstract
Face aging aims to estimate aged facial textures given a certain face image. A number of 2D face-aging methods have been developed, but there have been few studies on 3D face aging, which would be valuable in several real-world applications. The lack of 3D face-aging data has had a significant impact on the development of 3D face aging, but we hypothesized that the large amounts of 2D face-aging data on the internet could be leveraged for 3D aged facial textures. In this paper, we propose a novel 3D aging framework, which we call UV-transformation texture estimation based on generative adversarial networks (UVTE-GAN), to achieve 3D face aging. Specifically, the proposed framework has three parts: 1) a 3D vertex and texture estimator, which accurately estimates the face’s spatial vertices and textures; 2) a texture-aging GAN, which is responsible for aging the estimated texture map via adversarial learning; and 3) a 2D & 3D rendering rebuilder, which recovers 2D & 3D faces using the estimated facial vertex map and aged facial texture map. In addition, we also design a plugin layer that allows us to train the whole model in an end-to-end manner. Experimental results demonstrate the effectiveness of the proposed method in synthesizing visually pleasing 3D aged face pictures, and state-of-the-art performance is achieved on several public datasets.
Yiqiang Wu, Ruxin Wang 0002, Mingming Gong, Jun Cheng 0002, Zhengtao Yu 0001, Dapeng Tao
IEEE Trans. Circuits Syst. Video Technol.6
2022 Image Hallucination From Attribute Pairs
abstract
Recent image-generation methods have demonstrated that realistic images can be produced from captions. Despite the promising results achieved, existing caption-based generation methods confront a dilemma. On the one hand, the image generator should be provided with sufficient details for realistic hallucination, meaning that longer sentences with rich content are preferred, but on the other hand, the generator is meanwhile fragile to long sentences due to their complex semantics and syntax like long-range dependencies and the combinatorial explosion of object visual features. Toward alleviating this dilemma, a novel approach is proposed in this article to hallucinate images from attribute pairs, which can be extracted from natural language processing (NLP) toolsets in the presence of complex semantics and syntax. Attribute pairs, therefore, enable our image generator to tackle long sentences handily and alleviate the combinatorial explosion, and at the same time, allow us to enlarge the training dataset and to produce hallucinations from randomly combined attribute pairs at ease. Experiments on widely used datasets demonstrate that the proposed approach yields results superior to the state of the art.
Fuxiang Wu, Jun Cheng 0002, Xinchao Wang, Lei Wang 0018, Dapeng Tao
IEEE Trans. Cybern.5
2022 BiN-Flow: Bidirectional Normalizing Flow for Robust Image Dehazing
abstract
Image dehazing aims to remove haze in images to improve their image quality. However, most image dehazing methods heavily depend on strict prior knowledge and paired training strategy, which would hinder generalization and performance when dealing with unseen scenes. In this paper, to address the above problem, we propose Bidirectional Normalizing Flow (BiN-Flow), which exploits no prior knowledge and constructs a neural network through weakly-paired training with better generalization for image dehazing. Specifically, BiN-Flow designs 1) Feature Frequency Decoupling (FFD) for mining the various texture details through multi-scale residual blocks and 2) Bidirectional Propagation Flow (BPF) for exploiting the one-to-many relationships between hazy and haze-free images using a sequence of invertible Flow. In addition, BiN-Flow constructs a reference mechanism (RM) that uses a small number of paired hazy and haze-free images and a large number of haze-free reference images for weakly-paired training. Essentially, the mutual relationships between hazy and haze-free images could be effectively learned to further improve the generalization and performance for image dehazing. We conduct extensive experiments on five commonly-used datasets to validate the BiN-Flow. The experimental results that BiN-Flow outperforms all state-of-the-art competitors demonstrate the capability and generalization of our BiN-Flow. Besides, our BiN-Flow could produce diverse dehazing images for the same image by considering restoration diversity.
Yiqiang Wu, Dapeng Tao, Yibing Zhan, Chenyang Zhang 0003
IEEE Trans. Image Process.2
2022 Domain-invariant Graph for Adaptive Semi-supervised Domain Adaptation
abstract
Domain adaptation aims to generalize a model from a source domain to tackle tasks in a related but different target domain. Traditional domain adaptation algorithms assume that enough labeled data, which are treated as the prior knowledge are available in the source domain. However, these algorithms will be infeasible when only a few labeled data exist in the source domain, thus the performance decreases significantly. To address this challenge, we propose a Domain-invariant Graph Learning (DGL) approach for domain adaptation with only a few labeled source samples. Firstly, DGL introduces the Nyström method to construct a plastic graph that shares similar geometric property with the target domain. Then, DGL flexibly employs the Nyström approximation error to measure the divergence between the plastic graph and source graph to formalize the distribution mismatch from the geometric perspective. Through minimizing the approximation error, DGL learns a domain-invariant geometric graph to bridge the source and target domains. Finally, we integrate the learned domain-invariant graph with the semi-supervised learning and further propose an adaptive semi-supervised model to handle the cross-domain problems. The results of extensive experiments on popular datasets verify the superiority of DGL, especially when only a few labeled source samples are available.
Weifeng Liu 0001, Yicong Zhou, Jun Yu 0002, Dapeng Tao, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.5
2021 Adaptive colour restoration and detail retention for image enhancement
abstract
Abstract Computer vision‐based crowd understanding and analysis technology has been widely used in public safety due to the rapid growth of population and the frequent occurrence of various accidents. Improving imaging quality is the key to improve the performance of crowd analysis, density estimation, target recognition, segmentation, and detection in computer vision tasks. Due to the complex imaging environment such as fog and low illumination, some images taken in outdoor environment often have the problems of colour distortion, lack of details, and the poor imaging quality, which affect the subsequent visual tasks. To improve the imaging quality and visual effect, an adaptive colour restoration and detail retention‐based method is proposed for image enhancement. First, to overcome the problem of colour distortion caused by low illumination and fog, a multi‐channel fusion based adaptive image colour restoration method is proposed. To make the enhancement result more consistent with human observation, the detail retention‐based method is applied to enhance the details. Experimental results demonstrate that the authors' results are effective and outperform the compared methods both in visual and objective evaluations.
Kangjian He, Dapeng Tao, Dan Xu 0001
IET Image Process.2
2021 Equidistant distribution loss for person re-identification
Zhao Yang 0001, Jiehao Liu, Yuanxin Zhu, Li Wang 0067, Dapeng Tao
Neurocomputing6
2021 Semi-supervised classification by graph p-Laplacian convolutional networks
Sichao Fu, Weifeng Liu 0001, Kai Zhang 0029, Yicong Zhou, Dapeng Tao
Inf. Sci.5
2021 Cross adversarial consistency self-prediction learning for unsupervised domain adaptation person re-identification
Huafeng Li 0001, Jian Pang, Dapeng Tao, Zhengtao Yu 0001
Inf. Sci.3
2021 Deep features for person re-identification on metric learning
Wanyin Wu, Dapeng Tao, Zhao Yang 0001, Jun Cheng 0002
Pattern Recognit.2
2021 Attribute-Aligned Domain-Invariant Feature Learning for Unsupervised Domain Adaptation Person Re-Identification
abstract
Domain invariance and discrimination of learned features as two crucial factors affect the performance of unsupervised domain adaptation (UDA) person re-identification (Re-ID). Person attributes (such as “backpack”, “boots”, “handbag”, etc) remaining unchanged across multiple domains have been used as mid-level visual-semantic information in UDA person Re-ID. As two main challenges, both misalignment of attribute-related regions across multiple images and domain shift between source and target domains affect the learning of domain-invariant features (DIF). To address the above two challenges, this article proposes to take advantage of the stability of person attributes and the complementarity of person attributes and the corresponding low-level visual features to guide the learning of discriminative DIF. Specifically, the proposed solution contains the generation of latent attribute-correlated visual features (GLAVF), DIF learning under the guidance of person attributes, and the alignment of person attributes corresponding to the local regions of pedestrian images. Due to the gap between person attributes and visual features, person attributes are first converted into latent attribute-correlated visual features (LAVF) without any specific domain information in GLAVF, and then LAVF are used as the substitutions of person attributes to guide the learning of DIF. To enhance the discrimination of learned features, the proposed solution mainly explores the alignment between person attributes and corresponding local regions, and the alignment of the same person attributes across multiple pedestrian images. A fully connected layer is used to achieve the above two types of alignment in the proposed framework, which reduces the adverse impacts of inference information and ensures the semantic consistency between person attributes and corresponding local regions across multiple pedestrian images. The effectiveness of the proposed solution is confirmed on four existing datasets by comparative experiments.
Huafeng Li 0001, Dapeng Tao, Zhengtao Yu 0001, Guanqiu Qi
IEEE Trans. Inf. Forensics Secur.3
2021 A Contour Co-Tracking Method for Image Pairs
abstract
We proposed a contour co-tracking method for co-segmentation of image pairs based on active contour model. Our method comprehensively re-models objects and backgrounds signified by level set functions, and leverages Hellinger distance to measure the similarity between image regions encoded by probability distributions. The main contribution are as follows. 1) The new energy functional, combining a rewarding and a penalty term, relaxes the assumptions of co-segmentation methods. 2) Hellinger distance, fulfilling the triangle inequality, ensures a coherence measurement between probability distributions in metric space, and contributes to finding a unique solution to the energy functional. The proposed contour co-tracking method was carefully verified against five representative methods on four popular datasets, i.e., the images pair dataset (105 pairs), MSRC dataset (30 pairs), iCoseg dataset (66 pairs) and Coseg-rep dataset (25 pairs). The comparison experiments suggest that our method achieves the competitive and even better performance compared to the state-of-the-art co-segmentation methods.
Bin Wang 0027, Dapeng Tao, Yuan Yan Tang, Xinbo Gao 0001
IEEE Trans. Image Process.2
2021 Dynamic Graph Learning Convolutional Networks for Semi-supervised Classification
abstract
Over the past few years, graph representation learning (GRL) has received widespread attention on the feature representations of the non-Euclidean data. As a typical model of GRL, graph convolutional networks (GCN) fuse the graph Laplacian-based static sample structural information. GCN thus generalizes convolutional neural networks to acquire the sample representations with the variously high-order structures. However, most of existing GCN-based variants depend on the static data structural relationships. It will result in the extracted data features lacking of representativeness during the convolution process. To solve this problem, dynamic graph learning convolutional networks (DGLCN) on the application of semi-supervised classification are proposed. First, we introduce a definition of dynamic spectral graph convolution operation. It constantly optimizes the high-order structural relationships between data points according to the loss values of the loss function, and then fits the local geometry information of data exactly. After optimizing our proposed definition with the one-order Chebyshev polynomial, we can obtain a single-layer convolution rule of DGLCN. Due to the fusion of the optimized structural information in the learning process, multi-layer DGLCN can extract richer sample features to improve classification performance. Substantial experiments are conducted on citation network datasets to prove the effectiveness of DGLCN. Experiment results demonstrate that the proposed DGLCN obtains a superior classification performance compared to several existing semi-supervised classification models.
Sichao Fu, Weifeng Liu 0001, Weili Guan, Yicong Zhou, Dapeng Tao, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.5
2020 RiFeGAN: Rich Feature Generation for Text-to-Image Synthesis From Prior Knowledge
abstract
Text-to-image synthesis is a challenging task that generates realistic images from a textual sequence, which usually contains limited information compared with the corresponding image and so is ambiguous and abstractive. The limited textual information only describes a scene partly, which will complicate the generation with complementing the other details implicitly and lead to low-quality images. To address this problem, we propose a novel rich feature generating text-to-image synthesis, called RiFeGAN, to enrich the given description. In order to provide additional visual details and avoid conflicting, RiFeGAN exploits an attention-based caption matching model to select and refine the compatible candidate captions from prior knowledge. Given enriched captions, RiFeGAN uses self-attentional embedding mixtures to extract features across them effectually and handle the diverging features further. Then it exploits multi-captions attentional generative adversarial networks to synthesize images from those features. The experiments conducted on widely-used datasets show that the models can generate images from enriched captions effectually and improve the results significantly.
Jun Cheng 0002, Fuxiang Wu, Yanling Tian, Lei Wang 0018, Dapeng Tao
CVPR5
2020 Spatial-spectral weighted nuclear norm minimization for hyperspectral image denoising
Xinjian Huang, Bo Du 0001, Dapeng Tao, Liangpei Zhang 0001
Neurocomputing3
2020 Hetero-Center loss for cross-modality person Re-identification
Yuanxin Zhu, Zhao Yang 0001, Li Wang 0067, Sai Zhao, Dapeng Tao
Neurocomputing6
2020 HesGCN: Hessian graph convolutional networks for semi-supervised classification
Sichao Fu, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou, Liqiang Nie
Inf. Sci.3
2020 Domain Adaptation with Few Labeled Source Samples by Graph Regularization
Weifeng Liu 0001, Yicong Zhou, Dapeng Tao, Liqiang Nie
Neural Process. Lett.4
2020 Output Layer Multiplication for Class Imbalance Problem in Convolutional Neural Networks
Zhao Yang 0001, Yuanxin Zhu, Sai Zhao, Yunyan Wang, Dapeng Tao
Neural Process. Lett.6
2020 Online learning using projections onto shrinkage closed balls for adaptive brain-computer interface
Jun Cheng 0002, Dapeng Tao
Pattern Recognit.3
2020 Attribute-Identity Embedding and Self-Supervised Learning for Scalable Person Re-Identification
abstract
Due to the domain shift between source dataset and target dataset, most of the existing person re-identification (PRID) algorithms trained by a supervised learning framework often fail to be well generalized to another domain. To address this challenge, we propose a self-supervised learning algorithm based on attribute-identity embedding, which can incrementally optimize the model by selecting unlabeled samples from target domain. Thus the gap between source domain and target domain is bridged. Specifically, we first develop an attribute-identity joint prediction dictionary learning model for simultaneously learning a latent attribute space, a semantic attribute dictionary and an identifier. In our method, the predicted attribute from latent attribute space is used as a bridge to establish a preliminary link between different domains so as to predict the label of the target data sample. Second, to exploit the latent label contained in the predicted samples, we propose a prediction-training cycle self-supervised learning to tune the model variables to make them more adaptive in the target domain. Finally, the similarity measurement of pedestrians is achieved by combining the attribute space with latent identity space. The experiments show that the developed method outperforms some state-of-the-art supervised PRID methods and unsupervised PRID algorithms.
Huafeng Li 0001, Shuanglin Yan, Zhengtao Yu 0001, Dapeng Tao
IEEE Trans. Circuits Syst. Video Technol.4
2020 Constrained Discriminative Projection Learning for Image Classification
abstract
Projection learning is widely used in extracting discriminative features for classification. Although numerous methods have already been proposed for this goal, they barely explore the label information during projection learning and fail to obtain satisfactory performance. Besides, many existing methods can learn only a limited number of projections for feature extraction which may degrade the performance in recognition. To address these problems, we propose a novel constrained discriminative projection learning (CDPL) method for image classification. Specifically, CDPL can be formulated as a joint optimization problem over subspace learning and classification. The proposed method incorporates the low-rank constraint to learn a robust subspace which can be used as a bridge to seamlessly connect the original visual features and objective outputs. A regression function is adopted to explicitly exploit the class label information so as to enhance the discriminability of subspace. Unlike existing methods, we use two matrices to perform feature learning and regression, respectively, such that the proposed approach can obtain more projections and achieve superior performance in classification tasks. The experiments on several datasets show clearly the advantages of our method against other state-of-the-art methods.
Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu, Dapeng Tao
IEEE Trans. Image Process.5
2020 Saliency Detection via a Multiple Self-Weighted Graph-Based Manifold Ranking
abstract
As an important task in the process of image understanding and analysis, saliency detection has recently received increasing attention. In this paper, we propose an efficient multiple self-weighted graph-based manifold ranking method to construct salient maps. First, we extract several different views of features from superpixels, and generate original salient regions as foreground and background cues using boundary information via multiple graph-based manifold ranking. Furthermore, a set of hyperparameters is learned to distinguish the importance between different graphs, which can be viewed as an adaptive weighting of each graph, and then a centroid graph is generated by using these self-weighted multiple graphs. An iterative algorithm is proposed to simultaneously optimize the hyperparameters as well as the centroid graph connection. Thus, an ideal centroid graph can be obtained, offering a more clear profile of the separated structure. Finally, the saliency maps can be produced with an approximate binary image from the manifold ranking. Extensive experiments have demonstrated our method consistently achieves superior detection performance than several state-of-the-arts.
Cheng Deng 0002, Xu Yang 0019, Feiping Nie 0001, Dapeng Tao
IEEE Trans. Multim.4
2020 A Cuboid CNN Model With an Attention Mechanism for Skeleton-Based Action Recognition
abstract
The introduction of depth sensors such as Microsoft Kinect have driven research in human action recognition. Human skeletal data collected from depth sensors convey a significant amount of information for action recognition. While there has been considerable progress in action recognition, most existing skeleton-based approaches neglect the fact that not all human body parts move during many actions, and they fail to consider the ordinal positions of body joints. Here, and motivated by the fact that an action's category is determined by local joint movements, we propose a cuboid model for skeleton-based action recognition. Specifically, a cuboid arranging strategy is developed to organize the pairwise displacements between all body joints to obtain a cuboid action representation. Such a representation is well structured and allows deep CNN models to focus analyses on actions. Moreover, an attention mechanism is exploited in the deep model, such that the most relevant features are extracted. Extensive experiments on our new Yunnan University-Chinese Academy of Sciences-Multimodal Human Action Dataset (CAS-YNU MHAD), the NTU RGB+D dataset, the UTD-MHAD dataset, and the UTKinect-Action3D dataset demonstrate the effectiveness of our method compared to the current state-of-the-art.
Kaijun Zhu, Ruxin Wang 0002, Jun Cheng 0002, Dapeng Tao
IEEE Trans. Multim.5
2019 Semantic Adversarial Network with Multi-Scale Pyramid Attention for Video Classification
abstract
Two-stream architecture have shown strong performance in video classification task. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, there are some problems within such architecture. First, it relies on optical flow to model temporal information, which are often expensive to compute and store. Second, it has limited ability to capture details and local context information for video data. Third, it lacks explicit semantic guidance that greatly decrease the classification performance. In this paper, we proposed a new two-stream based deep framework for video classification to discover spatial and temporal information only from RGB frames, moreover, the multi-scale pyramid attention (MPA) layer and the semantic adversarial learning (SAL) module is introduced and integrated in our framework. The MPA enables the network capturing global and local feature to generate a comprehensive representation for video, and the SAL can make this representation gradually approximate to the real video semantics in an adversarial manner. Experimental results on two public benchmarks demonstrate our proposed methods achieves state-of-the-art results on standard video datasets.
De Xie, Cheng Deng 0002, Hao Wang 0062, Chao Li 0033, Dapeng Tao
AAAI5
2019 Student Becoming the Master: Knowledge Amalgamation for Joint Scene Parsing, Depth Estimation, and More
abstract
In this paper, we investigate a novel deep-model reusing task. Our goal is to train a lightweight and versatile student model, without human-labelled annotations, that amalgamates the knowledge and masters the expertise of two pre-trained teacher models working on heterogeneous problems, one on scene parsing and the other on depth estimation. To this end, we propose an innovative training strategy that learns the parameters of the student intertwined with the teachers, achieved by ``projecting'' its amalgamated features onto each teacher's domain and computing the loss. We also introduce two options to generalize the proposed training strategy to handle three or more tasks simultaneously. The proposed scheme yields very encouraging results. As demonstrated on several benchmarks, the trained student model achieves results even superior to those of the teachers in their own expertise domains and on par with the state-of-the-art fully supervised models relying on human-labelled annotations.
Jingwen Ye, Yixin Ji, Xinchao Wang, Kairi Ou, Dapeng Tao, Mingli Song
CVPR5
2019 Embedded Block Residual Network: A Recursive Restoration Model for Single-Image Super-Resolution
abstract
Single-image super-resolution restores the lost structures and textures from low-resolved images, which has achieved extensive attention from the research community. The top performers in this field include deep or wide convolutional neural networks, or recurrent neural networks. However, the methods enforce a single model to process all kinds of textures and structures. A typical operation is that a certain layer restores the textures based on the ones recovered by the preceding layers, ignoring the characteristics of image textures. In this paper, we believe that the lower-frequency and higher-frequency information in images have different levels of complexity and should be restored by models of different representational capacity. Inspired by this, we propose a novel embedded block residual network (EBRN) which is an incremental recovering progress for texture super-resolution. Specifically, different modules in the model restores information of different frequencies. For lower-frequency information, we use shallower modules of the network to recover; for higher-frequency information, we use deeper modules to restore. Extensive experiments indicate that the proposed EBRN model achieves superior performance and visual improvements against the state-of-the-arts.
Yajun Qiu, Ruxin Wang 0002, Dapeng Tao, Jun Cheng 0002
ICCV3
2019 Knowledge Amalgamation from Heterogeneous Networks by Common Feature Learning
abstract
An increasing number of well-trained deep networks have been released online by researchers and developers, enabling the community to reuse them in a plug-and-play way without accessing the training annotations. However, due to the large number of network variants, such public-available trained models are often of different architectures, each of which being tailored for a specific task or dataset. In this paper, we study a deep-model reusing task, where we are given as input pre-trained networks of heterogeneous architectures specializing in distinct tasks, as teacher models. We aim to learn a multitalented and light-weight student model that is able to grasp the integrated knowledge from all such heterogeneous-structure teachers, again without accessing any human annotation. To this end, we propose a common feature learning scheme, in which the features of all teachers are transformed into a common space and the student is enforced to imitate them all so as to amalgamate the intact knowledge. We test the proposed approach on a list of benchmarks and demonstrate that the learned student is able to achieve very promising performance, superior to those of the teachers in their specialized tasks.
Sihui Luo 0001, Xinchao Wang, Gongfan Fang, Dapeng Tao, Mingli Song
IJCAI5
2019 Pseudo Supervised Matrix Factorization in Discriminative Subspace
abstract
Non-negative Matrix Factorization (NMF) and spectral clustering have been proved to be efficient and effective for data clustering tasks and have been applied to various real-world scenes. However, there are still some drawbacks in traditional methods: (1) most existing algorithms only consider high-dimensional data directly while neglect the intrinsic data structure in the low-dimensional subspace; (2) the pseudo-information got in the optimization process is not relevant to most spectral clustering and manifold regularization methods. In this paper, a novel unsupervised matrix factorization method, Pseudo Supervised Matrix Factorization (PSMF), is proposed for data clustering. The main contributions are threefold: (1) to cluster in the discriminant subspace, Linear Discriminant Analysis (LDA) combines with NMF to become a unified framework; (2) we propose a pseudo supervised manifold regularization term which utilizes the pseudo-information to instruct the regularization term in order to find subspace that discriminates different classes; (3) an efficient optimization algorithm is designed to solve the proposed problem with proved convergence. Extensive experiments on multiple benchmark datasets illustrate that the proposed model outperforms other state-of-the-art clustering algorithms.
Jiaqi Ma 0002, Yipeng Zhang 0001, Lefei Zhang, Bo Du 0001, Dapeng Tao
IJCAI5
2019 Top distance regularized projection and dictionary learning for person re-identification
Huafeng Li 0001, Jinting Zhu, Dapeng Tao, Zhengtao Yu 0001
Inf. Sci.4
2019 Hessian-Regularized Multitask Dictionary Learning for Remote Sensing Image Recognition
abstract
Learning effective image representations is a vital issue for remote sensing (RS) image recognition tasks. Although numerous algorithms have been proposed, it is still challenging due to the limited labeled data. One representative work is the Laplacian-regularized multitask dictionary learning (LR-MTDL) that employs graph Laplacian regularization terms to fully utilize both the labeled and unlabeled information. However, it probably conduces to poor extrapolating power because Laplacian regularization biases the solution toward a constant function. In this letter, we propose a Hessian-regularized multitask dictionary learning to learn a source-data set-shared but target-data set-biased representation for RS image recognition. Particularly, Hessian can properly exploit the intrinsic local geometry of the data manifold and finally leverage the performance. Extensive experiments on four RS image data sets validate the effectiveness of the proposed method by comparing with baseline algorithms including single-task dictionary learning and LR-MTDL.
Guanhua Feng, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou
IEEE Geosci. Remote. Sens. Lett.4
2019 Effective human action recognition by combining manifold regularization and pairwise constraints
Xueqi Ma, Dapeng Tao, Weifeng Liu 0001
Multim. Tools Appl.2
2019 A flexible vehicle surround view camera system by central-around coordinate mapping model
Zhao Yang 0001, Lihua Zhou, Dapeng Tao
Multim. Tools Appl.6
2019 Hessian Regularized Distance Metric Learning for People Re-Identification
Guanhua Feng, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou
Neural Process. Lett.3
2019 A tensor framework for geosensor data forecasting of significant societal events
Lihua Zhou, Guowang Du, Ruxin Wang 0002, Dapeng Tao, Lizhen Wang 0001, Jun Cheng 0002
Pattern Recognit.4
2019 Action Parsing-Driven Video Summarization Based on Reinforcement Learning
abstract
How to manage, store, and index large numbers of videos is an urgent problem to be solved. Although there are many video summarization models achieving good results, models based on low-level features cannot summarize important semantic information and models based on semantic analysis need related text descriptions that do not exist for most videos. As a consequence, the mining semantic information contained in the video itself is a more feasible way. In this paper, we propose an action parsing-driven video summarization model based on reinforcement learning. The model is mainly divided into two parts, video cut by action parsing and video summarization based on reinforcement learning. In the first part, a sequential multiple instance learning model is trained with weakly annotated data to solve the problem of full annotation’s time consuming and weak annotation’s ambiguity. In the second part, we design a deep recurrent neural network-based video summarization model that selects the most distinguishable frames comparing with other actions. Meanwhile, the quality of the extracted key frames could be evaluated by the categorization accuracy. Experiments and comparison with state-of-the-art methods demonstrate the advantage of the proposed approach.
Jie Lei 0002, Qiao Luan, Xinhui Song, Xiao Liu 0012, Dapeng Tao, Mingli Song
IEEE Trans. Circuits Syst. Video Technol.5
2019 $p$ -Laplacian Regularization for Scene Recognition
abstract
The explosive growth of multimedia data on the Internet makes it essential to develop innovative machine learning algorithms for practical applications especially where only a small number of labeled samples are available. Manifold regularized semi-supervised learning (MRSSL) thus received intensive attention recently because it successfully exploits the local structure of data distribution including both labeled and unlabeled samples to leverage the generalization ability of a learning model. Although there are many representative works in MRSSL, including Laplacian regularization (LapR) and Hessian regularization, how to explore and exploit the local geometry of data manifold is still a challenging problem. In this paper, we introduce a fully efficient approximation algorithm of graph p -Laplacian, which significantly saving the computing cost. And then we propose p -LapR (pLapR) to preserve the local geometry. Specifically, p -Laplacian is a natural generalization of the standard graph Laplacian and provides convincing theoretical evidence to better preserve the local structure. We apply pLapR to support vector machines and kernel least squares and conduct the implementations for scene recognition. Extensive experiments on the Scene 67 dataset, Scene 15 dataset, and UC-Merced dataset validate the effectiveness of pLapR in comparison to the conventional manifold regularization methods.
Weifeng Liu 0001, Xueqi Ma, Yicong Zhou, Dapeng Tao, Jun Cheng 0002
IEEE Trans. Cybern.4
2019 Hypergraph $p$ -Laplacian Regularization for Remotely Sensed Image Recognition
abstract
Graph-based and manifold-regularization (MR)-based semisupervised learning, including Laplacian regularization (LapR) and hypergraph LapR (HLapR), have achieved prominent performance in preserving locality and similarity information. However, it is still a great challenge to exactly explore and exploit the local structure of the data distribution. In this paper, we present an efficient and effective approximation algorithm of hypergraph${p}$-Laplacian and then propose hypergraph${p}$-LapR (HpLapR) to preserve the geometry of the probability distribution. In particular, hypergraph is a generalization of a standard graph while hypergraph${p}$-Laplacian is a nonlinear generalization of the standard graph Laplacian. The proposed HpLapR shows great potential to exploit the local structures. We integrate HpLapR with logistic regression for remote sensing image recognition. Experiments on UC-Merced data set demonstrate that the proposed HpLapR has superior performance compared with several popular MR methods including LapR and HLapR.
Xueqi Ma, Weifeng Liu 0001, Dapeng Tao, Yicong Zhou
IEEE Trans. Geosci. Remote. Sens.4
2019 Domain-Weighted Majority Voting for Crowdsourcing
abstract
Crowdsourcing labeling systems provide an efficient way to generate multiple inaccurate labels for given observations. If the competence level or the "reputation," which can be explained as the probabilities of annotating the right label, for each crowdsourcing annotators is equal and biased to annotate the right label, majority voting (MV) is the optimal decision rule for merging the multiple labels into a single reliable one. However, in practice, the competence levels of annotators employed by the crowdsourcing labeling systems are often diverse very much. In these cases, weighted MV is more preferred. The weights should be determined by the competence levels. However, since the annotators are anonymous and the ground-truth labels are usually unknown, it is hard to compute the competence levels of the annotators directly. In this paper, we propose to learn the weights for weighted MV by exploiting the expertise of annotators. Specifically, we model the domain knowledge of different annotators with different distributions and treat the crowdsourcing problem as a domain adaptation problem. The annotators provide labels to the source domains and the target domain is assumed to be associated with the ground-truth labels. The weights are obtained by matching the source domains with the target domain. Although the target-domain labels are unknown, we prove that they could be estimated under mild conditions. Both theoretical and empirical analyses verify the effectiveness of the proposed method. Large performance gains are shown for specific data sets.
Dapeng Tao, Jun Cheng 0002, Zhengtao Yu 0001, Kun Yue, Lizhen Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Person re-identification by discriminant analytical least squares metric learning
Zhao Yang 0001, Fei Dai 0002, Jianxin Pang, Dapeng Tao
Mach. Vis. Appl.6
2018 Joint medical image fusion, denoising and enhancement via discriminative low-rank sparse dictionaries learning
Huafeng Li 0001, Xiaoge He, Dapeng Tao, Yuan Yan Tang, Ruxin Wang 0002
Pattern Recognit.3
2018 Reinforcement online learning for emotion prediction by using physiological signals
Weifeng Liu 0001, Lianbo Zhang, Dapeng Tao, Jun Cheng 0002
Pattern Recognit. Lett.3
2018 Skeleton embedded motion body partition for human action recognition using depth sequences
Xiaopeng Ji, Jun Cheng 0002, Wei Feng 0009, Dapeng Tao
Signal Process.4
2018 Ensemble One-Dimensional Convolution Neural Networks for Skeleton-Based Action Recognition
abstract
This letter proposes an ensemble neural network (Ensem-NN) for skeleton-based action recognition. The Ensem-NN is introduced based on the idea of ensemble learning, “two heads are better than one.” According to the property of skeleton sequences, we design one-dimensional convolution neural network with residual structure asBase-Net. From entirety to local, from focus to motion, we designed four different subnets based on theBase-Netto extract diverse features. The first subnet is aTwo-stream Entirety Net, which performs on the entirety skeleton and explores both temporal and spatial features. The second is aBody-part Net, which can extract fine-grained spatial and temporal features. The third is anAttention Net, in which a channel-wised attention mechanism can learn important frames and feature channels.Frame-difference Net, as the fourth subnet, aims at exploring motion features. Finally, the four subnets are fused as one ensemble network. Experimental results show that the proposed Ensem-NN performs better than state-of-the-art methods on three widely used datasets.
Yangyang Xu 0004, Jun Cheng 0002, Lei Wang 0018, Haiying Xia, Feng Liu 0013, Dapeng Tao
IEEE Signal Process. Lett.6
2018 Deep Multi-View Feature Learning for Person Re-Identification
abstract
Person re-identification aims to identify the same pedestrians across different camera views at different locations. This important yet difficult intelligent video analysis problem remains a vigorous area of research due to demands for performance improvements. Person re-identification involves two main steps: feature representation and metric learning. Handcrafted features, such as color and texture histograms, are frequently used for person re-identification, but most handcrafted features are limited by not being directly applicable to practical problems. Deep learning methods have obtained the state-of-the-art performance in a wide variety of applications, including image annotation, face recognition, and speech recognition. However, deep learning features are heavily dependent on large-scale labeling of samples. In this paper, by utilizing the Cross-view Quadratic Discriminant Analysis (XQDA) metric learning, we propose a novel scheme called deep multi-view feature learning (DMVFL), which exploits the collaboration between handcrafted and deep learning features in a simple but effective way. Furthermore, we prove that the XQDA is a robust algorithm. Extensive experiments on two challenging person re-identification data sets (VIPeR and GRID) demonstrate that DMVFL improves on current state-of-the-art methods.
Dapeng Tao, Yanan Guo 0003, Baosheng Yu, Jianxin Pang, Zhengtao Yu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Tensor Rank Preserving Discriminant Analysis for Facial Recognition
abstract
Facial recognition, one of the basic topics in computer vision and pattern recognition, has received substantial attention in recent years. However, for those traditional facial recognition algorithms, the facial images are reshaped to a long vector, thereby losing part of the original spatial constraints of each pixel. In this paper, a new tensor-based feature extraction algorithm termed tensor rank preserving discriminant analysis (TRPDA) for facial image recognition is proposed; the proposed method involves two stages: in the first stage, the low-dimensional tensor subspace of the original input tensor samples was obtained; in the second stage, discriminative locality alignment was utilized to obtain the ultimate vector feature representation for subsequent facial recognition. On the one hand, the proposed TRPDA algorithm fully utilizes the natural structure of the input samples, and it applies an optimization criterion that can directly handle the tensor spectral analysis problem, thereby decreasing the computation cost compared those traditional tensor-based feature selection algorithms. On the other hand, the proposed TRPDA algorithm extracts feature by finding a tensor subspace that preserves most of the rank order information of the intra-class input samples. Experiments on the three facial databases are performed here to determine the effectiveness of the proposed TRPDA algorithm.
Dapeng Tao, Yanan Guo 0003, Yaotang Li, Xinbo Gao 0001
IEEE Trans. Image Process.1
2017 An adaptive semi-supervised clustering approach via multiple density-based information
Yun Yang 0003, Wei Wang 0140, Dapeng Tao
Neurocomputing4
2017 Canonical correlation analysis networks for two-view image recognition
Xinghao Yang, Weifeng Liu 0001, Dapeng Tao, Jun Cheng 0002
Inf. Sci.3
2017 Support vector machine active learning by Hessian regularization
Weifeng Liu 0001, Lianbo Zhang, Dapeng Tao, Jun Cheng 0002
J. Vis. Commun. Image Represent.3
2017 The spatial Laplacian and temporal energy pyramid representation for human action recognition using depth sequences
Xiaopeng Ji, Jun Cheng 0002, Dapeng Tao, Xinyu Wu 0001, Wei Feng 0009
Knowl. Based Syst.3
2017 Multiview Canonical Correlation Analysis Networks for Remote Sensing Image Recognition
abstract
In the past decade, deep learning (DL) algorithms have been widely used for remote sensing (RS) image recognition tasks. As the most typical DL model, convolutional neural networks (CNNs) achieves outstand performance for big RS data classification. Recently, a variant of CNN, dubbed canonical correlation analysis network (CCANet), was proposed to abstract the two-view image features. Extensive experiments conducted on several benchmark databases validate the effectiveness of CCANet. However, the CCANet structure is powerless when the observations arrive from more than two sources. To serve the multiview purpose, in this letter, we propose multiview CCANets (MCCANets). Particularly, the MCCANet model learns the stacked multiperspective filter banks by the MCCA method and builds a deep convolutional structure. In the output stage, the binarization and the blockwise histogram are employed as nonlinear processing and feature pooling, respectively. To access the effectiveness of the MCCANet, we conduct a host of experiments on the RSSCN7 RS database. Extensive experimental results demonstrate that the MCCANet outperforms the two-view CCANet.
Xinghao Yang, Weifeng Liu 0001, Dapeng Tao, Jun Cheng 0002
IEEE Geosci. Remote. Sens. Lett.3
2017 Cauchy Estimator Discriminant Learning for RGB-D Sensor-based Scene Classification
Dapeng Tao, Xipeng Yang, Weifeng Liu 0001, Shuifa Sun, Yanan Guo 0003, Jianxin Pang
Multim. Tools Appl.1
2017 LMAE: A large margin Auto-Encoders for classification
Weifeng Liu 0001, Tengzhou Ma, Qiangsheng Xie, Dapeng Tao, Jun Cheng 0002
Signal Process.4
2017 Robust Sparse Coding for Mobile Image Labeling on the Cloud
abstract
With the rapid development of the mobile service and online social networking service, a large number of mobile images are generated and shared on the social networks every day. The visual content of these images contains rich knowledge for many uses, such as social categorization and recommendation. Mobile image labeling has, therefore, been proposed to understand the visual content and received intensive attention in recent years. In this paper, we present a novel mobile image labeling scheme on the cloud, in which mobile images are first and efficiently transmitted to the cloud by Hamming compressed sensing, such that the heavy computation for image understanding is transferred to the cloud for quick response to the queries of the users. On the cloud, we design a sparse correntropy framework for robustly learning the semantic content of mobile images, based on which the relevant tags are assigned to the query images. The proposed framework (called maximum correntropy-based mobile image labeling) is very insensitive to the noise and the outliers, and is optimized by a half-quadratic optimization technique. We theoretically show that our image labeling approach is more robust than the squared loss, absolute loss, Cauchy loss, and many other robust loss function-based sparse coding methods. To further understand the proposed algorithm, we also derive its robustness and generalization error bounds. Finally, we conduct experiments on the PASCAL VOC’07 data set and empirically demonstrate the effectiveness of the proposed robust sparse coding method for mobile image labeling.
Dapeng Tao, Jun Cheng 0002, Xinbo Gao 0001, Xuelong Li 0001, Cheng Deng 0002
IEEE Trans. Circuits Syst. Video Technol.1
2017 AllFocus: Patch-Based Video Out-of-Focus Blur Reconstruction
abstract
Amateur videos always contain focusing issues. A focusing mistake may produce out-of-focus blur, which seriously degrades the expressive force of the video. In this paper, we propose a patch-based method to remove the out-of-focus blur of a video and build an all-in-focus video. We assume that the out-of-focus blurry region in one frame will be clear in a portion of other frames; thus, the clear corresponding regions can be used to reconstruct the blurry one. We divide each video frame into a grid of patches and track each patch in the surrounding frames. We independently reconstruct each video frame by building a Markov random field model to identify the optimal target patches that are sharp, similar to the original patches, and are coherent with their neighboring patches within the overlapped regions. To recover an all-in-focus video, an iterative framework is utilized, in which the reconstructed video of each iteration is substituted in the next iteration. Finally, we employ the idea of a bilateral filter to temporally smooth the reconstructed video. The experimental results and the comparison with the previous works demonstrate the effectiveness of our method.
Yinting Wang, Zhenyang Wang, Dapeng Tao, Shaojie Zhuo, Xianghua Xu, Shiliang Pu, Mingli Song
IEEE Trans. Circuits Syst. Video Technol.3
2017 Latent Max-Margin Multitask Learning With Skelets for 3-D Action Recognition
abstract
Recent emergence of low-cost and easy-operating depth cameras has reinvigorated the research in skeleton-based human action recognition. However, most existing approaches overlook the intrinsic interdependencies between skeleton joints and action classes, thus suffering from unsatisfactory recognition performance. In this paper, a novel latent max-margin multitask learning model is proposed for 3-D action recognition. Specifically, we exploit skelets as the mid-level granularity of joints to describe actions. We then apply the learning model to capture the correlations between the latent skelets and action classes each of which accounts for a task. By leveraging structured sparsity inducing regularization, the common information belonging to the same class can be discovered from the latent skelets, while the private information across different classes can also be preserved. The proposed model is evaluated on three challenging action data sets captured by depth cameras. Experimental results show that our model consistently achieves superior performance over recent state-of-the-art approaches.
Yanhua Yang, Cheng Deng 0002, Dapeng Tao, Shaoting Zhang 0001, Wei Liu 0005, Xinbo Gao 0001
IEEE Trans. Cybern.3
2017 Coherent Semantic-Visual Indexing for Large-Scale Image Retrieval in the Cloud
abstract
The rapidly increasing number of images on the internet has further increased the need for efficient indexing for digital image searching of large databases. The design of a cloud service that provides high efficiency but compact image indexing remains challenging, partly due to the well-known semantic gap between user queries and the rich semantics of large-scale data sets. In this paper, we construct a novel joint semantic-visual space by leveraging visual descriptors and semantic attributes, which narrows the semantic gap by combining both attributes and indexing into a single framework. Such a joint space embraces the flexibility of coherent semantic-visual indexing, which employs binary codes to boost retrieval speed while maintaining accuracy. To solve the proposed model, we make the following contributions. First, we propose an interactive optimization method to find the joint semantic and visual descriptor space. Second, we prove convergence of our optimization algorithm, which guarantees a good solution after a certain number of iterations. Third, we integrate the semantic-visual joint space system with spectral hashing, which finds an efficient solution to search up to billion-scale data sets. Finally, we design an online cloud service to provide a more efficient online multimedia service. Experiments on two standard retrieval datasets (i.e., Holidays1M, Oxford5K) show that the proposed method is promising compared with the current state-of-the-art and that the cloud system significantly improves performance.
Richang Hong, Lei Li 0002, Dapeng Tao, Meng Wang 0001, Qi Tian 0001
IEEE Trans. Image Process.4
2017 Large Sparse Cone Non-negative Matrix Factorization for Image Annotation
abstract
Image annotation assigns relevant tags to query images based on their semantic contents. Since Non-negative Matrix Factorization (NMF) has the strong ability to learn parts-based representations, recently, a number of algorithms based on NMF have been proposed for image annotation and have achieved good performance. However, most of the efforts have focused on the representations of images and annotations. The properties of the semantic parts have not been well studied. In this article, we revisit the sparseness-constrained NMF (sNMF) proposed by Hoyer [2004]. By endowing the sparseness constraint with a geometric interpretation and sNMF with theoretical analyses of the generalization ability, we show that NMF with such a sparseness constraint has three advantages for image annotation tasks: (i) The sparseness constraint is more ℓ 0 -norm oriented than the ℓ 1 -norm-based sparseness, which significantly enhances the ability of NMF to robustly learn semantic parts. (ii) The sparseness constraint has a large cone interpretation and thus allows the reconstruction error of NMF to be smaller, which means that the learned semantic parts are more powerful to represent images for tagging. (iii) The learned semantic parts are less correlated, which increases the discriminative ability for annotating images. Moreover, we present a new efficient large sparse cone NMF (LsCNMF) algorithm to optimize the sNMF problem by employing the Nesterov’s optimal gradient method. We conducted experiments on the PASCAL VOC07 dataset and demonstrated the effectiveness of LsCNMF for image annotation.
Dapeng Tao, Dacheng Tao, Xuelong Li 0001, Xinbo Gao 0001
ACM Trans. Intell. Syst. Technol.1
2017 Discriminative Multi-instance Multitask Learning for 3D Action Recognition
abstract
As the prosperity of low-cost and easy-operating depth cameras, skeleton-based human action recognition has been extensively studied recently. However, most of the existing methods partially consider that all 3D joints of a human skeleton are identical. Actually, these 3D joints exhibit diverse responses to different action classes, and some joint configurations are more discriminative to distinguish a certain action. In this paper, we propose a discriminative multi-instance multitask learning (MIMTL) framework to discover the intrinsic relationship between joint configurations and action classes. First, a set of discriminative and informative joint configurations for the corresponding action class is captured in multi-instance learning model by regarding the action and the joint configurations as a bag and its instances, respectively. Then, a multitask learning model with group structure constraints is exploited to further reveal the intrinsic relationship between the joint configurations and different action classes. We conduct extensive evaluations of MIMTL using three benchmark 3D action recognition datasets. Experimental results show that our proposed MIMTL framework performs favorably compared with several state-of-the-art approaches.
Yanhua Yang, Cheng Deng 0002, Shangqian Gao, Wei Liu 0005, Dapeng Tao, Xinbo Gao 0001
IEEE Trans. Multim.5
2017 Multiview Cauchy Estimator Feature Embedding for Depth and Inertial Sensor-Based Human Action Recognition
abstract
The ever-growing popularity of Kinect and inertial sensors has prompted intensive research efforts on human action recognition. Since human actions were extracted from Kinect and inertial sensors, they can be characterized by multiple feature representations. By encoding the multiview features into a unified space, it could be optimal for human action recognition. In this paper, we propose a new unsupervised feature fusion method termed multiview Cauchy estimator feature embedding (MCEFE) for human action recognition. By minimizing empirical risk, MCEFE integrates the encoded complementary information in multiple views to find the unified data representation and the projection matrices. To enhance robustness to outliers, the Cauchy estimator is imposed on the reconstruction error. Furthermore, ensemble manifold regularization is enforced on the projection matrices to encode the correlations between different views and avoid overfitting. Experiments are conducted on the new Chinese Academy of Sciences—Yunnan University—multimodal human action database to demonstrate the effectiveness and robustness of MCEFE for human action recognition.
Yanan Guo 0003, Dapeng Tao, Weifeng Liu 0001, Jun Cheng 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2016 Hessian regularization by patch alignment framework
Weifeng Liu 0001, Dapeng Tao
Neurocomputing3
2016 Manifold regularized kernel logistic regression for web image annotation
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001, Ke Lu 0002
Neurocomputing3
2016 HSAE: A Hessian regularized sparse auto-encoders
Weifeng Liu 0001, Tengzhou Ma, Dapeng Tao, Jane You
Neurocomputing3
2016 Event-based large scale surveillance video summarization
Xinhui Song, Jie Lei 0002, Dapeng Tao, Guanhong Yuan, Mingli Song
Neurocomputing4
2016 Cauchy estimator discriminant analysis for face recognition
Xipeng Yang, Jun Cheng 0002, Wei Feng 0009, Zhengyao Bai, Dapeng Tao
Neurocomputing6
2016 Recent developments on deep big vision
Jun Yu 0002, Dapeng Tao, Richang Hong, Xinbo Gao 0001
Neurocomputing2
2016 Online tracking based on efficient transductive learning with sample matching costs
Peng Zhang 0005, Tao Zhuo, Yanning Zhang 0001, Dapeng Tao, Jun Cheng 0002
Neurocomputing4
2016 Multicolumn Bidirectional Long Short-Term Memory for Mobile Devices-Based Human Activity Recognition
abstract
The ever-growing popularity of mobile devices equipped with accelerometers has provided the opportunity to capture the semantic aspects of human activity and improve user experiences with behavior-based recommendations. These functions depend heavily on the accuracy of human activity recognition, and thus real applications that use mobile devices-based human activity recognition systems (MARSs) need to seamlessly incorporate the information carried by newly labeled training samples. Motivated by the success of the weightlessness feature, we propose a new two-directional feature for bidirectional long short-term memory (BLSTM) for incremental learning in human activity recognition. To further improve the performance, we also present a new ensemble classifier termed multicolumn BLSTM (MBLSTM), which effectively combines different acceleration signal features to further improve activity recognition accuracy. Experiments on the naturalistic mobile devices-based human activity dataset suggest that MBLSTM is superior to other state-of-the-art MARS methods.
Dapeng Tao, Yonggang Wen 0001, Richang Hong
IEEE Internet Things J.1
2016 Large-scale paralleled sparse principal component analysis
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001, Ke Lu 0002
Multim. Tools Appl.3
2016 Data-driven facial animation via semi-supervised local patch alignment
Jian Zhang 0026, Jun Yu 0002, Jane You, Dapeng Tao, Jun Cheng 0002
Pattern Recognit.4
2016 DeepChart: Combining deep convolutional networks and deep belief networks in chart classification
Binbin Tang, Xiao Liu 0012, Jie Lei 0002, Mingli Song, Dapeng Tao, Shuifa Sun, Fangmin Dong
Signal Process.5
2016 Real-time tracking-by-learning with high-order regularization fusion for big video abstraction
Peng Zhang 0005, Tao Zhuo, Yanning Zhang 0001, Lei Xie 0001, Dapeng Tao
Signal Process.5
2016 Principal Component 2-D Long Short-Term Memory for Font Recognition on Single Chinese Characters
abstract
Chinese character font recognition (CCFR) has received increasing attention as the intelligent applications based on optical character recognition becomes popular. However, traditional CCFR systems do not handle noisy data effectively. By analyzing in detail the basic strokes of Chinese characters, we propose that font recognition on a single Chinese character is a sequence classification problem, which can be effectively solved by recurrent neural networks. For robust CCFR, we integrate a principal component convolution layer with the 2-D long short-term memory (2DLSTM) and develop principal component 2DLSTM (PC-2DLSTM) algorithm. PC-2DLSTM considers two aspects: 1) the principal component layer convolution operation helps remove the noise and get a rational and complete font information and 2) simultaneously, 2DLSTM deals with the long-range contextual processing along scan directions that can contribute to capture the contrast between character trajectory and background. Experiments using the frequently used CCFR dataset suggest the effectiveness of PC-2DLSTM compared with other state-of-the-art font recognition methods.
Dapeng Tao, Xuelong Li 0001
IEEE Trans. Cybern.1
2016 Person Re-Identification by Dual-Regularized KISS Metric Learning
abstract
Person re-identification aims to match the images of pedestrians across different camera views from different locations. This is a challenging intelligent video surveillance problem that remains an active area of research due to the need for performance improvement. Person re-identification involves two main steps: feature representation and metric learning. Although the keep it simple and straightforward (KISS) metric learning method for discriminative distance metric learning has been shown to be effective for the person re-identification, the estimation of the inverse of a covariance matrix is unstable and indeed may not exist when the training set is small, resulting in poor performance. Here, we present dual-regularized KISS (DR-KISS) metric learning. By regularizing the two covariance matrices, DR-KISS improves on KISS by reducing overestimation of large eigenvalues of the two estimated covariance matrices and, in doing so, guarantees that the covariance matrix is irreversible. Furthermore, we provide theoretical analyses for supporting the motivations. Specifically, we first prove why the regularization is necessary. Then, we prove that the proposed method is robust for generalization. We conduct extensive experiments on three challenging person re-identification datasets, VIPeR, GRID, and CUHK 01, and show that DR-KISS achieves new state-of-the-art performance.
Dapeng Tao, Yanan Guo 0003, Mingli Song, Yaotang Li, Zhengtao Yu 0001, Yuan Yan Tang
IEEE Trans. Image Process.1
2016 Tensor Manifold Discriminant Projections for Acceleration-Based Human Activity Recognition
abstract
With the rapid development of wearable sensors and pervasive computing technologies including smartphones, acceleration-based human activity recognition is receiving increased attention for medical research applications. Motivated by the “weightlessness” feature, here we apply a bidirectional feature during the feature extraction phase of activity recognition; however, since the bidirectional feature has two components, they cannot simply be concatenated into a long vector, but can be naturally treated as a second-order tensor. Therefore, we propose a new tensor-based feature selection method termed tensor manifold discriminant projections (TMDP). TMDP simultaneously considers: 1) applying an optimization criterion that can directly process the tensor spectral analysis problem, thereby decreasing the computational cost compared to traditional tensor-based feature selection methods; 2) extracting local rank information by finding a tensor subspace that preserves the rank order information of the within-class input samples; and 3) extracting discriminant information by maximizing the sum of distances between every sample and their interclass sample mean. Experiments on the naturalistic mobile devices-based human activity 2.0 dataset are performed to demonstrate the effectiveness and robustness of TMDP.
Yanan Guo 0003, Dapeng Tao, Jun Cheng 0002, Alan William Dougherty, Yaotang Li, Kun Yue, Bob Zhang 0001
IEEE Trans. Multim.2
2016 Manifold Ranking-Based Matrix Factorization for Saliency Detection
abstract
Saliency detection is used to identify the most important and informative area in a scene, and it is widely used in various vision tasks, including image quality assessment, image matching, and object recognition. Manifold ranking (MR) has been used to great effect for the saliency detection, since it not only incorporates the local spatial information but also utilizes the labeling information from background queries. However, MR completely ignores the feature information extracted from each superpixel. In this paper, we propose an MR-based matrix factorization (MRMF) method to overcome this limitation. MRMF models the ranking problem in the matrix factorization framework and embeds query sample labels in the coefficients. By incorporating spatial information and embedding labels, MRMF enforces similar saliency values on neighboring superpixels and ranks superpixels according to the learned coefficients. We prove that the MRMF has good generalizability, and develops an efficient optimization algorithm based on the Nesterov method. Experiments using popular benchmark data sets illustrate the promise of MRMF compared with the other state-of-the-art saliency detection methods.
Dapeng Tao, Jun Cheng 0002, Mingli Song
IEEE Trans. Neural Networks Learn. Syst.1
2016 Ensemble Manifold Rank Preserving for Acceleration-Based Human Activity Recognition
abstract
With the rapid development of mobile devices and pervasive computing technologies, acceleration-based human activity recognition, a difficult yet essential problem in mobile apps, has received intensive attention recently. Different acceleration signals for representing different activities or even a same activity have different attributes, which causes troubles in normalizing the signals. We thus cannot directly compare these signals with each other, because they are embedded in a nonmetric space. Therefore, we present a nonmetric scheme that retains discriminative and robust frequency domain information by developing a novel ensemble manifold rank preserving (EMRP) algorithm. EMRP simultaneously considers three aspects: 1) it encodes the local geometry using the ranking order information of intraclass samples distributed on local patches; 2) it keeps the discriminative information by maximizing the margin between samples of different classes; and 3) it finds the optimal linear combination of the alignment matrices to approximate the intrinsic manifold lied in the data. Experiments are conducted on the South China University of Technology naturalistic 3-D acceleration-based activity dataset and the naturalistic mobile-devices based human activity dataset to demonstrate the robustness and effectiveness of the new nonmetric scheme for acceleration-based human activity recognition.
Dapeng Tao, Yuan Yuan 0001, Yang Xue 0001
IEEE Trans. Neural Networks Learn. Syst.1
2015 Chart classification by combining deep convolutional networks and deep belief networks
abstract
Chart classification is the foundation of chart analysis and document understanding. In this paper, we propose a novel framework to classify charts by combining convolutional networks and deep belief networks. In the framework, we firstly extract deep hidden features of charts, which are taken from the fully-connected layer of deep convolutional networks. We then utilize deep belief networks to predict the labels of the charts based on their deep hidden features. The convolutional networks are initialized using a large number of natural images and fine-tuned using the chart images to prevent overfitting. Compared with previous methods using primitive feature extraction, the deep features give our framework better scalability and stability. We collect a 5-class chart dataset with more than 5000 images and show that the proposed framework outperforms existing methods greatly.
Xiao Liu 0012, Binbin Tang, Zhenyang Wang, Xianghua Xu, Shiliang Pu, Dapeng Tao, Mingli Song
ICDAR6
2015 Local mean spatio-temporal feature for depth image-based speed-up action recognition
abstract
With the promptly growing population of the low-cost Microsoft Kinect sensor, action recognition, which is a hard yet important problem in computer vision, has been received substantial attention. However, most existing approaches in action recognition spend much time on feature detection even though these methods can achieve high recognition rates. In this paper, we propose a local mean spatio-temporal feature (LMSF) to speed up depth image based action recognition. In particular, we solve the problem from three aspects: (1) associate the 4D normals by a local mean spatio-temporal neighborhood; (2) extract motion frames by detecting the differences between consecutive frames; (3) reduce redundant normals extracted from depth cloud points by sparse coding. The proposed approach is tested on two public benchmark datasets, i.e., MSRAction3D and MSRGesture3D. Experimental results demonstrate the advantages of our improvement method and the state-of-the-art performance on processing speed.
Xiaopeng Ji, Jun Cheng 0002, Dapeng Tao
ICIP3
2015 Hessian Regularized Sparse Coding for Human Action Recognition
Weifeng Liu 0001, Zhen Wang 0004, Dapeng Tao, Jun Yu 0002
MMM (2)3
2015 A general framework for co-training and its applications
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001
Neurocomputing3
2015 DLANet: A manifold-learning-based discriminative feature learning network for scene classification
Ziyong Feng, Dapeng Tao, Shuangping Huang
Neurocomputing3
2015 Upper limb motion tracking with the integration of IMU and Kinect
Yushuang Tian, Xiaoli Meng, Dapeng Tao, Dongquan Liu
Neurocomputing3
2015 Anti-counterfeiting digital watermarking algorithm for printed QR barcode
Rongsheng Xie, Shunzhi Zhu, Dapeng Tao
Neurocomputing4
2015 Human pose recovery by supervised spectral embedding
Jun Yu 0002, Yukun Guo, Dapeng Tao, Jian Wan 0001
Neurocomputing3
2015 Monocular face reconstruction with global and local shape constraints
Jian Zhang 0026, Dapeng Tao, Xiangjuan Bian, Xiaosi Zhan
Neurocomputing2
2015 Discriminative dictionary learning via Fisher discrimination K-SVD algorithm
Dapeng Tao
Neurocomputing2
2015 Multi-view ensemble manifold regularization for 3D object recognition
Jun Yu 0002, Jane You, Dapeng Tao
Inf. Sci.5
2015 Local structure preserving discriminative projections for RGB-D sensor-based scene classification
Dapeng Tao, Jun Cheng 0002
Inf. Sci.1
2015 Multiview Hessian regularized logistic regression for action recognition
Weifeng Liu 0001, Dapeng Tao, Yanjiang Wang 0001, Ke Lu 0002
Signal Process.3
2015 Semantic embedding for indoor scene recognition by weighted hypergraph learning
Jun Yu 0002, Dapeng Tao, Meng Wang 0001
Signal Process.3
2015 Person Reidentification by Minimum Classification Error-Based KISS Metric Learning
abstract
In recent years, person reidentification has received growing attention with the increasing popularity of intelligent video surveillance. This is because person reidentification is critical for human tracking with multiple cameras. Recently, keep it simple and straightforward (KISS) metric learning has been regarded as a top level algorithm for person reidentification. The covariance matrices of KISS are estimated by maximum likelihood (ML) estimation. It is known that discriminative learning based on the minimum classification error (MCE) is more reliable than classical ML estimation with the increasing of the number of training samples. When considering a small sample size problem, direct MCE KISS does not work well, because of the estimate error of small eigenvalues. Therefore, we further introduce the smoothing technique to improve the estimates of the small eigenvalues of a covariance matrix. Our new scheme is termed the minimum classification error-KISS (MCE-KISS). We conduct thorough validation experiments on the VIPeR and ETHZ datasets, which demonstrate the robustness and effectiveness of MCE-KISS for person reidentification.
Dapeng Tao, Yongfei Wang, Xuelong Li 0001
IEEE Trans. Cybern.1
2014 Genetic algorithm for spanning tree construction in P2P distributed interactive applications
Yusen Li, Jun Yu 0002, Dapeng Tao
Neurocomputing3
2014 Sparse Discriminative Information Preservation for Chinese character font categorization
Dapeng Tao, Shuye Zhang, Zhao Yang 0001, Yongfei Wang
Neurocomputing1
2014 Sparse frontal face image synthesis from an arbitrary profile image
Lin Zhao 0003, Xinbo Gao 0001, Yuan Yuan 0001, Dapeng Tao
Neurocomputing4
2014 Semantic preserving distance metric learning and applications
Jun Yu 0002, Dapeng Tao, Jonathan Li 0001, Jun Cheng 0002
Inf. Sci.2
2014 Grassmann multimodal implicit feature selection
Dapeng Tao, Xiao Liu 0012, Mingli Song, Chun Chen 0001
Multim. Syst.2
2014 Motionlet LLC coding for discriminative human pose estimation
Mingli Song, Dapeng Tao, Jiajun Bu, Chun Chen 0001
Multim. Tools Appl.3
2014 Similar handwritten Chinese character recognition by kernel discriminative locality alignment
Dapeng Tao, Lingyu Liang, Yan Gao 0011
Pattern Recognit. Lett.1
2014 Rank Preserving Discriminant Analysis for Human Behavior Recognition on Wireless Sensor Networks
abstract
With the rapid development of the intelligent sensing and the prompt growing industrial safety demands, human behavior recognition has received a great deal of attentions in industrial informatics. To deploy an utmost scalable, flexible, and robust human behavior recognition system, we need both innovative sensing electronics and suitable intelligence algorithms. Wireless sensor networks (WSNs) open a novel way for human behavior recognition, because the heavy computation can be immediately transferred to a network server. In this paper, a new scheme for human behavior recognition on WSNs is proposed, which transmits activities' signals compressed by Hamming compressed sensing to the network server and conducts behavior recognition through a collaboration between a new dimension reduction algorithm termed rank preserving discriminant analysis (RPDA) and a nearest neighbor classifier. RPDA encodes local rank information of within-class samples and discriminative information of the between-class under the framework of Patch Alignment Framework. Experiments are conducted on the SCUT Naturalistic 3D Acceleration-based Activity (SCUT NAA) dataset and demonstrate the effectiveness of RPDA for human behavior recognition.
Dapeng Tao, Yongfei Wang, Xuelong Li 0001
IEEE Trans. Ind. Informatics1
2013 A Faster Method for Chinese Font Recognition Based on Harris Corner
abstract
A new fast font recognition method based on Harris corners is proposed in this paper. According to our observation, we find out that the font discriminative information is mainly hidden in some interesting points of characters. Based on this discovery, we extract the texture feature of the interesting points. The Harris corners are extracted as the interesting points. Experimental results show that our method is 20 times faster than the traditional Gabor features based method while keeping an almost equal accuracy.
Shuye Zhang, Dapeng Tao, Zhao Yang 0001
SMC3
2013 Color-to-gray based on chance of happening preservation
Mingli Song, Dapeng Tao, Chun Chen 0001, Jiajun Bu, Yezhou Yang
Neurocomputing2
2013 Skeleton correspondence construction and its applications in animation style reusing
Zhijun Song, Jun Yu 0002, Changle Zhou, Dapeng Tao
Neurocomputing4
2013 High-level attributes modeling for indoor scenes classification
Chaojie Wang 0003, Jun Yu 0002, Dapeng Tao
Neurocomputing3
2013 Automatic local exposure correction using bright channel prior for under-exposed images
Yinting Wang, Shaojie Zhuo, Dapeng Tao, Jiajun Bu
Signal Process.3
2013 Person Re-Identification by Regularized Smoothing KISS Metric Learning
abstract
With the rapid development of the intelligent video surveillance (IVS), person re-identification, which is a difficult yet unavoidable problem in video surveillance, has received increasing attention in recent years. That is because computer capacity has shown remarkable progress and the task of person re-identification plays a critical role in video surveillance systems. In short, person re-identification aims to find an individual again that has been observed over different cameras. It has been reported that KISS metric learning has obtained the state of the art performance for person re-identification on the VIPeR dataset. However, given a small size training set, the estimation to the inverse of a covariance matrix is not stable and thus the resulting performance can be poor. In this paper, we present regularized smoothing KISS metric learning (RS-KISS) by seamlessly integrating smoothing and regularization techniques for robustly estimating covariance matrices. RS-KISS is superior to KISS, because RS-KISS can enlarge the underestimated small eigenvalues and can reduce the overestimated large eigenvalues of the estimated covariance matrix in an effective way. By providing additional data, we can obtain a more robust model by RS-KISS. However, retraining RS-KISS on all the available examples in a straightforward way is time consuming, so we introduce incremental learning to RS-KISS. We thoroughly conduct experiments on the VIPeR dataset and verify that 1) RS-KISS completely beats all available results for person re-identification and 2) incremental RS-KISS performs as well as RS-KISS but reduces the computational cost significantly.
Dapeng Tao, Yongfei Wang, Yuan Yuan 0001, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2013 Rank Preserving Sparse Learning for Kinect Based Scene Classification
abstract
With the rapid development of the RGB-D sensors and the promptly growing population of the low-cost Microsoft Kinect sensor, scene classification, which is a hard, yet important, problem in computer vision, has gained a resurgence of interest recently. That is because the depth of information provided by the Kinect sensor opens an effective and innovative way for scene classification. In this paper, we propose a new scheme for scene classification, which applies locality-constrained linear coding (LLC) to local SIFT features for representing the RGB-D samples and classifies scenes through the cooperation between a new rank preserving sparse learning (RPSL) based dimension reduction and a simple classification method. RPSL considers four aspects: 1) it preserves the rank order information of the within-class samples in a local patch; 2) it maximizes the margin between the between-class samples on the local patch; 3) the L1-norm penalty is introduced to obtain the parsimony property; and 4) it models the classification error minimization by utilizing the least-squares error minimization. Experiments are conducted on the NYU Depth V1 dataset and demonstrate the robustness and effectiveness of RPSL for scene classification.
Dapeng Tao, Zhao Yang 0001, Xuelong Li 0001
IEEE Trans. Cybern.1
2013 Hessian Regularized Support Vector Machines for Mobile Image Annotation on the Cloud
abstract
With the rapid development of the cloud computing and mobile service, users expect a better experience through multimedia computing, such as automatic or semi-automatic personal image and video organization and intelligent user interface. These functions heavily depend on the success of image understanding, and thus large-scale image annotation has received intensive attention in recent years. The collaboration between mobile and cloud opens a new avenue for image annotation, because the heavy computation can be transferred to the cloud for immediately responding user actions. In this paper, we present a scheme for image annotation on the cloud, which transmits mobile images compressed by Hamming compressed sensing to the cloud and conducts semantic annotation through a novel Hessian regularized support vector machine on the cloud. We carefully explained the rationality of Hessian regularization for encoding the local geometry of the compact support of the marginal distribution and proved that Hessian regularized support vector machine in the reproducing kernel Hilbert space is equivalent to conduct Hessian regularized support vector machine in the space spanned by the principal components of the kernel principal component analysis. We conducted experiments on the PASCAL VOC'07 dataset and demonstrated the effectiveness of Hessian regularized support vector machine for large-scale image annotation.
Dapeng Tao, Weifeng Liu 0001, Xuelong Li 0001
IEEE Trans. Multim.1
2012 Discriminative information preservation for face recognition
Dapeng Tao
Neurocomputing1
2011 Similar Handwritten Chinese Character Recognition Using Discriminative Locality Alignment Manifold Learning
abstract
The discriminant analysis for Similar Handwritten Chinese Character Recognition (SHCR) is essential for the improvement of handwritten Chinese character recognition performance. In this paper, a new manifold based subspace learning algorithm, Discriminative Locality Alignment (DLA), is introduced into SHCR. Experimental results demonstrate that DLA is consistently superior to LDA (Linear Discriminant Analysis) in terms of discriminate information extraction, dimension reduction and recognition accuracy. In addition, DLA reveals some attractive intrinsic properties for numeric calculation, e.g. it can overcome the matrix singular problem and small sample size problem in SHCR.
Dapeng Tao, Lingyu Liang, Yan Gao 0011
ICDAR1