Zhenmin Tang

dblp:13/6728 · also Zhen-Min Tang, Zhenming Tang · DBLP profile ↗
← Back
88ranked-venue papers
0as first author
28since 2021 · last 2025
0000-0001-6708-2205ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 53 · 19 since 2021Artificial intelligence and machine learning · 36 · 11 since 2021Databases, data management, data science and information retrieval · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Security and privacy · 1
YearPublicationVenuePosition
2025 Semi-Paired Semi-Supervised Deep Hashing for cross-view retrieval
abstract
Abstract Due to its fast computational speed and low storage cost, hashing has been effectively applied to large‐scale multimedia retrieval tasks, such as medical video and security video retrieval. Most existing cross‐view hashing methods require good matching information, however, this exact pairing relationship is difficult to fully realise in practice. The association between views is incomplete, as is the label information. This task of missing paired and labelled information is very challenging, but less explored in research. In this study, a semi‐supervised semi‐paired deep hashing for large‐scale data is proposed, named Semi‐Paired Semi‐Supervised Deep Hashing (SPSDH) to solve this challenging task. SPSDH is a novel end‐to‐end deep neural network model with high‐order affinity. A non‐local higher‐order affinity measure that better considers the multimodal neighbourhood structure is proposed. A common representation to associate different modalities is introduced, which combined with the labelled information greatly maintains the consistency within the modalities. SPSDH is evaluated on three benchmark datasets for large‐scale cross‐view approximate nearest neighbour search and compared with several state‐of‐the‐art hashing methods. Extensive experimental results demonstrate the superior performance of our proposed SPSDH in semi‐supervised semi‐paired retrieval tasks.
Yi Wang 0104, Xiaobo Shen 0001, Zhenmin Tang, Ming Zhang 0033
IET Comput. Vis.3
2025 A spatio-frequency cross fusion model for deepfake detection and segmentation
Junshuai Zheng, Ning Zhang 0033, Xiyuan Hu, Kaiwen Xu, Dongyang Gao, Zhenmin Tang
Neurocomputing7
2024 Deepfake Detection With Combined Unsupervised-Supervised Contrastive Learning
abstract
The malicious dissemination of fake images has caused a societal trust crisis, deepfake detection becomes a hot topic now. Through existing detection methods achieve good results in intra-dataset, their performance are poor for unknown manipulations or datasets. To deal with this problem, this paper proposes a new deepfake detection model with combined unsupervised-supervised contrastive learning. By combining unsupervised contrastive learning and supervised contrastive learning with deepfake detection together, the model can discover the essence of fake images from both individual and class features. In addition, a multi-scale attention fusion module is proposed, which helps to enhance the model stability by fusion global and local features of the image. Finally, lots of experiments prove that our method has good performance and generalization ability in intra-dataset, cross-dataset and cross-manipulation scenarios.
Junshuai Zheng, Xiyuan Hu, Zhenmin Tang
ICIP4
2024 Delving Deeper Into Clean Samples for Combating Noisy Labels
Yiyou Gao, Zeren Sun, Yazhou Yao, Xiruo Jiang, Zhenmin Tang
PRCV (9)5
2024 A new deepfake detection model for responding to perception attacks in embodied artificial intelligence
Junshuai Zheng, Xiyuan Hu, Chen Chen 0036, Dongyang Gao, Zhenmin Tang
Image Vis. Comput.6
2023 DT-TransUNet: A Dual-Task Model for Deepfake Detection and Segmentation
Junshuai Zheng, Xiyuan Hu, Zhenmin Tang
PRCV (7)4
2023 Rapid Person Re-Identification via Sub-space Consistency Regularization
Qingze Yin, Guan'an Wang, Guodong Ding, Qilei Li, Shaogang Gong, Zhenmin Tang
Neural Process. Lett.6
2023 Depth and Video Segmentation Based Visual Attention for Embodied Question Answering
abstract
Embodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real-world environment. It has attracted increasing research interests due to its broad applications in personal assistants and in-home robots. Most of the existing methods perform poorly in terms of answering and navigation accuracy due to the absence of fine-level semantic information, stability to the ambiguity, and 3D spatial information of the virtual environment. To tackle these problems, we propose a depth and segmentation based visual attention mechanism for Embodied Question Answering. First, we extract local semantic features by introducing a novel high-speed video segmentation framework. Then guided by the extracted semantic features, a depth and segmentation based visual attention mechanism is proposed for the Visual Question Answering (VQA) sub-task. Further, a feature fusion strategy is designed to guide the navigator's training process without much additional computational cost. The ablation experiments show that our method effectively boosts the performance of the VQA module and navigation module, leading to 4.9 % and 5.6 % overall improvement in EQA accuracy on House3D and Matterport3D datasets respectively.
Haonan Luo 0002, Guosheng Lin, Yazhou Yao, Fayao Liu, Zichuan Liu, Zhenmin Tang
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Robust learning from noisy web data for fine-Grained recognition
Zhenhuang Cai, Guosen Xie, Xingguo Huang, Yazhou Yao, Zhenmin Tang
Pattern Recognit.6
2023 Co-mining: Mining informative samples with noisy labels
Zhenhuang Cai, Huafeng Liu 0004, Yazhou Yao, Zhenmin Tang
Signal Process.5
2023 Guided by Meta-Set: A Data-Driven Method for Fine-Grained Visual Recognition
abstract
The lack of sufficient training data has been one obstacle to fine-grained visual classification research because labeling subcategories generally requires specialist knowledge. As one optional approach to alleviating the data-hunger problem, leveraging web images as training data is drawing increasing attention. Nevertheless, web images potentially have false labels, which can misguide the training process. Although several works have been proposed to deal with label noise, it still can be difficult for the network to tackle complex real-world noisy labels without any prior knowledge. In the literature, we propose to leverage a small and clean meta-set to provide reliable prior knowledge for tackling noisy web images. Specifically, our method trains a network with two peer predicting heads, which learn from noisy web images (web head) and meta ones (meta head), respectively. The meta head produces pseudo soft labels for web images to revise their training loss, which can overcome the high noise ratio problem. Furthermore, a selection net is trained in a meta-learning strategy to identify in- and out-of-distribution noisy images. Then in-distribution ones are reused for training with pseudo soft labels produced by the meta head as supervision, while out-of-distribution ones are discarded. In this manner, the misguidance caused by label noise is remarkably alleviated and in-distribution noisy samples are properly exploited to boost model performance. The superiority of our proposed approach is demonstrated by mathematical theory with great interpretability as well as extensive experimental results on the real-world dataset WebFG-496.
Chuanyi Zhang, Guosheng Lin, Qiong Wang 0003, Fumin Shen, Yazhou Yao, Zhenmin Tang
IEEE Trans. Multim.6
2022 Hierarchical Feature Alignment Network for Unsupervised Video Object Segmentation
Gensheng Pei, Fumin Shen, Yazhou Yao, Guosen Xie, Zhenmin Tang, Jinhui Tang 0001
ECCV (34)5
2022 A Novel Multi-Sample Data Augmentation Method for Oriented Object Detection in Remote Sensing Images
abstract
Data augmentation is widely used in computer vision tasks for enhancing the diversity of training data. However, due to sample redundancy and lack of object background, it is challenging to apply traditional data augmentation techniques to oriented object detection in remote sensing images. In this work, we propose SSMup (specifically synthetic mineral oversampling with mosaic and mixup), a multi-sample data augmentation method, to improve object detection performance in remote sensing images. Our method integrates Mosaic, Mixup, and SSMOTE to enable even distribution of target objects in augmented samples. Moreover, it equips the augmented samples with rich background information. Compared to existing state-of-the-art methods, our proposed method can remarkably improve the detection and generalization performance in remote sensing images. Comprehensive experiments are provided to demonstrate the effectiveness of the proposed method.
Guhua Chen, Gensheng Pei, Tao Chen 0012, Zhenmin Tang
MMSP5
2022 Dense Semantics-Assisted Networks for Video Action Recognition
abstract
Most existing action recognition approaches directly leverage the video-level features to recognize human actions from videos. Although these methods have made remarkable progress, the accuracy is still unsatisfied. When the test video involves complex backgrounds and activities, existing methods usually suffer from a significant drop in accuracy. Human action is inherently a high-level concept. Merely applying a video classification model without a detailed semantic understanding of the video content, e.g., objects, scene context, object motions, object interactions, is inadequate to tackle the challenges for action recognition. Fine-level semantic understanding of videos generates elementary semantic concepts from the raw video data, such as the semantics of objects and background regions. It can be employed to bridge the gap between the raw video data and the high-level concept of human actions. In this work, we leverage dense semantic segmentation masks, which encode rich semantic details, provide extra information for the network training, and improve the performance of action recognition. We propose a novel deep architecture which is named as Dense Semantics-Assisted Convolutional Neural Networks (DSA-CNNs) to effectively utilize dense semantic information of video by a bottom-up attention way in the spatial stream, while by the way of branch fusion in the temporal stream. To verify the effectiveness of our approach, we conduct extensive experiments on publicly available datasets – UCF101, HMDB51, and Kinetics. The experimental results demonstrate that our approach substantially improves existing methods and achieves very competitive performance. It also shows that our approach is superior to other related methods that utilize extra information for action recognition.
Haonan Luo 0002, Guosheng Lin, Yazhou Yao, Zhenmin Tang, Qingyao Wu, Xian-Sheng Hua 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Enhanced Feature Alignment for Unsupervised Domain Adaptation of Semantic Segmentation
abstract
Unsupervised domain adaptation for semantic segmentation aims to transfer knowledge from a labeled source domain to another unlabeled target domain. However, due to the label noise and domain mismatch, learning directly from source domain data tends to have poor performance. Though adversarial learning methods strive to reduce domain discrepancies by aligning feature distributions, traditional methods suffer from the training imbalance and feature distortion problems. Besides, due to the absence of target domain labels, the classifier is blind to features from the target domain during training. Consequently, the final classifier overfits the source domain features and usually fails to predict the structured outputs of the target domain. To alleviate these problems, we focus on enhancing the adversarial learning based feature alignment from three perspectives. First, a classification constrained discriminator is proposed to balance the adversarial training and alleviate the feature distortion problem. Next, to alleviate the classifier overfitting problem, self-training is collaboratively used to learn a domain robust classifier with target domain pseudo labels. Moreover, an efficient class centroid calculation module is proposed and the domain discrepancy is further reduced by aligning the feature centroids of the same class from different domains. Experimental evaluations on GTA5$\rightarrow$Cityscapes and SYNTHIA$\rightarrow$Cityscapes demonstrate state-of-the-art results compared to other counterpart methods. The source code and models have been made available at.11[Online]. Available:https://github.com/NUST-Machine-Intelligence-Laboratory/EFA.
Tao Chen 0012, Shuihua Wang, Qiong Wang 0003, Zheng Zhang 0006, Guosen Xie, Zhenmin Tang
IEEE Trans. Multim.6
2022 Semantically Meaningful Class Prototype Learning for One-Shot Image Segmentation
abstract
One-shot semantic image segmentation aims to segment the object regions for the novel class with only one annotated image. Recent works adopt the episodic training strategy to mimic the expected situation at testing time. However, these existing approaches simulate the test conditions too strictly during the training process, and thus cannot make full use of the given label information. Besides, these approaches mainly focus on the foreground-background target class segmentation setting. They only utilize binary mask labels for training. In this paper, we propose to leverage the multi-class label information during the episodic training. It will encourage the network to generate more semantically meaningful features for each category. After integrating the target class cues into the query features, we then propose a pyramid feature fusion module to mine the fused features for the final classifier. Furthermore, to take more advantage of the support image-mask pair, we propose a self-prototype guidance branch to support image segmentation. It can constrain the network for generating more compact features and a robust prototype for each semantic class. For inference, we propose a fused prototype guidance branch for the segmentation of the query image. Specifically, we leverage the prediction of the query image to extract the pseudo-prototype and combine it with the initial prototype. Then we utilize the fused prototype to guide the final segmentation of the query image. Extensive experiments demonstrate the superiority of our proposed approach. The source codes and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/SMCP.
Tao Chen 0012, Guosen Xie, Yazhou Yao, Qiong Wang 0003, Fumin Shen, Zhenmin Tang, Jian Zhang 0002
IEEE Trans. Multim.6
2022 Exploiting Web Images for Fine-Grained Visual Recognition via Dynamic Loss Correction and Global Sample Selection
abstract
To distinguish subtle differences among fine-grained categories, a large amount of well-labeled images are typically required. However, acquiring manual annotations for fine-grained categories is an extremely difficult task as it usually has a high demand for professional knowledge. To this end, directly leveraging web images for learning fine-grained models becomes a natural choice. Nevertheless, due to the existence of label noise, this learning paradigm tends to have a poor performance. In this work, we propose an end-to-end approach by combining dynamic loss correction and global sample selection to alleviate the problem of label noise. Specifically, we leverage the network to predict all samples, record the predictions of recent several epochs, and calculate the uncertainly-based dynamic loss for global sample selection. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed approach. The source code of our approach has been released on the website:https://github.com/NUST-Machine-Intelligence-Laboratory/dlc.
Huafeng Liu 0004, Haofeng Zhang 0001, Jianfeng Lu 0003, Zhenmin Tang
IEEE Trans. Multim.4
2022 Exploiting Web Images for Fine-Grained Visual Recognition by Eliminating Open-Set Noise and Utilizing Hard Examples
abstract
Labeling objects at a subordinate level typically requires expert knowledge, which is not always available when using random annotators. As such, learning directly from web images for fine-grained recognition has attracted broad attention. However, the presence of label noise and hard examples in web images are two obstacles for training robust fine-grained recognition models. Therefore, in this paper, we propose a novel approach for removing irrelevant samples from real-world web images during training, while employing useful hard examples to update the network. Thus, our approach can alleviate the harmful effects of irrelevant noisy web images and hard examples to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is far superior to current state-of-the-art web-supervised methods. The data and source code of this work have been made publicly available at:https://github.com/NUST-Machine-Intelligence-Laboratory/Advanced-Softly-Update-Drop.
Huafeng Liu 0004, Chuanyi Zhang, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Zhenmin Tang, Jian Zhang 0002
IEEE Trans. Multim.6
2022 Co-LDL: A Co-Training-Based Label Distribution Learning Method for Tackling Label Noise
abstract
Performances of deep neural networks are prone to be degraded by label noise due to their powerful capability in fitting training data. Deeming low-loss instances as clean data is one of the most promising strategies in tackling label noise and has been widely adopted by state-of-the-art methods. However, prior works tend to drop high-loss instances directly, neglecting their valuable information. To address this issue, we propose an end-to-end framework named Co-LDL, which incorporates the low-loss sample selection strategy with label distribution learning. Specifically, we simultaneously train two deep neural networks and let them communicate useful knowledge by selecting low-loss and high-loss samples for each other. Low-loss samples are leveraged conventionally for updating network parameters. On the contrary, high-loss samples are trained in a label distribution learning manner to update network parameters and label distributions concurrently. Moreover, we propose a self-supervised module to further boost the model performance by enhancing the learned representations. Comprehensive experiments on both synthetic and real-world noisy datasets are provided to demonstrate the superiority of our Co-LDL method over state-of-the-art approaches in learning with noisy labels. The source code and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/CoLDL.
Zeren Sun, Huafeng Liu 0004, Qiong Wang 0003, Tianfei Zhou, Qi Wu 0001, Zhenmin Tang
IEEE Trans. Multim.6
2022 Robust Learning From Noisy Web Images Via Data Purification for Fine-Grained Recognition
abstract
Manually labeling fine-grained datasetsis laborious and typically requires domain-specific expert knowledge. Conversely, a vast amount of web data is relatively easy to obtain with nearly no human effort. Therefore, learning from noisy web data for fine-grained tasks is attracting increasing attention in recent years. However, the presence of noise in web images is a huge obstacle for training robust fine-grained recognition models. To this end, we propose a novel approach to identify noisy images as well as specifically distinguish in- and out-of-distribution samples. It can purify the noisy web training set by discarding out-of-distribution noise and relabeling in-distribution noisy samples. Then we can train the model on the purified dataset to alleviate the harmful effects of noise and make the most of web images to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is far superior to current state-of-the-art web-supervised methods. The data and source code of this work have been made publicly available at:https://github.com/NUST-Machine-Intelligence-Laboratory/Dataset-Purification.
Chuanyi Zhang, Qiong Wang 0003, Guosen Xie, Qi Wu 0001, Fumin Shen, Zhenmin Tang
IEEE Trans. Multim.6
2021 Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of segmentation. However, existing works mainly concentrate on expanding the seed of pseudo labels within the image’s salient region. In this work, we propose a non-salient region object mining approach for weakly supervised semantic segmentation. We introduce a graph-based global reasoning unit to strengthen the classification network’s ability to capture global relations among disjoint and distant regions. This helps the network activate the object features outside the salient area. To further mine the non-salient region objects, we propose to exert the segmentation network’s self-correction ability. Specifically, a potential object mining module is proposed to reduce the false-negative rate in pseudo labels. Moreover, we propose a non-salient region masking module for complex images to generate masked pseudo labels. Our non-salient region masking module helps further discover the objects in the non-salient region. Extensive experiments on the PASCAL VOC dataset demonstrate state-of-the-art results compared to current methods. The source codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/nsrom.
Yazhou Yao, Tao Chen 0012, Guosen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu 0001, Zhenmin Tang, Jian Zhang 0002
CVPR7
2021 Jo-SRC: A Contrastive Approach for Combating Noisy Labels
abstract
Due to the memorization effect in Deep Neural Networks (DNNs), training with noisy labels usually results in inferior model performance. Existing state-of-the-art methods primarily adopt a sample selection strategy, which selects small-loss samples for subsequent training. However, prior literature tends to perform sample selection within each mini-batch, neglecting the imbalance of noise ratios in different mini-batches. Moreover, valuable knowledge within high-loss samples is wasted. To this end, we propose a noise-robust approach named Jo-SRC (Joint Sample Selection and Model Regularization based on Consistency). Specifically, we train the network in a contrastive learning manner. Predictions from two different views of each sample are used to estimate its "likelihood" of being clean or out-of-distribution. Furthermore, we propose a joint loss to advance the model generalization performance by introducing consistency regularization. Extensive experiments have validated the superiority of our approach over existing state-of-the-art methods. The source code and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/Jo-SRC.
Yazhou Yao, Zeren Sun, Chuanyi Zhang, Fumin Shen, Qi Wu 0001, Jian Zhang 0002, Zhenmin Tang
CVPR7
2021 Extracting Useful Knowledge from Noisy Web Images via Data Purification for Fine-Grained Recognition
abstract
Fine-grained visual recognition tasks typically require training data with reliable acquisition and annotation processes. Acquiring such datasets with precise fine-grained annotations is very expensive and time-consuming. Conversely, a vast amount of web data is relatively easy to obtain with nearly no human effort. Nevertheless, the presence of label noise in web images becomes a huge obstacle for training robust fine-grained recognition models. In this work, we investigate the noisy label problem and propose a method that can specifically distinguish in- and out-of-distribution noisy samples. It can purify the web training data by discarding out-of-distribution noisy images and relabeling in-distribution ones. After purification, we can train the model on a less noisy web training set to achieve better robustness and performance. Extensive experiments on three real-world web datasets for fine-grained visual recognition demonstrate the superiority of our approach.
Chuanyi Zhang, Yazhou Yao, Xing Xu 0001, Jie Shao 0001, Jingkuan Song, Zechao Li, Zhenmin Tang
ACM Multimedia7
2021 Local Self-Attention on Fine-grained Cross-media Retrieval
abstract
Due to the heterogeneity gap, the data representation of different media is inconsistent and belongs to different feature spaces. Therefore, it is challenging to measure the fine-grained gap between them. To this end, we propose an attention space training method to learn common representations of different media data. Specifically, we utilize local self-attention layers to learn the common attention space between different media data. We propose a similarity concatenation method to understand the content relationship between features. To further improve the robustness of the model, we also train a local position encoding to capture the spatial relationships between features. In this way, our proposed method can effectively reduce the gap between different feature distributions on cross-media retrieval tasks. It also improves the fine-grained recognition performance by attaching attention to high-level semantic information. Extensive experiments and ablation studies demonstrate that our proposed method achieves state-of-the-art performance. At the same time, our approach provides a new pipeline for fine-grained cross-media retrieval. The source code and models are publicly available at: https://github.com/NUST-Machine-Intelligence-Laboratory/SAFGCMHN.
Yazhou Yao, Qiong Wang 0003, Zhenmin Tang
MMAsia4
2021 Information-theoretic measures of uncertainty for interval-set decision tables
Xiuyi Jia, Zhenmin Tang
Inf. Sci.3
2021 Exploiting textual queries for dynamically visual disambiguation
abstract
Due to the high cost of manual annotation, learning directly from the web has attracted broad attention. One issue that limits the performance of current webly supervised models is the problem of visual polysemy. In this work, we present a novel framework that resolves visual polysemy by dynamically matching candidate text queries with retrieved images. Specifically, our proposed framework includes three major steps: we first discover and then dynamically select the text queries according to the keyword-based image search results, we employ the proposed saliency-guided deep multi-instance learning (MIL) network to remove outliers and learn classification models for visual disambiguation. Compared to existing methods, our proposed approach can figure out the right visual senses, adapt to dynamic changes in the search results, remove outliers, and jointly learn the classification models . Extensive experiments and ablation studies on CMU-Poly-30 and MIT-ISD datasets demonstrate the effectiveness of our proposed approach.
Zeren Sun, Yazhou Yao, Jimin Xiao, Lei Zhang 0054, Jian Zhang 0002, Zhenmin Tang
Pattern Recognit.6
2021 Robust gait recognition using hybrid descriptors based on Skeleton Gait Energy Image
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang, Wankou Yang
Pattern Recognit. Lett.5
2021 Multi-View Label Prediction for Unsupervised Learning Person Re-Identification
abstract
Person re-identification (ReID) aims to match pedestrian images across disjoint cameras. Existing supervised ReID methods utilize deep networks and train them with identity-labeled images, which suffer from limited annotations. Recently, clustering-based unsupervised ReID attracts more and more attention. It first clusters unlabeled images and assigns cluster index to the pseudo-identity-labels, then trains a ReID model with the pseudo-identity-labels. However, considering the slight inter-class variations and significant intra-class variations, pseudo-identity-labels learned from clustering algorithms are usually noisy and coarse. To alleviate the problems above, besides clustering pseudo-identity-labels, we propose to learn pseudo-patch-labels, which brings two advantages: (1) Patch naturally alleviates the effect of backgrounds, occlusions, and carryings since they usually occupy small parts in images, thus overcome noisy labels. (2) It is plausible that patches from different pedestrians belong to the same pseudo-identity-label. For example, pedestrians have a high probability of wearing either the same shoes or pants but a low possibility of wearing both. The experiments demonstrate our proposed method achieves the best performance by a large margin on both image- and video-based datasets.
Qingze Yin, Guan'an Wang, Guodong Ding, Shaogang Gong, Zhenmin Tang
IEEE Signal Process. Lett.5
2020 Web-Supervised Network with Softly Update-Drop Training for Fine-Grained Visual Classification
abstract
Labeling objects at the subordinate level typically requires expert knowledge, which is not always available from a random annotator. Accordingly, learning directly from web images for fine-grained visual classification (FGVC) has attracted broad attention. However, the existence of noise in web images is a huge obstacle for training robust deep neural networks. In this paper, we propose a novel approach to remove irrelevant samples from the real-world web images during training, and only utilize useful images for updating the networks. Thus, our network can alleviate the harmful effects caused by irrelevant noisy web images to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art webly supervised methods. The data and source code of this work have been made anonymously available at: https://github.com/z337-408/WSNFGVC.
Chuanyi Zhang, Yazhou Yao, Huafeng Liu 0004, Guosen Xie, Xiangbo Shu, Tianfei Zhou, Zheng Zhang 0006, Fumin Shen, Zhenmin Tang
AAAI9
2020 Classification Constrained Discriminator For Domain Adaptive Semantic Segmentation
abstract
Unsupervised domain adaptation for semantic segmentation aims to transfer knowledge from label-rich synthetic datasets to real-world images without any annotation. The traditional adversarial learning methods for domain adaptation learn to extract domain-invariant feature representations by aligning the feature distributions of both domains. However, these methods suffer from an imbalance in adversarial training and feature distortion. In this work, we propose a classification constrained discriminator to alleviate these problems. Specifically, we first propose to balance the adversarial training by eliminating any pooling layers or strided convolutions in the discriminator. Then, we propose to constrain the discriminator with an auxiliary classification loss to help the feature generator extract the domain-invariant features that are useful for segmentation rather than just ambiguous features to fool the domain discriminator. Extensive experiments demonstrate the superiority of our proposed approach. The source code and models have been made available at https://github.com/NUSTMachine-Intelligence-Laboratory/ccd.
Tao Chen 0012, Jian Zhang 0002, Guosen Xie, Yazhou Yao, Xiaoshui Huang, Zhenmin Tang
ICME6
2020 Hsi Road: A Hyper Spectral Image Dataset For Road Segmentation
abstract
Road segmentation is a challenging task in the field of self-driving research. This paper present a road dataset built by hyper spectral imaging (HSI) cameras instead of the widely-used RGB cameras. HSI image is informative in spectrums and full of potential for natural environment perception. In this article, a first-of-its-kind HSI road segmentation dataset is built with careful annotation in both urban and rural scenes. It contains 3799 scenes with RGB and NIR bands as well as their respective masks. Unlike many existing datasets that provide urban scenes in RGB images only, our dataset expands the sensing spectrum to 28 bands and includes various kinds of road surfaces, such as asphalt, cement, dirt and sand, under rural and natural scenes. We also provide benchmark performances based on the recently popular segmentation algorithms on this dataset. The dataset is released at github‡.‡https://github.com/NUST-Machine-Intelligence-Laboratory/hsi_road
Jiarou Lu, Huafeng Liu 0004, Yazhou Yao, Shuyin Tao, Zhenmin Tang, Jianfeng Lu 0003
ICME5
2020 Web-Supervised Network for Fine-Grained Visual Classification
abstract
Fine-grained visual classification (FGVC) is a tough task due to its high annotation cost of the fine-grained subcategories. To build a large-scale dataset at low manual cost, straightforwardly learning from web images for FGVC has attracted broad attention. However, there exist two characteristics in the need of concerning for the web dataset: 1) Noisy images; 2) A large proportion of hard examples. In this paper, we propose a simple yet effective approach to deal with noisy images and hard examples during training. Our method is a pure web-supervised method for FGVC. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to the state-of-the-art web-supervised methods. The data and source code of this work have been posted available at: https://github.com/NUST-Machine-Intelligence-Laboratory/WSNFG.
Chuanyi Zhang, Yazhou Yao, Jiachao Zhang, Jian Zhang 0002, Zhenmin Tang
ICME7
2020 Data-driven Meta-set Based Fine-Grained Visual Recognition
abstract
Constructing fine-grained image datasets typically requires domain-specific expert knowledge, which is not always available for crowd-sourcing platform annotators. Accordingly, learning directly from web images becomes an alternative method for fine-grained visual recognition. However, label noise in the web training set can severely degrade the model performance. To this end, we propose a data-driven meta-set based approach to deal with noisy web images for fine-grained recognition. Specifically, guided by a small amount of clean meta-set, we train a selection net in a meta-learning manner to distinguish in- and out-of-distribution noisy images. To further boost the robustness of the model, we also learn a labeling net to correct the labels of in-distribution noisy data. In this way, our proposed method can alleviate the harmful effects caused by out-of-distribution noise and properly exploit the in-distribution noisy samples for training. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art noise-robust methods.
Chuanyi Zhang, Yazhou Yao, Xiangbo Shu, Zechao Li, Zhenmin Tang, Qi Wu 0001
ACM Multimedia5
2020 Road segmentation with image-LiDAR data fusion in deep neural network
Huafeng Liu 0004, Yazhou Yao, Zeren Sun, Xiangrui Li, Ke Jia, Zhenmin Tang
Multim. Tools Appl.6
2020 Feature mask network for person re-identification
Guodong Ding, Salman Khan 0001, Zhenmin Tang, Fatih Porikli
Pattern Recognit. Lett.3
2020 True-Color and Grayscale Video Person Re-Identification
abstract
Person re-identification is an important task in forensics applications. Most existing person re-identification methods focus on matching persons captured by different true-color cameras. In practice, the captured pedestrian videos may be grayscale in some cases due to camera malfunction or special treatment for gray mode. In these cases, the person re-identification between true-color and grayscale pedestrian videos, which we call color to gray video person re-identification (CGVPR), will be needed. Since the color information that is very important to represent a pedestrian is usually intensity information and monochrome in grayscale videos, the CGVPR problem is very challenging. To relieve the difficulties in CGVPR, we propose an asymmetric within-video projection based Semicoupled Dictionary Pair Learning (SDPL) approach. SDPL simultaneously learns two within-video projection matrices, a pair of true-color and grayscale dictionaries, as well as a semi-coupled mapping matrix. The learned within-video projection matrices can make each video (true-color or grayscale) more compact. The learned dictionary pair and the mapping matrix can work together to bridge the gap between features of true-color and grayscale videos. To date there exists no true-color and grayscale pedestrian video dataset, so we contribute a new one, called true-color and grayscale video person re-identification dataset (CGVID). Our dataset is collected under a real-world scenario and consists of over 50K frames. Extensive evaluations demonstrate that the collected CGVID dataset is very challenging and can be used for further research on person re-identification. The experimental results show that our approach outperforms the compared methods on the CGVPR task.
Fei Ma 0004, Xiaoyuan Jing, Zhenmin Tang, Zhiping Peng
IEEE Trans. Inf. Forensics Secur.4
2019 Dispersion based Clustering for Unsupervised Person Re-identification
Guodong Ding, Salman Khan 0001, Zhenmin Tang
BMVC3
2019 SegEQA: Video Segmentation Based Visual Attention for Embodied Question Answering
abstract
Embodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real world environment. It has attracted increasing research interests due to its broad applications in automatic driving system, in-home robots, and personal assistants. Most of the existing methods perform poorly in terms of answering and navigation accuracy due to the absence of local details and vulnerability to the ambiguity caused by complicated vision conditions. To tackle these problems, we propose a segmentation based visual attention mechanism for Embodied Question Answering. Firstly, We extract the local semantic features by introducing a novel high-speed video segmentation framework. Then by the guide of extracted semantic features, a bottom-up visual attention mechanism is proposed for the Visual Question Answering (VQA) sub-task. Further, a feature fusion strategy is proposed to guide the training of the navigator without much additional computational cost. The ablation experiments show that our method boosts the performance of VQA module by 4.2% (68.99% vs 64.73%) and leads to 3.6% (48.59% vs 44.98%) overall improvement in EQA accuracy.
Haonan Luo 0002, Guosheng Lin, Zichuan Liu, Fayao Liu, Zhenmin Tang, Yazhou Yao
ICCV5
2019 Uncertainty measures for interval set information tables based on interval δ-similarity relation
Xiuyi Jia, Zhenmin Tang, Xianzhong Long
Inf. Sci.3
2019 Deep representation learning for road detection using Siamese network
Huafeng Liu 0004, Xiangrui Li, Yazhou Yao, Zhenmin Tang
Multim. Tools Appl.6
2019 Statistical performance of convex low-rank and sparse tensor recovery
Xiangrui Li, Andong Wang, Jianfeng Lu 0003, Zhenmin Tang
Pattern Recognit.4
2019 Exploiting textual and visual features for image categorization
Yazhou Yao, Wankou Yang, Qiong Wang 0003, Yunfei Cai, Zhenmin Tang
Pattern Recognit. Lett.6
2019 Extracting Privileged Information for Enhancing Classifier Learning
abstract
The accuracy of data-driven learning approaches is often unsatisfactory when the training data is inadequate either in quantity or quality. Manually labeled privileged information (PI), e.g., attributes, tags or properties, is usually incorporated to improve classifier learning. However, the process of manually labeling is time-consuming and labor-intensive. Moreover, due to the limitations of personal knowledge, manually labeled PI may not be rich enough. To address these issues, we propose to enhance classifier learning by exploring PI from untagged corpora, which can effectively eliminate the dependency on manually labeled data and obtain much richer PI. In detail, we treat each selected PI as a subcategory and learn one classifier for each subcategory independently. The classifiers for all subcategories are integrated together to form a more powerful category classifier. Particularly, we propose a novel instancelevel multi-instance learning (MIL) model to simultaneously select a subset of training images from each subcategory and learn the optimal SVM classifiers based on the selected images. Extensive experiments on four benchmark datasets demonstrate the superiority of our proposed approach.
Yazhou Yao, Fumin Shen, Jian Zhang 0002, Li Liu 0004, Zhenmin Tang, Ling Shao 0001
IEEE Trans. Image Process.5
2019 Feature Affinity-Based Pseudo Labeling for Semi-Supervised Person Re-Identification
abstract
Vision-based person re-identification aims to match a person's identity across multiple images, which is a fundamental task in multimedia content analysis and retrieval. Deep neural networks have recently manifested great potential in this task. However, a major bottleneck of existing supervised deep networks is their reliance on a large amount of annotated training data. Manual labeling for person identities in large-scale surveillance camera systems is quite challenging and incurs significant costs. Some recent studies adopt generative model outputs as training data augmentation. To more effectively use these synthetic data for an improved feature learning and re-identification performance, this paper proposes a novel feature affinity-based pseudo labeling method with two possible label encodings. To the best of our knowledge, this is the first study that employs pseudo-labeling by measuring the affinity of unlabeled samples with the underlying clusters of labeled data samples using the intermediate feature representations from deep networks. We propose training the network with the joint supervision of cross-entropy loss together with a center regularization term, which not only ensures discriminative feature representation learning but also simultaneously predicts pseudo-labels for unlabeled data. We show that both label encodings can be learned in a unified manner and help improve the overall performance. Our extensive experiments on three person re-identification datasets: Market-1501, DukeMTMC-reID, and CUHK03, demonstrate significant performance boost over the state-of-the-art person re-identification approaches.
Guodong Ding, Shanshan Zhang 0001, Salman Khan 0001, Zhenmin Tang, Jian Zhang 0002, Fatih Porikli
IEEE Trans. Multim.4
2019 Extracting Multiple Visual Senses for Web Learning
abstract
Labeled image datasets have played a critical role in high-level image understanding. However, the process of manual labeling is both time consuming and labor intensive. To reduce the dependence on manually labeled data, there have been increasing research efforts on learning visual classifiers by directly exploiting web images. One issue that limits their performance is the problem of polysemy. Existing unsupervised approaches attempt to reduce the influence of visual polysemy by filtering out irrelevant images, but do not directly address polysemy. To this end, in this paper, we present a multimodal framework that solves the problem of polysemy by allowing sense-specific diversity in search results. Specifically, we first discover a list of possible semantic senses from untagged corpora to retrieve sense-specific images. Then, we merge visual similar semantic senses and prune noise by using the retrieved images. Finally, we train one visual classifier for each selected semantic sense and use the learned sense-specific classifiers to distinguish multiple visual senses. Extensive experiments on classifying images into sense-specific categories and reranking search results demonstrate the superiority of our proposed approach.
Yazhou Yao, Fumin Shen, Jian Zhang 0002, Li Liu 0004, Zhenmin Tang, Ling Shao 0001
IEEE Trans. Multim.5
2018 Discovering and Distinguishing Multiple Visual Senses for Polysemous Words
abstract
To reduce the dependence on labeled data, there have been increasing research efforts on learning visual classifiers by exploiting web images. One issue that limits their performance is the problem of polysemy. To solve this problem, in this work, we present a novel framework that solves the problem of polysemy by allowing sense-specific diversity in search results. Specifically, we first discover a list of possible semantic senses to retrieve sense-specific images. Then we merge visual similar semantic senses and prune noises by using the retrieved images. Finally, we train a visual classifier for each selected semantic sense and use the learned sense-specific classifiers to distinguish multiple visual senses. Extensive experiments on classifying images into sense-specific categories and re-ranking search results demonstrate the superiority of our proposed approach.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Wankou Yang, Zhenmin Tang
AAAI6
2018 Extracting Privileged Information from Untagged Corpora for Classifier Learning
abstract
The performance of data-driven learning approaches is often unsatisfactory when the training data is inadequate either in quantity or quality. Manually labeled privileged information (PI), \eg attributes, tags or properties, is usually incorporated to improve classifier learning. However, the process of manually labeling is time-consuming and labor-intensive. To address this issue, we propose to enhance classifier learning by extracting PI from untagged corpora, which can effectively eliminate the dependency on manually labeled data. In detail, we treat each selected PI as a subcategory and learn one classifier for per subcategory independently. The classifiers for all subcategories are then integrated together to form a more powerful category classifier. Particularly, we propose a new instance-level multi-instance learning (MIL) model to simultaneously select a subset of training images from each subcategory and learn the optimal classifiers based on the selected images. Extensive experiments demonstrate the superiority of our approach.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Wankou Yang, Xian-Sheng Hua 0001, Zhenmin Tang
IJCAI6
2017 A new web-supervised method for image dataset constructions
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
Neurocomputing6
2017 Exploiting Web Images for Dataset Construction: A Domain Robust Approach
abstract
Labeled image datasets have played a critical role in high-level image understanding. However, the process of manual labeling is both time-consuming and labor intensive. To reduce the cost of manual labeling, there has been increased research interest in automatically constructing image datasets by exploiting web images. Datasets constructed by existing methods tend to have a weak domain adaptation ability, which is known as the “dataset bias problem.” To address this issue, we present a novel image dataset construction framework that can be generalized well to unseen target domains. Specifically, the given queries are first expanded by searching the Google Books Ngrams Corpus to obtain a rich semantic description, from which the visually nonsalient and less relevant expansions are filtered out. By treating each selected expansion as a “bag” and the retrieved images as “instances,” image selection can be formulated as a multi-instance learning problem with constrained positive bags. We propose to solve the employed problems by the cutting-plane and concave-convex procedure algorithm. By using this approach, images from different distributions can be kept while noisy images are filtered out. To verify the effectiveness of our proposed approach, we build an image dataset with 20 categories. Extensive experiments on image classification, cross-dataset generalization, diversity comparison, and object detection demonstrate the domain robustness of our dataset.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
IEEE Trans. Multim.6
2016 Automatic image dataset construction with multiple textual metadata
abstract
The goal of this work is to automatically collect a large number of highly relevant images from the Internet for given queries. A novel image dataset construction framework is proposed by employing multiple textual metadata. In specific, the given queries are first expanded by searching in the Google Books Ngrams Corpora to obtain a richer semantic description, from which the visually non-salient and less relevant expansions are then filtered. After retrieving images from the Internet with filtered expansions, we further filter noisy images by clustering and progressively Convolutional Neural Networks (CNN). To verify the effectiveness of our proposed method, we construct a dataset with 10 categories, which is not only much larger than but also have comparable cross-dataset generalization ability with manually labeled dataset STL-10 and CIFAR-10.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
ICME6
2016 A Domain Robust Approach For Image Dataset Construction
abstract
There have been increasing research interests in automatically constructing image dataset by collecting images from the Internet. However, existing methods tend to have a weak domain adaptation ability, known as the "dataset bias problem". To address this issue, in this work, we propose a novel image dataset construction framework which can generalize well to unseen target domains. In specific, the given queries are first expanded by searching in the Google Books Ngrams Corpora (GBNC) to obtain a richer semantic description, from which the noisy query expansions are then filtered out. By treating each expansion as a "bag" and the retrieved images therein as "instances", we formulate image filtering as a multi-instance learning (MIL) problem with constrained positive bags. By this approach, images from different data distributions will be kept while with noisy images filtered out. Comprehensive experiments on two challenging tasks demonstrate the effectiveness of our proposed approach.
Yazhou Yao, Xian-Sheng Hua 0001, Fumin Shen, Jian Zhang 0002, Zhenmin Tang
ACM Multimedia5
2016 Extracting Visual Knowledge from the Internet: Making Sense of Image Data
Yazhou Yao, Jian Zhang 0002, Xian-Sheng Hua 0001, Fumin Shen, Zhenmin Tang
MMM (1)5
2016 Fast total variation deconvolution for blurred image contaminated by Poisson noise
Shuyin Tao, Wende Dong, Zhi-hai Xu, Zhenmin Tang
J. Vis. Commun. Image Represent.4
2015 A novel SRC fusion method using hierarchical multi-scale LBP and greedy search strategy
Zi Liu, Xiaoning Song, Zhenmin Tang
Neurocomputing3
2015 Proportional fair resource allocation based on hybrid ant colony optimization for slow adaptive OFDMA system
Lei Xu 0015, Qianmu Li, Yuwang Yang, Zhenmin Tang, Xiaofei Zhang 0001
Inf. Sci.5
2015 Fusing hierarchical multi-scale local binary patterns and virtual mirror samples to perform face recognition
Zi Liu, Xiaoning Song, Zhenmin Tang
Neural Comput. Appl.3
2015 Hashing on Nonlinear Manifolds
abstract
Learning-based hashing methods have attracted considerable attention due to their ability to greatly increase the scale at which existing algorithms may operate. Most of these methods are designed to generate binary codes preserving the Euclidean similarity in the original space. Manifold learning techniques, in contrast, are better able to model the intrinsic structure embedded in the original high-dimensional data. The complexities of these models, and the problems with out-of-sample data, have previously rendered them unsuitable for application to large-scale embedding, however. In this paper, how to learn compact binary embeddings on their intrinsic manifolds is considered. In order to address the above-mentioned difficulties, an efficient, inductive solution to the out-of-sample data problem, and a process by which nonparametric manifold learning may be used as the basis of a hashing method are proposed. The proposed approach thus allows the development of a range of new hashing techniques exploiting the flexibility of the wide variety of manifold learning approaches available. It is particularly shown that hashing on the basis of t-distributed stochastic neighbor embedding outperforms state-of-the-art hashing methods on large-scale benchmark data sets, and is very effective for image classification with very short code lengths. It is shown that the proposed framework can be further improved, for example, by minimizing the quantization error with learned orthogonal rotations without much computation overhead. In addition, a supervised inductive manifold hashing framework is developed by incorporating the label information, which is shown to greatly advance the semantic retrieval performance.
Fumin Shen, Chunhua Shen, Qinfeng Shi, Anton van den Hengel, Zhenmin Tang, Heng Tao Shen
IEEE Trans. Image Process.5
2015 Local Structure-Based Sparse Representation for Face Recognition
abstract
This article presents a simple yet effective face recognition method, called local structure-based sparse representation classification (LS_SRC). Motivated by the “divide-and-conquer” strategy, we first divide the face into local blocks and classify each local block, then integrate all the classification results to make the final decision. To classify each local block, we further divide each block into several overlapped local patches and assume that these local patches lie in a linear subspace. This subspace assumption reflects the local structure relationship of the overlapped patches, making sparse representation-based classification (SRC) feasible even when encountering the single-sample-per-person (SSPP) problem. To lighten the computing burden of LS_SRC, we further propose the local structure-based collaborative representation classification (LS_CRC). Moreover, the performance of LS_SRC and LS_CRC can be further improved by using the confusion matrix of the classifier. Experimental results on four public face databases show that our methods not only generalize well to SSPP problem but also have strong robustness to occlusion; little pose variation; and the variations of expression, illumination, and time.
Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Liyan Zhang 0001, Zhenmin Tang
ACM Trans. Intell. Syst. Technol.5
2014 Local structure based sparse representation for face recognition with single sample per person
abstract
In this paper, we propose local structure based sparse representation classification (LS SRC) to solve single sample per person (SSPP) problem. By adopting the “divide-conquer-aggregate” strategy, we successfully alleviate the dilemma of high data dimensionality and small samples, where we first divide the face into local blocks, and classify each local block, and then integrate all the classification results by voting. For each block, we further divide it into overlapped patches and assume that these patches lie in a linear subspace. This subspace assumption reflects local structure relationship of the overlapped patches and makes SRC feasible for SSPP problem. To lighten the computing burden, we further propose local structure based collaborative representation classification (LS CRC). Experimental results on three public face databases show that our methods not only generalize well to SSPP problem but also have strong robustness to expression, illumination, little pose variation, occlusion and time variation.
Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Xinguang Xiang, Zhenmin Tang
ICIP5
2014 Proportional fairness resource allocation scheme based on quantised feedback for multiuser orthogonal frequency division multiplexing system
abstract
This work addresses the resource allocation problem with the proportional fair constraint condition based on quantised feedback for multiuser orthogonal frequency division multiplexing access system. The resource allocation problem is converted as an optimisation problem with maximising the lower bound of the total average throughput and this formulation provides the low complexity of solving the above resource allocation problem. Tailored for the above optimisation problem, the authors design the codebook of equivalent channel quantisation threshold and the codebook of power and rate according to the equal probability quantiser and the Lagrange multiplier method, respectively. Further, they develop a suboptimal algorithm based on the stochastic approximate method. The proposed algorithm not only satisfies the constraint condition of the proportional fair very well, but also reduces the feedback overhead of the resource allocation result greatly. Moreover, the average throughput of the proposed algorithm is very close to that of the optimal resource allocation algorithm with full feedback when the equivalent channel gain in every subcarrier is quantised by 4 bit.
Lei Xu 0015, Yuwang Yang, Xiaofei Zhang 0001, Zhenmin Tang, Shaohua Lan
IET Commun.5
2014 On an optimization representation of decision-theoretic rough set model
Xiuyi Jia, Zhenmin Tang, Wenhe Liao, Lin Shang 0001
Int. J. Approx. Reason.2
2014 Discriminant similarity and variance preserving projection for feature extraction
Cai-Kou Chen, Zhenmin Tang, Zhangjing Yang
Neurocomputing3
2014 Feature extraction using local structure preserving discriminant analysis
Cai-Kou Chen, Zhenmin Tang, Zhangjing Yang
Neurocomputing3
2014 Exploiting Universum data in AdaBoost using gradient descent
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang
Image Vis. Comput.4
2014 Local maximal margin discriminant embedding for face recognition
Zhenmin Tang, Cai-Kou Chen, Zhangjing Yang
J. Vis. Commun. Image Represent.2
2014 Boosting Separability in Semisupervised Learning for Object Classification
abstract
Boosting algorithms, especially AdaBoost, have attracted great attention in computer vision. In the early version of boosting algorithms, the weak classifier selection and the strong classifier learning are linked together. It has been demonstrated that decoupling of these two processes can provide more flexibility for training a better classifier. In these studies, linear discriminant analysis (LDA) has been adopted to select weak classifiers independently based on class separability rather than a training error that occurs normally in AdaBoost. It is observed that LDA is successful only if a large number of labeled training samples is available. However, a large-scale labeled training set is not always available in many computer vision applications such as object classification. To tackle this problem, this paper proposes semisupervised subspace learning combined with a boosting framework for object classification, through which unlabeled data can participate in the boosting training to compensate for the lack of enough labeled data. With the proposed framework, this paper develops three various approaches that utilize unlabeled data in different ways. According to the experiments on several public image data sets, the proposed methods achieve superior performance over AdaBoost and existing semisupervised algorithms.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Fumin Shen, Zhenmin Tang
IEEE Trans. Circuits Syst. Video Technol.5
2013 Inductive Hashing on Manifolds
abstract
Learning based hashing methods have attracted considerable attention due to their ability to greatly increase the scale at which existing algorithms may operate. Most of these methods are designed to generate binary codes that preserve the Euclidean distance in the original space. Manifold learning techniques, in contrast, are better able to model the intrinsic structure embedded in the original high-dimensional data. The complexity of these models, and the problems with out-of-sample data, have previously rendered them unsuitable for application to large-scale embedding, however. In this work, we consider how to learn compact binary embeddings on their intrinsic manifolds. In order to address the above-mentioned difficulties, we describe an efficient, inductive solution to the out-of-sample data problem, and a process by which non-parametric manifold learning may be used as the basis of a hashing method. Our proposed approach thus allows the development of a range of new hashing techniques exploiting the flexibility of the wide variety of manifold learning approaches available. We particularly show that hashing on the basis of t-SNE [29] outperforms state-of-the-art hashing methods on large-scale benchmark datasets, and is very effective for image classification with very short code lengths.
Fumin Shen, Chunhua Shen, Qinfeng Shi, Anton van den Hengel, Zhenmin Tang
CVPR5
2013 Pavement crack detection based on saliency and statistical features
abstract
Traditional pavement crack detection methods can not cope well with the complexity and diversity of noises in large image area. To solve this problem, we propose a novel unsupervised crack detection approach based on saliency and statistical features. The saliency is initially represented by a conspicuity map built from the intensity rarity and local contrast of image regions. Then spatial continuity of candidate crack pixels is measured based on the statistical features extracted in their neighborhood. This is followed by a Bayesian model to automatically update the saliency map. Finally, cracks are extracted after adaptive saliency map binarization. Experiments show that proposed method has generated consistent results as those by human visual inspection. The results have also proved the effectiveness of the proposed method in suppressing noises compared with several alternative methods.
Zhenmin Tang, Jun Zhou 0001, Jundi Ding
ICIP2
2013 Training boosting-like algorithms with semi-supervised subspace learning
abstract
Boosting algorithms have attracted great attention since the first real-time face detector by Viola & Jones through feature selection and strong classifier learning simultaneously. On the other hand, researchers have proposed to decouple such two procedures to improve the performance of Boosting algorithms. Motivated by this, we propose a boosting-like algorithm framework by embedding semi-supervised subspace learning methods. It selects weak classifiers based on class-separability. Combination weights of selected weak classifiers can be obtained by subspace learning. Three typical algorithms are proposed under this framework and evaluated on public data sets. As shown by our experimental results, the proposed methods obtain superior performances over their supervised counterparts and AdaBoost.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Fumin Shen, Zhenmin Tang
ICIP5
2013 WLBP: Weber local binary pattern for local image description
Fan Liu 0003, Zhenmin Tang, Jinhui Tang 0001
Neurocomputing2
2013 Locality constrained representation based classification with spatial pyramid patches
Fumin Shen, Zhenmin Tang, Jingsong Xu
Neurocomputing2
2013 Improving protein-ATP binding residues prediction by boosting SVMs with random under-sampling
Dongjun Yu, Jun Hu 0011, Zhenmin Tang, Hong-Bin Shen, Jian Yang 0003, Jing-Yu Yang 0001
Neurocomputing3
2013 Asymptotic stability of bidirectional associative memory neural networks with time-varying delays via delta operator approach
Zhengli Zhao, Fangzhou Liu 0001, Xiaochen Xie, Xiaohui Liu 0001, Zhenmin Tang
Neurocomputing5
2013 Minimum cost attribute reduction in decision-theoretic rough set models
Xiuyi Jia, Wenhe Liao, Zhenmin Tang, Lin Shang 0001
Inf. Sci.3
2013 Approximate Least Trimmed Sum of Squares Fitting and Applications in Image Analysis
abstract
The least trimmed sum of squares (LTS) regression estimation criterion is a robust statistical method for model fitting in the presence of outliers. Compared with the classical least squares estimator, which uses the entire data set for regression and is consequently sensitive to outliers, LTS identifies the outliers and fits to the remaining data points for improved accuracy. Exactly solving an LTS problem is NP-hard, but as we show here, LTS can be formulated as a concave minimization problem. Since it is usually tractable to globally solve a convex minimization or concave maximization problem in polynomial time, inspired by , we instead solve LTS' approximate complementary problem, which is convex minimization. We show that this complementary problem can be efficiently solved as a second order cone program. We thus propose an iterative procedure to approximately solve the original LTS problem. Our extensive experiments demonstrate that the proposed method is robust, efficient and scalable in dealing with problems where data are contaminated with outliers. We show several applications of our method in image analysis.
Fumin Shen, Chunhua Shen, Anton van den Hengel, Zhenmin Tang
IEEE Trans. Image Process.4
2012 Object Detection Based on Co-occurrence GMuLBP Features
abstract
Image co-occurrence has shown great powers on object classification because it captures the characteristic of individual features and spatial relationship between them simultaneously. For example, Co-occurrence Histogram of Oriented Gradients (CoHOG) has achieved great success on human detection task. However, the gradient orientation in CoHOG is sensitive to noise. In addition, CoHOG does not take gradient magnitude into account which is a key component to reinforce the feature detection. In this paper, we propose a new LBP feature detector based image co-occurrence. Building on uniform Local Binary Patterns, the new feature detector detects Co-occurrence Orientation through Gradient Magnitude calculation. It is known as CoGMuLBP. An extension version of the GoGMuLBP is also presented. The experimental results on the UIUC car data set show that the proposed features outperform state-of-the-art methods.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang
ICME4
2012 PLDA Modeling in I-Vector and Supervector Space for Speaker Verification
abstract
In this paper, we advocate the use of uncompressed form of i-vector. We employ the probabilistic linear discriminant analysis (PLDA) to handle speaker and session variability for speaker verification task. An i-vector is a low-dimensional vector containing both speaker and channel information acquired from a speech segment. When PLDA is used on i-vector, dimension reduction is performed twice – first in the i-vector extraction process and second in the PLDA model. Keeping the full dimensionality of i-vector in the supervector space for PLDA modeling and scoring would avoid unnecessary loss of information. The drawback of using PLDA on uncompressed i-vector is the inversion of large matrices, which we show can be solved rather efficiently by portioning large matrix into smaller blocks. We also introduce the Gaussianized rank-norm, as an alternative to whitening, for feature normalization prior to PLDA modeling. Index Terms: speaker verification, i-vector, probabilistic LDA 1.
Kong-Aik Lee, Zhenmin Tang, Bin Ma 0001, Anthony Larcher, Haizhou Li 0001
INTERSPEECH3
2012 Abnormal behavior recognition system for ATM monitoring by RGB-D camera
abstract
In this demo, we present an effective real-time system for ATM intelligent monitoring by using Kinect of Microsoft. With Kinect, we can easily detect people in ATM room and get their position information. By analyzing position information and video content, the system detects abnormal behaviors such as face-hiding, peeping and wandering, while records the time of these abnormal videos. Therefore, it not only prevents crimes but also helps to find suspects quickly after crimes have happened. The experimental results show that the system has the advantages of robustness, and provides a new mean for preventing financial crimes.
Fan Liu 0003, Jinhui Tang 0001, Ruizhen Zhao, Zhenmin Tang
ACM Multimedia4
2012 Fast and Accurate Human Detection Using a Cascade of Boosted MS-LBP Features
abstract
In this letter, a new scheme for generating local binary patterns (LBP) is presented. This Modified Symmetric LBP (MS-LBP) feature takes advantage of LBP and gradient features. It is then applied into a boosted cascade framework for human detection. By combining MS-LBP with Haar-like feature into the boosted framework, the performances of heterogeneous features based detectors are evaluated for the best trade-off between accuracy and speed. Two feature training schemes, namely Single AdaBoost Training Scheme (SATS) and Dual AdaBoost Training Scheme (DATS) are proposed and compared. On the top of AdaBoost, two multidimensional feature projection methods are described. A comprehensive experiment is presented. Apart from obtaining higher detection accuracy, the detection speed based on DATS is 17 times faster than HOG method.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang
IEEE Signal Process. Lett.4
2011 A Variable Muitlgranulation Rough Sets Approach
Ming Zhang 0033, Zhenmin Tang, Weiyan Xu, Xibei Yang
ICIC (3)2
2011 Face Recognition with Multi-feature Joint Representation
abstract
Different methods have been proposed over the last few years to improve the recognition rate for face images. In this paper, the merits of multi-feature joint representation based for face recognition is studied. The whole approach of face recognition can be separated into two phases: training phase and recognition phase. At first, given a query image, we train the recognition system by using the gabor and gradient features together to represent the face images. In the second phase, modular LRC classification will be used to classify the face images rather than an NN classification. Unlike the traditional LRC algorithm which operates directly on the whole face image patterns, the modular method operates on sub-blocks partitioned from an original whole face image. Experiments are carried on two face databases, the results show that the combination of the gabor information and the gradient information by modular LRC are better than the method using the single information.
Zhenmin Tang
ICIG2
2011 Nearest-neighbor classifier motivated marginal discriminant projections for face recognition
Zhenmin Tang, Cai-Kou Chen, Xintian Cheng
Frontiers Comput. Sci. China2
2010 A modification of kernel discriminant analysis for high-dimensional data - with application to face recognition
Dake Zhou, Zhenmin Tang
Signal Process.2
2010 Kernel-based improved discriminant analysis and its application to face recognition
Dake Zhou, Zhenmin Tang
Soft Comput.2
2008 Particle swarm based stereo algorithm and disparity map evaluation
abstract
In this paper, a new particle swarm based stereo algorithm is presented. Our motivation is to improve the accuracy of the disparity map by removing the mismatches caused by both occlusions and false targets. In our approach, the stereo matching problem is divided into two steps, including partial matching of segmented image and particle swarm optimization of the rest. The algorithm first takes advantage of SAD and Dynamic Programming to remove the mismatches mainly caused by visibility problems; after the first step, the algorithm selects all the rest image segmented regions, takes them as a particle and uses particle swarm to optimization it. In the second step, the cost function is defined on the pixel level, as well as on the segmented level, while the pixel level measures the data similarity based the current disparity map, the segmented level incorporates a smooth term. Results obtained for benchmark indicate that the proposed method is able to get rather accurate disparity maps.
Haofeng Zhang 0001, Chunxia Zhao, Zhenmin Tang, Jing-Yu Yang 0001
ICARCV3
2008 Hierarchical initialization approach for K-Means clustering
Jianfeng Lu 0003, J. B. Tang, Zhenmin Tang, Jing-Yu Yang 0001
Pattern Recognit. Lett.3
2007 Transformation-Based GMM with Improved Cluster Algorithm for Speaker Identification
Limin Xu, Zhenmin Tang, Keke He, Bo Qian 0004
PAKDD2
2001 A theorem on the uncorrelated optimal discriminant vectors
Zhong Jin, Jing-Yu Yang 0001, Zhenmin Tang, Zhong-Shan Hu
Pattern Recognit.3