Jingsong Xu

dblp:118/9311 · DBLP profile ↗
← Back
36ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-9102-3616ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Trajectory anomaly detection via spatial-temporal complementary reconstruction diffusion model
Tonglong Wei, Tianzi Yan, Jingsong Xu, Qisen Xu, Shengnan Guo 0001, Huaiyu Wan, Youfang Lin
GeoInformatica4
2022 Collaborative Feature Learning for Gait Recognition Under Cloth Changes
abstract
Since gait can be utilized to identify individuals from a far distance without their interaction and coordination, recently many gait recognition methods have been proposed. However, due to a real-world scenario of clothing changes, a degradation occurs for most of these methods. Thus in this paper, a more efficient gait recognition method is proposed to address the problem of clothing variances. First, part-based gait features are formulated from two different perspectives,i.e., the separated body parts that are more robust to clothing changes and the estimated human skeleton key-point regions. It is reasonable to formulate such features for cloth-changing gait recognition, because these two perspectives are both less vulnerable to clothing changes. Given that each feature has its own advantages and disadvantages, a more efficient gait feature is generated in this paper by assembling these two features together. Moreover, since local features are more discriminative than global features, in this paper more attention is focused on the local short-range features. Also, unlike most methods, in our method we treat the estimated key-point features as a set of word embeddings, and a transformer encoder is specifically used to learn the dependence of each correlative key-points. The robustness and effectiveness of our proposed method are certified by experiments on CASIA Gait Dataset B, and it has achieved the state-of-the-art performance on this dataset.
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.4
2022 Alleviating Modality Bias Training for Infrared-Visible Person Re-Identification
abstract
The task of infrared-visible person re-identification (IV-reID) is to recognize people across two modalities (i.e., RGB and IR). Existing cutting-edge approaches normally use a pair of images that have the same IDs (i.e., ID-tied cross-modality image pairs) and input them into an ImageNet-trained ResNet50. The ResNet50 backbone model can learn shared features across modalities to tolerate modality discrepancies between RGB and IR. This work will unveil a Modality Bias Training (MBT) problem that is less discussed in IV-reID, which will demonstrate that MBT significantly compromises the performance of IV-reID. Due to MBT, IR information can be overwhelmed by RGB information during training when the ResNet50 model is pretrained based on a large amount of RGB images from ImageNet. Thus, the trained models are more inclined to RGB information. Accordingly, the cross-modality generalization ability of the model is also compromised. To tackle this issue, we present a Dual-level Learning Strategy (DLS) that 1) enforces the focus of the network on ID-exclusive (rather than ID-tied) labels of cross-modality image pairs to mitigate the problem of MBT and 2) introduces third modality data that contain both RGB and IR information to further prevent the information from the IR modality from being overwhelmed during training. Our third modality images are generated by a generative adversarial network. A dynamic ID-exclusive Smooth (dIDeS) label is proposed for the generated third modality data. In experiments, comprehensive experiments are carried out to demonstrate the success of DLS in tackling the MBT issue exposed in IV-reID.
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Peng Zhang 0057, Zhaoxiang Zhang 0001
IEEE Trans. Multim.3
2022 Multimodal Marketing Intent Analysis for Effective Targeted Advertising
abstract
People’s daily information sharing and acquisition through the Internet has become more and more popular. The comprehensive multimodal marketing advertorial generated by ‘We Media’ accounts besides the normal social news is gaining its importance on social media platforms. In order to achieve effective advertising, the marketing intent understanding is a key step towards generating targeted advertising strategies (push advertorials to specific people at a specific time). However, advertorials in real are usually designed to pretend as normal social news with a wide range of contents. This poses big challenges to the platforms on accurately recognizing and analyzing the marketing intents behind the advertorials. As a pioneering study, we address this new problem of multimodal-based marketing intent analysis and answer three core questions: (1) does a piece of social news contain marketing intent? (2) what is the topic of marketing intent? (3) what is the extent of marketing intent? Towards this end, we propose a novel Multimodal-based Marketing Intent Analysis scheme (MMIA) to estimate the marketing intent embedded in the multimodal contents. Specifically, a novel supervised neural autoregressive model (SmiDocNADE) is proposed to enhance the discriminative capacity of the learned hidden features so that a single system is capable of solving the three questions. In order to effectively model inter-correlations between images and text in advertorials, we fuse multimodal data and extract features by Graph Convolution Networks as an enhancement to SmiDocNADE. The extensive evaluations demonstrate the advantages of our proposed system in multimodal-based marketing intent analysis from multiple aspects.
Lu Zhang 0062, Jialie Shen 0001, Jian Zhang 0002, Jingsong Xu, Zhibin Li 0002, Yazhou Yao, Litao Yu
IEEE Trans. Multim.4
2022 Unsupervised Image and Text Fusion for Travel Information Enhancement
abstract
With the explosive growth of the shared information on social media platforms, people are increasingly interested in sharing and making their travel plans by referring to others’ travel experiences. However, different social media sources render the heterogeneity of these valuable data, bringing difficulties for data collection and fusion. Thus, facing massive information online, one of the biggest challenges to enhance travel information is how to integrate and match these multi-source data without clear labels. In this paper, we propose an unsupervised method to fuse and match images and travelogues. We first use the three textual components (title, tag, and description) of the descriptive texts of images as three criteria to embed travelogues and the descriptive texts of images, and further introduce images into our method by joint embedding texts and images. Finally, a multiple kernel clustering approach is adopted for matching travelogues and images. Extensive experiments conducted on the real dataset crawled from two websites (Flickr and TripAdvisor) demonstrate the effectiveness and robustness of our proposed method.
Lu Zhang 0062, Jingsong Xu, Yongshun Gong, Litao Yu, Jian Zhang 0002, Jialie Shen 0001
IEEE Trans. Multim.2
2022 Recognizing Gaits Across Walking and Running Speeds
abstract
For decades, very few methods were proposed for cross-mode (i.e., walking vs. running) gait recognition. Thus, it remains largely unexplored regarding how to recognize persons by the way they walk and run. Existing cross-mode methods handle the walking-versus-running problem in two ways, either by exploring the generic mapping relation between walking and running modes or by extracting gait features which are non-/less vulnerable to the changes across these two modes. However, for the first approach, a mapping relation fit for one person may not be applicable to another person. There is no generic mapping relation given that walking and running are two highly self-related motions. The second approach does not give more attention to the disparity between walking and running modes, since mode labels are not involved in their feature learning processes. Distinct from these existing cross-mode methods, in our method, mode labels are used in the feature learning process, and a mode-invariant gait descriptor is hybridized for cross-mode gait recognition to handle this walking-versus-running problem. Further research is organized in this article to investigate the disparity between walking and running. Running is different from walking not only in the speed variances but also, more significantly, in prominent gesture/motion changes. According to these rationales, in our proposed method, we give more attention to the differences between walking and running modes, and a robust gait descriptor is developed to hybridize the mode-invariant spatial and temporal features. Two multi-task learning-based networks are proposed in this method to explore these mode-invariant features. Spatial features describe the body parts non-/less affected by mode changes, and temporal features depict the instinct motion relation of each person. Mode labels are also adopted in the training phase to guide the network to give more attention to the disparity across walking and running modes. In addition, relevant experiments on OU-ISIR Treadmill Dataset A have affirmed the effectiveness and feasibility of the proposed method. A state-of-the-art result can be achieved by our proposed method on this dataset.
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2021 Clothing Status Awareness for Long-Term Person Re-Identification
abstract
Long-Term person re-identification (LT-reID) exposes extreme challenges because of the longer time gaps between two recording footages where a person is likely to change clothing. There are two types of approaches for LT-reID: biometrics-based approach and data adaptation based approach. The former one is to seek clothing irrelevant biometric features. However, seeking high quality biometric feature is the main concern. The latter one adopts fine-tuning strategy by using data with significant clothing change. However, the performance is compromised when it is applied to cases without clothing change. This work argues that these approaches in fact are not aware of clothing status (i.e., change or no-change) of a pedestrian. Instead, they blindly assume all footages of a pedestrian have different clothes. To tackle this issue, a Regularization via Clothing Status Awareness Network (RCSANet) is proposed to regularize descriptions of a pedestrian by embedding the clothing status awareness. Consequently, the description can be enhanced to maintain the best ID discriminative feature while improving its robustness to real-world LT-reID where both clothing-change case and no-clothing-change case exist. Experiments show that RCSANet performs reasonably well on three LT-reID datasets.
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Zhaoxiang Zhang 0001
ICCV3
2021 Incorporating Multimodal Cues for Advertorial Discovery
abstract
Commercial advertorials shared on websites are usually designed to pretend as normal social news for commercial benefits. The analysis of the commercial intents embedded in advertorials can greatly help media platforms personalize content. However, commercial intents are not only concealed in news texts but also conveyed by news images explicitly or implicitly. Consequently, how to effectively extract and incorporate the crucial cues of multiple modalities has been emerging as an important but challenging problem. Motivated by this observation, we propose a framework Multimodal Advertorial Discovery Model (MADM) to estimate the commercial intents embedded in the multimodal social news. Specifically, a novel Cross-graph Fusion (CGF) strategy is developed to achieve a soft assignment to incorporate images and text and generate comprehensive multimodal representations. The extensive evaluations demonstrate the superiority of our proposed system in multimodal-based advertorial detection and analysis.
Lu Zhang 0062, Jian Zhang 0002, Jialie Shen 0001, Jingsong Xu, Zhibin Li 0002, Litao Yu
ICME4
2021 Unsupervised Domain Adaptation with Background Shift Mitigating for Person Re-Identification
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Zhaoxiang Zhang 0001
Int. J. Comput. Vis.3
2021 Beyond modality alignment: Learning part-level representation for visible-infrared person re-identification
Peng Zhang 0057, Qiang Wu 0001, Xunxiang Yao, Jingsong Xu
Image Vis. Comput.4
2021 Low-Rank Pairwise Alignment Bilinear Network For Few-Shot Fine-Grained Image Classification
abstract
Deep neural networks have demonstrated advanced abilities on various visual classification tasks, which heavily rely on the large-scale training samples with annotated ground-truth. However, it is unrealistic always to require such annotation in real-world applications. Recently, Few-Shot learning (FS), as an attempt to address the shortage of training samples, has made significant progress in generic classification tasks. Nonetheless, it is still challenging for current FS models to distinguish the subtle differences between fine-grained categories given limited training data. To filling the classification gap, in this paper, we address the Few-Shot Fine-Grained (FSFG) classification problem, which focuses on tackling the fine-grained classification under the challenging few-shot learning setting. A novel low-rank pairwise bilinear pooling operation is proposed to capture the nuanced differences between the support and query images for learning an effective distance metric. Moreover, a feature alignment layer is designed to match the support image features with query ones before the comparison. We name the proposed model Low-Rank Pairwise Alignment Bilinear Network (LRPABN), which is trained in an end-to-end fashion. Comprehensive experimental results on four widely used fine-grained classification data sets demonstrate that our LRPABN model achieves the superior performances compared to state-of-the-art methods.
Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Jingsong Xu, Qiang Wu 0001
IEEE Trans. Multim.4
2021 Learning Spatial-Temporal Representations Over Walking Tracklet for Long-Term Person Re-Identification in the Wild
abstract
Long-term person re-identification (re-ID) aims to build identity correspondence of the Target Subject of Interest (TSI) exposed under surveillance cameras over a long time interval. Compared to the conventional short-term re-ID studied by most existing works, it suffers an additional problem: significant dressing change observed with time lapsing. Unfortunately, this variation in long-term person re-ID case contradicts the assumption of prior short-term re-ID approaches, and thus causes significant difficulties if conventional short-term re-ID methods are applied. To address the problem, this paper proposes to learn hybrid feature representation via a two-stream network named SpTSkM, including a spatial-temporal stream and a skeleton motion stream. The former performs directly on image sequences, which tends to learn identity-related spatial-temporal patterns such as body geometric structure and body movement. The latter operates on normalized 3D skeletons by adapting graph convolutional network, which tends to learn pure motion patterns from skeleton sequences. Both streams extract fine-grained level time-gap stable information that is robust to appearance changes in long-term re-ID and meanwhile maintains sufficient discriminability to differentiate different people. The final matching metric is obtained by mixing information of the two streams in a score-level fusion strategy. In addition, we collect a Cloth-Varying vIDeo re-ID (CVID-reID) dataset particularly for long-term re-ID. It contains video tracklets of celebrities posted on the Internet. These videos are snapshots under extremely different scenarios that include highly dynamic background, diverse camera views and abundant cloth variations on each TSI. These factors cause CVID-reID more complicated and closer to practice. Our experiments demonstrate the difficulty of long-term person re-ID and also validate the effectiveness of the proposed SpTSkM, showing the best performance.
Peng Zhang 0057, Jingsong Xu, Qiang Wu 0001, Yan Huang 0023, Xianye Ben
IEEE Trans. Multim.2
2020 Towards Better Graph Representation: Two-Branch Collaborative Graph Neural Networks For Multimodal Marketing Intention Detection
abstract
Inspired by the fact that spreading and collecting information through the Internet becomes the norm, more and more people choose to post for-profit contents (images and texts) in social networks. Due to the difficulty of network censors, malicious marketing may be capable of harming the society. Therefore, it is meaningful to detect marketing intentions online automatically. However, gaps between multimodal data make it difficult to fuse images and texts for content marketing detection. To this end, this paper proposes Two-Branch Collaborative Graph Neural Networks to collaboratively represent multimodal data by Graph Convolution Networks (GCNs) in an end-to-end fashion. We first separately embed groups of images and texts by GCNs layers from two views and further adopt the proposed multimodal fusion strategy to learn the graph representation collaboratively. Experimental results demonstrate that our proposed method achieves superior graph classification performance for marketing intention detection.
Lu Zhang 0062, Jian Zhang 0002, Zhibin Li 0002, Jingsong Xu
ICME4
2020 Part-based Collaborative Spatio-temporal Feature Learning for Cloth-changing Gait Recognition
abstract
In decades many gait recognition methods have been proposed using different techniques. However, due to a real-world scenario of clothing variations, a reduction of the recognition rate occurs for most of these methods. Thus in this paper, a part-based spatio-temporal feature learning method is proposed to tackle the problem of clothing variations for gait recognition. First, based on the anatomical properties, human bodies are segmented into two regions, which are affected and unaffected by clothing variations. A learning network is particularly proposed in this paper to grasp principal spatio-temporal features from those unaffected regions. Different from most part-based methods with spatial or temporal features solely being utilized, in our method these two features are associated in a more collaborative manner. Snapshots are created for each gait sequence from the H-W and T-W views. Stable spatial information is embedded in the H-W view and adequate temporal information is embedded in the T-W view. An inherent relationship exists between these two views. Thus, a collaborative spatio-temporal feature will be hybridized by concatenating these correlative spatial and temporal information. The robustness and efficiency of our proposed method are validated by experiments on CASIA Gait Dataset B and OU-ISIR Treadmill Gait Dataset B. Our proposed method can both achieve the state-of-the-art results on these two databases.
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Jingsong Xu
ICPR5
2020 Automatic Sheep Counting by Multi-object Tracking
abstract
Animal counting is a highly skilled yet tedious task in livestock transportation and trading. To effectively free up the human labour and provide accurate counts for sheep loading/unloading, we develop an auto sheep counting system based on multi-object detection, tracking and extrapolation techniques. Our system has demonstrated more than 99.9% accuracy with sheep moving freely in a race under optimal visual conditions.
Jingsong Xu, Litao Yu, Jian Zhang 0002, Qiang Wu 0001
VCIP1
2020 Generated Data With Sparse Regularized Multi-Pseudo Label for Person Re-Identification
abstract
Recently, Generative Adversarial Network (GAN) has been adopted to improve person re-identification (person re-ID) performance through data augmentation. However, directly leveraging generated data to train a re-ID model may easily lead to over-fitting issue on these extra data and decrease the generalisability of model to learn true ID-related features from real data. Inspired by the previous approach which assigns multi-pseudo labels on the generated data to reduce the risk of over-fitting, we propose to take sparse regularization into consideration. We attempt to further improve the performance of current re-ID models by using the unlabeled generated data. The proposed Sparse Regularized Multi-Pseudo Label (SRMpL) can effectively prevent the over-fitting issue when some larger weights are assigned to the generated data. Our experiments are carried out on two publicly available person re-ID datasets (e.g., Market-1501 and DukeMTMC-reID). Compared with existing unlabeled generated data re-ID solutions, our approach achieves competitive performance. Two classical re-ID models are used to verify our sparse regularization label on generated data, i.e., an ID-embedding network and a two-stream network.
Liqin Huang, Junyi Wu 0001, Yan Huang 0023, Qiang Wu 0001, Jingsong Xu
IEEE Signal Process. Lett.6
2020 Beyond Scalar Neuron: Adopting Vector-Neuron Capsules for Long-Term Person Re-Identification
abstract
Current person re-identification (re-ID) works mainly focus on the short-term scenario where a person is less likely to change clothes. However, in the long-term re-ID scenario, a person has a great chance to change clothes. A sophisticated re-ID system should take such changes into account. To facilitate the study of long-term re-ID, this paper introduces a large-scale re-ID dataset called “Celeb-reID” to the community. Unlike previous datasets, the same person can change clothes in the proposed Celeb-reID dataset. Images of Celeb-reID are acquired from the Internet using street snap-shots of celebrities. There is a total of 1,052 IDs with 34,186 images making Celeb-reID being the largest long-term re-ID dataset so far. To tackle the challenge of cloth changes, we propose to use vector-neuron (VN) capsules instead of the traditional scalar neurons (SN) to design our network. Compared with SN, one extra-dimensional information in VN can perceive cloth changes of the same person. We introduce a well-designed ReIDCaps network and integrate capsules to deal with the person re-ID task. Soft Embedding Attention (SEA) and Feature Sparse Representation (FSR) mechanisms are adopted in our network for performance boosting. Experiments are conducted on the proposed long-term re-ID dataset and two common short-term re-ID datasets. Comprehensive analyses are given to demonstrate the challenge exposed in our datasets. Experimental results show that our ReIDCaps can outperform existing state-of-the-art methods by a large margin in the long-term scenario.The new dataset and code will be released to facilitate future researches.
Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Yi Zhong 0002, Peng Zhang 0057, Zhaoxiang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Top-Push Constrained Modality-Adaptive Dictionary Learning for Cross-Modality Person Re-Identification
abstract
Person re-identification aims to match person captured by multiple non-overlapping cameras that mainly mean standard RGB cameras. In contemporary surveillance, cameras of different modalities such as infrared cameras and depth cameras are introduced because of their unique advantages in poor illumination scenarios. However, re-identifying the persons across such cameras of different modalities is extremely difficult and, unfortunately, seldom discussed. It is mainly caused by extremely different appearances of the person shown under such different camera modalities. In this paper, we tackle this challenging cross-modality people re-identification through a top-push constrained modality-adaptive dictionary learning. The proposed model asymmetrically projects the heterogeneous features from dissimilar modalities onto a common space. In this way, the modality-specific bias is mitigated. Thus, the heterogeneous data can be simultaneously enforced by a shared dictionary in a canonical space. Moreover, a top-push ranking graph regularization is embedded in the proposed model to improve the discriminability, which efficiently further boosts the matching accuracy. In order to implement the proposed model, an iterative process is developed in this paper to optimize these two processes jointly. Extensive experiments on the benchmark SYSU-MM01 and BIWI RGBD-ID person re-identification datasets show promising results which outperform state-of-the-art methods.
Peng Zhang 0057, Jingsong Xu, Qiang Wu 0001, Yan Huang 0023, Jian Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.2
2019 SBSGAN: Suppression of Inter-Domain Background Shift for Person Re-Identification
abstract
Cross-domain person re-identification (re-ID) is challenging due to the bias between training and testing domains. We observe that if backgrounds in the training and testing datasets are very different, it dramatically introduces difficulties to extract robust pedestrian features, and thus compromises the cross-domain person re-ID performance. In this paper, we formulate such problems as a background shift problem. A Suppression of Background Shift Generative Adversarial Network (SBSGAN) is proposed to generate images with suppressed backgrounds. Unlike simply removing backgrounds using binary masks, SBSGAN allows the generator to decide whether pixels should be preserved or suppressed to reduce segmentation errors caused by noisy foreground masks. Additionally, we take ID-related cues, such as vehicles and companions into consideration. With high-quality generated images, a Densely Associated 2-Stream (DA-2S) network is introduced with Inter Stream Densely Connection (ISDC) modules to strengthen the complementarity of the generated data and ID-related cues. The experiments show that the proposed method achieves competitive performance on three re-ID datasets, i.e., Market-1501, DukeMTMC-reID, and CUHK03, under the cross-domain person re-ID scenario.
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002
ICCV3
2019 Compare More Nuanced: Pairwise Alignment Bilinear Network for Few-Shot Fine-Grained Learning
abstract
The recognition ability of human beings is developed in a progressive way. Usually, children learn to discriminate various objects from coarse to fine-grained with limited supervision. Inspired by this learning process, we propose a simple yet effective model for the Few-Shot Fine-Grained (FSFG) recognition, which tries to tackle the challenging fine-grained recognition task using meta-learning. The proposed method, named Pairwise Alignment Bilinear Network (PABN), is an end-to-end deep neural network. Unlike traditional deep bilinear networks for fine-grained classification, which adopt the self-bilinear pooling to capture the subtle features of images, the proposed model uses a novel pairwise bilinear pooling to compare the nuanced differences between base images and query images for learning a deep distance metric. In order to match base image features with query image features, we design feature alignment losses before the proposed pairwise bilinear pooling. Experiment results on four fine-grained classification datasets and one generic few-shot dataset demonstrate that the proposed model outperforms both the state-of-the-art few-shot fine-grained and general few-shot methods.
Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Qiang Wu 0001, Jingsong Xu
ICME5
2019 Celebrities-ReID: A Benchmark for Clothes Variation in Long-Term Person Re-Identification
abstract
This paper considers person re-identification (re-ID) in the case of long-time gap (i.e., long-term re-ID) that concentrates on the challenge of clothes variation of each person. We introduce a new dataset, named Celebrities-reID to handle that challenge. Compared with current datasets, the proposed Celebrities-reID dataset is featured in two aspects. First, it contains 590 persons with 10,842 images, and each person does not wear the same clothing twice, making it the largest clothes variation person re-ID dataset to date. Second, a comprehensive evaluation using state of the arts is carried out to verify the feasibility and new challenge exposed by this dataset. In addition, we propose a benchmark approach to the dataset where a two-step fine-tuning strategy on human body parts is introduced to tackle the challenge of clothes variation. In experiments, we evaluate the feasibility and quality of the proposed Celebrities-reID dataset. The experimental results demonstrate that the proposed benchmark approach is not only able to best tackle clothes variation shown in our dataset but also achieves competitive performance on a widely used person re-ID dataset Market1501, which further proves the reliability of the proposed benchmark approach.
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002
IJCNN3
2019 VT-GAN: View Transformation GAN for Gait Recognition Across Views
abstract
Recognizing gaits without human cooperation is of importance in surveillance and forensics because of the benefits that gait is unique and collected remotely. However, change of camera view angle severely degrades the performance of gait recognition. To address the problem, previous methods usually learn mappings for each pair of views which incurs abundant independently built models. In this paper, we proposed a View Transformation Generative Adversarial Networks (VT-GAN) to achieve view transformation of gaits across two arbitrary views using only one uniform model. In specific, we generated gaits in target view conditioned on input images from any views and the corresponding target view indicator. In addition to the classical discriminator in GAN which makes the generated images look realistic, a view classifier is imposed. This controls the consistency of generated images and conditioned target view indicator and ensures to generate gaits in the specified target view. On the other hand, retaining identity information while performing view transformation is another challenge. To solve the issue, an identity distilling module with triplet loss is integrated, which constrains the generated images inheriting identity information from inputs and yields discriminative feature embeddings. The proposed VT-GAN generates visually promising gaits and achieves promising performances for cross-view gait recognition, which exhibits great effectiveness of the proposed VT-GAN.
Peng Zhang 0057, Qiang Wu 0001, Jingsong Xu
IJCNN3
2019 VN-GAN: Identity-preserved Variation Normalizing GAN for Gait Recognition
abstract
Gait is recognized as a unique biometric characteristic to identify a walking person remotely across surveillance networks. However, the performance of gait recognition severely suffers challenges from view angle diversity. To address the problem, an identity-preserved Variation Normalizing Generative Adversarial Network (VN-GAN) is proposed for learning purely identity-related representations. It adopts a coarse-to-fine manner which firstly generates initial coarse images by normalizing view to an identical one and then refines the coarse images by injecting identity-related information. In specific, Siamese structure with discriminators for both camera view angles and human identities is utilized to achieve variation normalization and identity preservation of two stages, respectively. In addition to discriminators, reconstruction loss and identity-preserving loss are integrated, which forces the generated images to be the same in view and to be discriminative in identity. This ensures to generate identity-related images in an identical view of good visual effect for gait recognition. Extensive experiments on benchmark datasets demonstrate that the proposed VN-GAN can generate visually interpretable results and achieve promising performance for gait recognition.
Peng Zhang 0057, Qiang Wu 0001, Jingsong Xu
IJCNN3
2019 Improving Person Re-Identification Performance Using Body Mask Via Cross-Learning Strategy
abstract
The task of person re-identification (re-id) is to find the same pedestrian across non-overlapping cameras. Normally, the performance of person re-id can be affected by background clutters. However, existing segmentation algorithms are hard to obtain perfect foreground person images. To effectively leverage the body (foreground) cue, and in the meantime pay attention to discriminative information in the background (e.g., companion or vehicle), we propose to use a cross-learning strategy to take both foreground and other discriminative information into account. In addition, since currently existing foreground segmentation result always involves noise, we use Label Smoothing Regularization (LSR) to strengthen the generalization capability during our learning process. In experiments, we pick up two state-of-the-art person re-id methods to verify the effectiveness of our proposed cross-learning strategy. Our experiments are carried out on two publicly available person re-id datasets. Obvious performance improvements can be observed on both datasets.
Junyi Wu 0001, Lingxiang Yao, Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Liqin Huang
VCIP4
2019 Multi-Pseudo Regularized Label for Generated Data in Person Re-Identification
abstract
Sufficient training data normally is required to train deeply learned models. However, due to the expensive manual process for labelling large number of images (i.e., annotation), the amount of available training data (i.e., real data) is always limited. To produce more data for training a deep network, Generative Adversarial Network (GAN) can be used to generate artificial sample data (i.e., generated data). However, the generated data usually does not have annotation labels. To solve this problem, in this paper, we propose a virtual label called Multi-pseudo Regularized Label (MpRL) and assign it to the generated data. With MpRL, the generated data will be used as the supplementary of real training data to train a deep neural network in a semi-supervised learning fashion. To build the corresponding relationship between the real data and generated data, MpRL assigns each generated data a proper virtual label which reflects the likelihood of the affiliation of the generated data to predefined training classes in the real data domain. Unlike the traditional label which usually is a single integral number, the virtual label proposed in this work is a set of weight-based values each individual of which is a number in (0,1] called multi-pseudo label and reflects the degree of relation between each generated data to every pre-defined class of real data. A comprehensive evaluation is carried out by adopting two state-of-the-art convolutional neural networks (CNNs) in our experiments to verify the effectiveness of MpRL. Experiments demonstrate that by assigning MpRL to generated data, we can further improve the person re-ID performance on five re-ID datasets, i.e., Market-1501, DukeMTMC-reID, CUHK03, VIPeR, and CUHK01. The proposed method obtains +6.29%, +6.30%, +5.58%, +5.84%, and +3.48% improvements in rank-1 accuracy over a strong CNN baseline on the five datasets respectively, and outperforms state-of-the-art methods.
Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Zhedong Zheng, Zhaoxiang Zhang 0001, Jian Zhang 0002
IEEE Trans. Image Process.2
2018 Long-Term Person Re-identification Using True Motion from Videos
abstract
Most person re-identification approaches and benchmarks assume that pedestrians go across the surveillance network without significant appearance changes in a brief period, which explicitly restricts person re-identification to a short-term event and incurs inter-sample similarity measurement by appearance matching. However, pedestrians are likely to reappear in the surveillance network after a long-time interval (long-term) and change their wearing in many real-world scenarios. These scenarios inevitably cause appearances between subjects more ambiguous and indistinguishable. In this paper we consider these scenarios and propose a unified feature representation based on true motion cues from videos named FIne moTion encoDing (FITD). Our hypothesis is that people keep constant motion patterns under non-distraction walking condition. Therefore, the motion characteristics are more reliable than static appearance feature to describe a walking person. Particularly, we extract motion patterns hierarchically by encoding trajectory-aligned descriptors with Fisher vectors in a spatial-aligned pyramid. To verify benefits of the proposed FITD, we collect a new dataset typically for the long-term situations. Extensive experiments demonstrate the merits of our FITD especially for the long-term scenarios.
Peng Zhang 0057, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002
WACV3
2017 User relationship strength modeling for friend recommendation on Instagram
Dongyan Guo, Jingsong Xu, Jian Zhang 0002, Min Xu 0001, Xiangjian He
Neurocomputing2
2017 A new web-supervised method for image dataset constructions
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
Neurocomputing5
2017 Exploiting Web Images for Dataset Construction: A Domain Robust Approach
abstract
Labeled image datasets have played a critical role in high-level image understanding. However, the process of manual labeling is both time-consuming and labor intensive. To reduce the cost of manual labeling, there has been increased research interest in automatically constructing image datasets by exploiting web images. Datasets constructed by existing methods tend to have a weak domain adaptation ability, which is known as the “dataset bias problem.” To address this issue, we present a novel image dataset construction framework that can be generalized well to unseen target domains. Specifically, the given queries are first expanded by searching the Google Books Ngrams Corpus to obtain a rich semantic description, from which the visually nonsalient and less relevant expansions are filtered out. By treating each selected expansion as a “bag” and the retrieved images as “instances,” image selection can be formulated as a multi-instance learning problem with constrained positive bags. We propose to solve the employed problems by the cutting-plane and concave-convex procedure algorithm. By using this approach, images from different distributions can be kept while noisy images are filtered out. To verify the effectiveness of our proposed approach, we build an image dataset with 20 categories. Extensive experiments on image classification, cross-dataset generalization, diversity comparison, and object detection demonstrate the domain robustness of our dataset.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
IEEE Trans. Multim.5
2016 Automatic image dataset construction with multiple textual metadata
abstract
The goal of this work is to automatically collect a large number of highly relevant images from the Internet for given queries. A novel image dataset construction framework is proposed by employing multiple textual metadata. In specific, the given queries are first expanded by searching in the Google Books Ngrams Corpora to obtain a richer semantic description, from which the visually non-salient and less relevant expansions are then filtered. After retrieving images from the Internet with filtered expansions, we further filter noisy images by clustering and progressively Convolutional Neural Networks (CNN). To verify the effectiveness of our proposed method, we construct a dataset with 10 categories, which is not only much larger than but also have comparable cross-dataset generalization ability with manually labeled dataset STL-10 and CIFAR-10.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
ICME5
2014 Exploiting Universum data in AdaBoost using gradient descent
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang
Image Vis. Comput.1
2014 Boosting Separability in Semisupervised Learning for Object Classification
abstract
Boosting algorithms, especially AdaBoost, have attracted great attention in computer vision. In the early version of boosting algorithms, the weak classifier selection and the strong classifier learning are linked together. It has been demonstrated that decoupling of these two processes can provide more flexibility for training a better classifier. In these studies, linear discriminant analysis (LDA) has been adopted to select weak classifiers independently based on class separability rather than a training error that occurs normally in AdaBoost. It is observed that LDA is successful only if a large number of labeled training samples is available. However, a large-scale labeled training set is not always available in many computer vision applications such as object classification. To tackle this problem, this paper proposes semisupervised subspace learning combined with a boosting framework for object classification, through which unlabeled data can participate in the boosting training to compensate for the lack of enough labeled data. With the proposed framework, this paper develops three various approaches that utilize unlabeled data in different ways. According to the experiments on several public image data sets, the proposed methods achieve superior performance over AdaBoost and existing semisupervised algorithms.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Fumin Shen, Zhenmin Tang
IEEE Trans. Circuits Syst. Video Technol.1
2013 Training boosting-like algorithms with semi-supervised subspace learning
abstract
Boosting algorithms have attracted great attention since the first real-time face detector by Viola & Jones through feature selection and strong classifier learning simultaneously. On the other hand, researchers have proposed to decouple such two procedures to improve the performance of Boosting algorithms. Motivated by this, we propose a boosting-like algorithm framework by embedding semi-supervised subspace learning methods. It selects weak classifiers based on class-separability. Combination weights of selected weak classifiers can be obtained by subspace learning. Three typical algorithms are proposed under this framework and evaluated on public data sets. As shown by our experimental results, the proposed methods obtain superior performances over their supervised counterparts and AdaBoost.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Fumin Shen, Zhenmin Tang
ICIP1
2013 Locality constrained representation based classification with spatial pyramid patches
Fumin Shen, Zhenmin Tang, Jingsong Xu
Neurocomputing3
2012 Object Detection Based on Co-occurrence GMuLBP Features
abstract
Image co-occurrence has shown great powers on object classification because it captures the characteristic of individual features and spatial relationship between them simultaneously. For example, Co-occurrence Histogram of Oriented Gradients (CoHOG) has achieved great success on human detection task. However, the gradient orientation in CoHOG is sensitive to noise. In addition, CoHOG does not take gradient magnitude into account which is a key component to reinforce the feature detection. In this paper, we propose a new LBP feature detector based image co-occurrence. Building on uniform Local Binary Patterns, the new feature detector detects Co-occurrence Orientation through Gradient Magnitude calculation. It is known as CoGMuLBP. An extension version of the GoGMuLBP is also presented. The experimental results on the UIUC car data set show that the proposed features outperform state-of-the-art methods.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang
ICME1
2012 Fast and Accurate Human Detection Using a Cascade of Boosted MS-LBP Features
abstract
In this letter, a new scheme for generating local binary patterns (LBP) is presented. This Modified Symmetric LBP (MS-LBP) feature takes advantage of LBP and gradient features. It is then applied into a boosted cascade framework for human detection. By combining MS-LBP with Haar-like feature into the boosted framework, the performances of heterogeneous features based detectors are evaluated for the best trade-off between accuracy and speed. Two feature training schemes, namely Single AdaBoost Training Scheme (SATS) and Dual AdaBoost Training Scheme (DATS) are proposed and compared. On the top of AdaBoost, two multidimensional feature projection methods are described. A comprehensive experiment is presented. Apart from obtaining higher detection accuracy, the detection speed based on DATS is 17 times faster than HOG method.
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang
IEEE Signal Process. Lett.1