EDBT 2026 Demo / reviewers in the wild / expert
Bingpeng Ma
dblp:62/1822
· DBLP profile ↗
84ranked-venue papers
7as first author
37since 2021 · last 2026
0000-0001-8984-205XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 6 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 2 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bilateral Transformation of Biased Pseudo-Labels under Distribution Inconsistency
Ruibing Hou, Hong Chang 0001, Minyang Hu, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | Learning invariances via correlation and marginal alignments for out-of-distribution generalization
Zong Guo, Bingpeng Ma |
Neurocomputing | 2 |
| 2026 | Prototype-guided text-based person search on rich Chinese descriptions
Ziqiang Wu, Bingpeng Ma |
Pattern Recognit. | 2 |
| 2025 | Temporal Reordering for Video Person Re-identification Based on Feature Reappearance Score
Bingpeng Ma |
ICIG (2) | 2 |
| 2025 | Visible-Infrared Person Re-Identification via Mutual Reinforcement of Prompts and Image EncodersabstractContrastive Language-Image Pre-training (CLIP) has achieved good results in Visible-Infrared Person Re-IDentification (VI-ReID) task. However, CLIP does not focus on person-related information, so prompts generated by original CLIP can not accurately describe identity information of a person. We argue that compared to original CLIP, encoders familiar with person-related information can generate prompts which are more suitable for VI-ReID. Based on such idea, we design a novel network that helps prompts focus on person-related information through alternately optimizing the prompts and image encoders. Specifically, when optimizing prompts, we introduce modality knowledge propagation loss. The loss aligns the predicted class probability of text and image features, so that the knowledge in image encoders is transferred to prompts. When optimizing encoders, we design modality alignment loss. The loss considers text features as a bridge between two modalities, aligning features from two modalities with text features. In this way, modality discrepancies are effectively reduced. Finally, through the mutual reinforcement of two parts, the quality of both prompts and image encoders is improved in a positive feedback manner. Experiments on two widely used datasets show that the proposed network outperforms state-of-the-art methods. Hongde Zhang, Bingpeng Ma |
ICMR | 2 |
| 2025 | Clothes-Changing Person Re-Identification With Feasibility-Aware Intermediary MatchingabstractCurrent clothes-changing person re-identification (re-id) approaches usually perform retrieval based on clothes-irrelevant features, while neglecting the potential of clothes-relevant features. However, we observe that relying solely on clothes-irrelevant features for clothes-changing re-id is limited, since they often lack adequate identity information and suffer from large intra-class variations. On the contrary, clothes-relevant features can be used to discover same-clothes intermediaries that possess informative identity clues. Based on this observation, we propose a Feasibility-Aware Intermediary Matching (FAIM) framework to additionally utilizeclothes-relevant featuresfor retrieval. First, an Intermediary Matching (IM) module is designed to perform an intermediary-assisted matching process. This process involves using clothes-relevant features to find informative intermediates, and then using clothes-irrelevant features of these intermediates to complete the matching. Second, in order to reduce the negative effect of low-quality intermediaries, an Intermediary-Based Feasibility Weighting (IBFW) module is designed to evaluate the feasibility of intermediary matching process by assessing the quality of intermediaries. Extensive experiments demonstrate that our method outperforms state-of-the-art methods on several widely-used clothes-changing re-id benchmarks. Jiahe Zhao, Ruibing Hou, Hong Chang 0001, Xinqian Gu, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | An Information Theoretical View for Out-of-Distribution Detection
Jinjing Hu, Wenrui Liu 0004, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
ECCV (55) | 4 |
| 2024 | Scalable Modular Network: A Framework for Adaptive Learning via Agreement RoutingabstractIn this paper, we propose a novel modular network framework, called Scalable Modular Network (SMN), which enables adaptive learning capability and supports integration of new modules after pre-training for better adaptation.
This adaptive capability comes from a novel design of router within SMN, named agreement router, which selects and composes different specialist modules through an iterative message passing process.
The agreement router iteratively computes the agreements among a set of input and outputs of all modules to allocate inputs to specific module.
During the iterative routing, messages of modules are passed to each other, which improves the module selection process with consideration of both local interactions (between a single module and input) and global interactions involving multiple other modules.
To validate our contributions, we conduct experiments on two problems: a toy min-max game and few-shot image classification task.
Our experimental results demonstrate that SMN can generalize to new distributions and exhibit sample-efficient adaptation to new tasks.
Furthermore, SMN can achieve a better adaptation capability when new modules are introduced after pre-training.
Our code is available at https://github.com/hu-my/ScalableModularNetwork. Minyang Hu, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
ICLR | 3 |
| 2024 | Triplet Adaptation Framework for Robust Semi-Supervised LearningabstractSemi-supervised learning (SSL) suffers from severe performance degradation when labeled and unlabeled data come from inconsistent and imbalanced distribution. Nonetheless, there is a lack of theoretical guidance regarding a remedy for this issue. To bridge the gap between theoretical insights and practical solutions, we embark to an analysis of generalization bound of classic SSL algorithms. This analysis reveals that distribution inconsistency between unlabeled and labeled data can cause a significant generalization error bound. Motivated by this theoretical insight, we present a Triplet Adaptation Framework (TAF) to reduce the distribution divergence and improve the generalization of SSL models. TAF comprises three adapters: Balanced Residual Adapter, aiming to map the class distribution of labeled and unlabeled data to a uniform distribution for reducing class distribution divergence; Representation Adapter, aiming to map the representation distribution of unlabeled data to labeled one for reducing representation distribution divergence; and Pseudo-Label Adapter, aiming to align the predicted pseudo-labels with the class distribution of unlabeled data, thereby preventing erroneous pseudo-labels from exacerbating representation divergence. These three adapters collaborate synergistically to reduce the generalization bound, ultimately achieving a more robust and generalizable SSL model. Extensive experiments across various robust SSL scenarios validate the efficacy of our method. Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Incorporating texture and silhouette for video-based person re-identification
Shutao Bai, Hong Chang 0001, Bingpeng Ma |
Pattern Recognit. | 3 |
| 2024 | Enhancing identification for person search with multi-scale multi-grained representation learning
Zhixiong Han, Bingpeng Ma |
Pattern Recognit. | 2 |
| 2024 | A Comprehensive Framework for Long-Tailed Learning via Pretraining and NormalizationabstractData in the visual world often present long-tailed distributions. However, learning high-quality representations and classifiers for imbalanced data is still challenging for data-driven deep learning models. In this work, we aim at improving the feature extractor and classifier for long-tailed recognition via contrastive pretraining and feature normalization, respectively. First, we carefully study the influence of contrastive pretraining under different conditions, showing that current self-supervised pretraining for long-tailed learning is still suboptimal in both performance and speed. We thus propose a new balanced contrastive loss and a fast contrastive initialization scheme to improve previous long-tailed pretraining. Second, based on the motivative analysis on the normalization for classifier, we propose a novel generalized normalization classifier that consists of generalized normalization and grouped learnable scaling. It outperforms traditional inner product classifier as well as cosine classifier. Both the two components proposed can improve recognition ability on tail classes without the expense of head classes. We finally build a unified framework that achieves competitive performance compared with state of the arts on several long-tailed recognition benchmarks and maintains high efficiency. Nan Kang, Hong Chang 0001, Bingpeng Ma, Shiguang Shan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Predictive Consistency Learning for Long-Tailed Recognition
Nan Kang, Hong Chang 0001, Bingpeng Ma, Shutao Bai, Shiguang Shan, Xilin Chen 0001 |
BMVC | 3 |
| 2023 | Diversity-Measurable Anomaly DetectionabstractReconstruction-based anomaly detection models achieve their purpose by suppressing the generalization ability for anomaly. However, diverse normal patterns are consequently not well reconstructed as well. Although some efforts have been made to alleviate this problem by modeling sample diversity, they suffer from shortcut learning due to undesired transmission of abnormal information. In this paper, to better handle the tradeoff problem, we propose Diversity-Measurable Anomaly Detection (DMAD) framework to enhance reconstruction diversity while avoid the undesired generalization on anomalies. To this end, we design Pyramid Deformation Module (PDM), which models diverse normals and measures the severity of anomaly by estimating multi-scale deformation fields from reconstructed reference to original input. Integrated with an information compression module, PDM essentially decouples deformation from prototypical embedding and makes the final anomaly score more reliable. Experimental results on both surveillance videos and industrial images demonstrate the effectiveness of our method. In addition, DMAD works equally well in front of contaminated data and anomaly-like normal samples. Wenrui Liu 0004, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
CVPR | 3 |
| 2023 | Understanding Few-Shot Learning: Measuring Task Relatedness and Adaptation Difficulty via AttributesabstractFew-shot learning (FSL) aims to learn novel tasks with very few labeled samples by leveraging experience from \emph{related} training tasks.
In this paper, we try to understand FSL by exploring two key questions:
(1) How to quantify the relationship between \emph{ training} and \emph{novel} tasks?
(2) How does the relationship affect the \emph{adaptation difficulty} on novel tasks for different models?
To answer the first question, we propose Task Attribute Distance (TAD) as a metric to quantify the task relatedness via attributes.
Unlike other metrics, TAD is independent of models, making it applicable to different FSL models.
To address the second question, we utilize TAD metric to establish a theoretical connection between task relatedness and task adaptation difficulty.
By deriving the generalization error bound on a novel task, we discover how TAD measures the adaptation difficulty on novel tasks for different models.
To validate our theoretical results, we conduct experiments on three benchmarks.
Our experimental results confirm that TAD metric effectively quantifies the task relatedness and reflects the adaptation difficulty on novel tasks for various FSL methods, even if some of them do not learn attributes explicitly or human-annotated attributes are not provided.
Our code is available at
\href{https://github.com/hu-my/TaskAttributeDistance}{https://github.com/hu-my/TaskAttributeDistance}. Minyang Hu, Hong Chang 0001, Zong Guo, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
NeurIPS | 4 |
| 2023 | Generalized Semi-Supervised Learning via Self-Supervised Feature AdaptationabstractTraditional semi-supervised learning (SSL) assumes that the feature distributions of labeled and unlabeled data are consistent which rarely holds in realistic scenarios.
In this paper, we propose a novel SSL setting, where unlabeled samples are drawn from a mixed distribution that deviates from the feature distribution of labeled samples.
Under this setting, previous SSL methods tend to predict wrong pseudo-labels with the model fitted on labeled data, resulting in noise accumulation. To tackle this issue, we propose \emph{Self-Supervised Feature Adaptation} (SSFA), a generic framework for improving SSL performance when labeled and unlabeled data come from different distributions.
SSFA decouples the prediction of pseudo-labels from the current model to improve the quality of pseudo-labels. Particularly, SSFA incorporates a self-supervised task into the SSL framework and uses it to adapt the feature extractor of the model to the unlabeled data. In this way, the extracted features better fit the distribution of unlabeled data, thereby generating high-quality pseudo-labels. Extensive experiments show that our proposed SSFA is applicable to various pseudo-label-based SSL learners and significantly improves performance in labeled, unlabeled, and even unseen distributions. Jiachen Liang, Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
NeurIPS | 4 |
| 2023 | Dual Compensation Residual Networks for Class Imbalanced LearningabstractLearning generalizable representation and classifier for class-imbalanced data is challenging for data-driven deep models. Most studies attempt to re-balance the data distribution, which is prone to overfitting on tail classes and underfitting on head classes. In this work, we propose Dual Compensation Residual Networks to better fit both tail and head classes. First, we propose dual Feature Compensation Module (FCM) and Logit Compensation Module (LCM) to alleviate the overfitting issue. The design of these two modules is based on the observation: an important factor causing overfitting is that there is severe feature drift between training and test data on tail classes. In details, the test features of a tail category tend to drift towards feature cloud of multiple similar head categories. So FCM estimates a multi-mode feature drift direction for each tail category and compensate for it. Furthermore, LCM translates the deterministic feature drift vector estimated by FCM along intra-class variations, so as to cover a larger effective compensation space, thereby better fitting the test features. Second, we propose a Residual Balanced Multi-Proxies Classifier (RBMC) to alleviate the under-fitting issue. Motivated by the observation that re-balancing strategy hinders the classifier from learning sufficient head knowledge and eventually causes underfitting, RBMC utilizes uniform learning with a residual path to facilitate classifier learning. Comprehensive experiments on Long-tailed and Class-Incremental benchmarks validate the efficacy of our method. Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Person Search by a Bi-Directional Task-Consistent Learning ModelabstractTwo-stage person search methods achieve the state-of-the-art performance by separate detection and re-ID stages, but neglect the consistency needs between these two stages. The re-ID stage needs more accurate query bounding boxes and fewer boxes of distractors; The detection stage needs the re-ID stage to have robustness against unavailable detection errors. In this paper, we introduce a novel Bi-directional Task-Consistent Learning (BTCL) person search framework, including a Target-Specific Detector (TSD) and a re-ID model with Dynamic Adaptive Learning Structure (DALS). For the former consistency need, we add a verification head for predicting the similarity scores between query and proposals in parallel with the existing heads for bounding box recognition. Thus, TSD generates accurate boxes for the query-like pedestrians, which are suitable for the re-ID stage. For the re-ID robustness need, DALS dynamically generates a large number of possible detection results in line with the real distribution. By training the re-ID model on data with different types of detection errors, DLAS improves the model robustness to detection inputs. Experimental results show our framework achieves state-of-the-art performance on two widely-used person search datasets. Cheng Wang 0043, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Refined Knowledge Transfer for Language-Based Person SearchabstractThis paper proposes a novel method, named Refined Knowledge Transfer (RKT), for language-based person search. Existing state-of-the-art methods do not deal with knowledge imbalance between image and text. In detail, textual identity knowledge is limited, but the image contains more identity knowledge. We propose Cross-Modal Knowledge Transfer (CMKT) to enhance textual identity knowledge by image to address this problem. Besides, multiple texts of one image include more identity knowledge than a single text. Thus, we propose Intra-Modal Knowledge Transfer (IMKT) to enhance textual identity knowledge by other texts. These two types of knowledge transfer will enhance the identity knowledge in text. Additionally, by considering that identity-irrelevant knowledge is transferred to text, we propose Knowledge Refiner (KR) to refine the knowledge in text. KR is capable of preserving identity knowledge and discarding identity-irrelevant knowledge. By combining CMKT, IMKT, and KR, RKT makes textual identity knowledge more salient. Extensive experiments show the state-of-the-art performance of RKT on the CUHK-PEDES and our proposed PRW-PEDES-CN datasets. In addition, the decent generalization ability of RKT is also validated on the Flickr30K, CUB, and Flowers datasets. Ziqiang Wu, Bingpeng Ma, Hong Chang 0001, Shiguang Shan |
IEEE Trans. Multim. | 2 |
| 2022 | Salient-to-Broad Transition for Video Person Re-identificationabstractDue to the limited utilization of temporal relations in video re-id, the frame-level attention regions of mainstream methods are partial and highly similar. To address this problem, we propose a Salient-to-Broad Module (SBM) to enlarge the attention regions gradually. Specifically, in SBM, while the previous frames have focused on the most salient regions, the later frames tend to focus on broader regions. In this way, the additional information in broad regions can supplement salient regions, incurring more powerful video-level representations. To further improve SBM, an Integration-and-Distribution Module (IDM) is introduced to enhance frame-level representations. IDM first integrates features from the entire feature space and then distributes the integrated features to each spatial location. SBM and IDM are mutually beneficial since they enhance the representations from video-level and frame-level, respectively. Extensive experiments on four prevalent benchmarks demonstrate the effectiveness and superiority of our method. The source code is available at https://github.com/baist/SINet. Shutao Bai, Bingpeng Ma, Hong Chang 0001, Rui Huang 0001, Xilin Chen 0001 |
CVPR | 2 |
| 2022 | Clothes-Changing Person Re-identification with RGB Modality OnlyabstractThe key to address clothes-changing person re-identification (re-id) is to extract clothes-irrelevant features, e.g., face, hairstyle, body shape, and gait. Most current works mainly focus on modeling body shape from multi-modality information (e.g., silhouettes and sketches), but do not make full use of the clothes-irrelevant information in the original RGB images. In this paper, we propose a Clothes-based Adversarial Loss (CAL) to mine clothes-irrelevant features from the original RGB images by penalizing the predictive power of re-id model w.r.t. clothes. Extensive experiments demonstrate that using RGB images only, CAL outperforms all state-of-the-art methods on widely-used clothes-changing person re-id benchmarks. Besides, compared with images, videos contain richer appearance and additional temporal information, which can be used to model proper spatiotemporal patterns to assist clothes-changing re-id. Since there is no publicly available clothes-changing video re-id dataset, we contribute a new dataset named CCVID and show that there exists much room for improvement in modeling spatiotemporal information. The code and new dataset are available at: h t t$p$s: //github.com/guxinqian/Simple-CCReID. Xinqian Gu, Hong Chang 0001, Bingpeng Ma, Shutao Bai, Shiguang Shan, Xilin Chen 0001 |
CVPR | 3 |
| 2022 | Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework
Botao Ye, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
ECCV (22) | 3 |
| 2022 | Gradual Domain Adaptation with Sample Transferability Exploitation for Person Re-IdentificationabstractIn this paper, we propose a novel gradual domain adaptation method with sample transferability exploitation to tackle the unsupervised domain adaptation (UDA) for person re-identification (re-id). Due to the direct but rough adaptation scheme, existing UDA for person re-id methods usually suffer from source domain-specific characteristics. To filter out the source domain-specific characteristics, motivated by the curriculum learning strategy, we conduct gradual domain adaptation by domain-level re-weighting with polynomial weight decay. Furthermore, we exploit sample transferability via maximum mean discrepancy based sample-level re-weighting strategy to diminish the domain gap. The sample transferability exploitation spotlights samples with higher importance to the adaptation process in each domain, hence enhance the adaptation performance. By combining the gradual domain adaptation with the sample transferability exploitation, our method achieves the state-of-the-art performance on transferring between two common person re-id datasets. Zong Guo, Bingpeng Ma, Hong Chang 0001, Xilin Chen 0001 |
ICME | 2 |
| 2022 | Learning Continuous Graph Structure with Bilevel Programming for Graph Neural NetworksabstractLearning graph structure for graph neural networks (GNNs) is crucial to facilitate the GNN-based downstream learning tasks. It is challenging due to the non-differentiable discrete graph structure and lack of ground-truth. In this paper, we address these problems and propose a novel graph structure learning framework for GNNs. Firstly, we directly model the continuous graph structure with dual-normalization, which implicitly imposes sparse constraint and reduces the influence of noisy edges. Secondly, we formulate the whole training process as a bilevel programming problem, where the inner objective is to optimize the GNNs given learned graphs, while the outer objective is to optimize the graph structure to minimize the generalization error of downstream task. Moreover, for bilevel optimization, we propose an improved Neumann-IFT algorithm to obtain an approximate solution, which is more stable and accurate than existing optimization methods. Besides, it makes the bilevel optimization process memory-efficient and scalable to large graphs. Experiments on node classification and scene graph generation show that our method can outperform related methods, especially with noisy graphs. Minyang Hu, Hong Chang 0001, Bingpeng Ma, Shiguang Shan |
IJCAI | 3 |
| 2022 | Optimal Positive Generation via Latent Transformation for Contrastive LearningabstractContrastive learning, which learns to contrast positive with negative pairs of samples, has been popular for self-supervised visual representation learning. Although great effort has been made to design proper positive pairs through data augmentation, few works attempt to generate optimal positives for each instance. Inspired by semantic consistency and computational advantage in latent space of pretrained generative models, this paper proposes to learn instance-specific latent transformations to generate Contrastive Optimal Positives (COP-Gen) for self-supervised contrastive learning. Specifically, we formulate COP-Gen as an instance-specific latent space navigator which minimizes the mutual information between the generated positive pair subject to the semantic consistency constraint. Theoretically, the learned latent transformation creates optimal positives for contrastive learning, which removes as much nuisance information as possible while preserving the semantics. Empirically, using generated positives by COP-Gen consistently outperforms other latent transformation methods and even real-image-based methods in self-supervised contrastive learning. Yinqi Li 0001, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
NeurIPS | 3 |
| 2022 | Extending generalized unsupervised manifold alignment
Xiaoyi Yin, Zhen Cui 0001, Hong Chang 0001, Bingpeng Ma, Shiguang Shan |
Sci. China Inf. Sci. | 4 |
| 2022 | Feature Completion for Occluded Person Re-IdentificationabstractPerson re-identification (reID) plays an important role in computer vision. However, existing methods suffer from performance degradation in occluded scenes. In this work, we propose an occlusion-robust block, Region Feature Completion (RFC), for occluded reID. Different from most previous works that discard the occluded regions, RFC block can recover the semantics of occluded regions in feature space. First, a Spatial RFC (SRFC) module is developed. SRFC exploits the long-range spatial contexts from non-occluded regions to predict the features of occluded regions. The unit-wise prediction task leads to an encoder/decoder architecture, where the region-encoder models the correlation between non-occluded and occluded region, and the region-decoder utilizes the spatial correlation to recover occluded region features. Second, we introduce Temporal RFC (TRFC) module which captures the long-term temporal contexts to refine the prediction of SRFC. RFC block is lightweight, end-to-end trainable and can be easily plugged into existing CNNs to form RFCnet. Extensive experiments are conducted on occluded and commonly holistic reID benchmarks. Our method significantly outperforms existing methods on the occlusion datasets, while remains top even superior performance on holistic datasets. The source code is available at https://github.com/blue-blue272/OccludedReID-RFCnet. Ruibing Hou, Bingpeng Ma, Hong Chang 0001, Xinqian Gu, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | SANet: Statistic Attention Network for Video-Based Person Re-IdentificationabstractCapturing long-range dependencies during feature extraction is crucial for video-based person re-identification (re-id) since it would help to tackle many challenging problems such as occlusion and dramatic pose variation. Moreover, capturing subtle differences, such as bags and glasses, is indispensable to distinguish similar pedestrians. In this paper, we propose a novel and efficacious Statistic Attention (SA) block which can capture both the long-range dependencies and subtle differences. SA block leverages high-order statistics of feature maps, which contain both long-range and high-order information. By modeling relations with these statistics, SA block can explicitly capture long-range dependencies with less time complexity. In addition, high-order statistics usually concentrate on details of feature maps and can perceive the subtle differences between pedestrians. In this way, SA block is capable of discriminating pedestrians with subtle differences. Furthermore, this lightweight block can be conveniently inserted into existing deep neural networks at any depth to form Statistic Attention Network (SANet). To evaluate its performance, we conduct extensive experiments on two challenging video re-id datasets, showing that our SANet outperforms the state-of-the-art methods. Furthermore, to show the generalizability of SANet, we evaluate it on three image re-id datasets and two more general image classification datasets, including ImageNet. The source code is available athttp://vipl.ict.ac.cn/resources/codes/code/SANet_code.zip. Shutao Bai, Bingpeng Ma, Hong Chang 0001, Rui Huang 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | PRDP: Person Reidentification With Dirty and Poor DataabstractIn this article, we propose a novel method to simultaneously solve the data problem of dirty quality and poor quantity for person reidentification (ReID). Dirty quality refers to the wrong labels in image annotations. Poor quantity means that some identities have very few images (FewIDs). Training with these mislabeled data or FewIDs with triplet loss will lead to low generalization performance. To solve the label error problem, we propose a weighted label correction based on cross-entropy (wLCCE) strategy. Specifically, according to the influence range of the wrong labels, we first classify the mislabeled images into point label error and set label error. Then, we propose a weighted triplet loss (WTL) to correct the two label errors, respectively. To alleviate the poor quantity issue, we propose a feature simulation based on autoencoder (FSAE) method to generate some virtual samples for FewID. For the authenticity of the simulated features, we transfer the difference pattern of identities with multiple images (MultIDs) to FewIDs by training an autoencoder (AE)-based simulator. In this way, the FewIDs obtain richer expressions to distinguish from other identities. By dealing with a dirty and poor data problem, we can learn more robust ReID models using the triplet loss. We conduct extensive experiments on two public person ReID datasets: 1) Market-1501 and 2) DukeMTMC-reID, to verify the effectiveness of our approach. Furong Xu, Bingpeng Ma, Hong Chang 0001, Shiguang Shan |
IEEE Trans. Cybern. | 2 |
| 2022 | Motion Feature Aggregation for Video-Based Person Re-IdentificationabstractMost video-based person re-identification (re-id) methods only focus on appearance features but neglect motion features. In fact, motion features can help to distinguish the target persons that are hard to be identified only by appearance features. However, most existing temporal information modeling methods cannot extract motion features effectively or efficiently for v ideo-based re-id. In this paper, we propose a more efficient Motion Feature Aggregation (MFA) method to model and aggregate motion information in the feature map level for video-based re-id. The proposed MFA consists of (i) a coarse-grained motion learning module, which extracts coarse-grained motion features based on the position changes of body parts over time, and (ii) a fine-grained motion learning module, which extracts fine-grained motion features based on the appearance changes of body parts over time. These two modules can model motion information from different granularities and are complementary to each other. It is easy to combine the proposed method with existing network architectures for end-to-end training. Extensive experiments on four widely used datasets demonstrate that the motion features extracted by MFA are crucial complements to appearance features for video-based re-id, especially for the scenario with large appearance changes. Besides, the results on LS-VID, the current largest publicly available video-based re-id dataset, surpass the state-of-the-art methods by a large margin. The code is available at: https://github.com/guxinqian/Simple-ReID. Xinqian Gu, Hong Chang 0001, Bingpeng Ma, Shiguang Shan |
IEEE Trans. Image Process. | 3 |
| 2022 | Interactive Regression and Classification for Dense Object DetectorabstractIn object detection, enhancing feature representation using localization information has been revealed as a crucial procedure to improve detection performance. However, the localization information (i.e., regression feature and regression offset) captured by the regression branch is still not well utilized. In this paper, we propose a simple but effective method called Interactive Regression and Classification (IRC) to better utilize localization information. Specifically, we propose Feature Aggregation Module (FAM) and Localization Attention Module (LAM) to leverage localization information to the classification branch during forward propagation. Furthermore, the classifier also guides the learning of the regression branch during backward propagation, to guarantee that the localization information is beneficial to both regression and classification. Thus, the regression and classification branches are learned in an interactive manner. Our method can be easily integrated into anchor-based and anchor-free object detectors without increasing computation cost. With our method, the performance is significantly improved on many popular dense object detectors, including RetinaNet, FCOS, ATSS, PAA, GFL, GFLV2, OTA, GA-RetinaNet, RepPoints, BorderDet and VFNet. Based on ResNet-101 backbone, IRC achieves 47.2% AP on COCO test-dev, surpassing the previous state-of-the-art PAA (44.8% AP), GFL (45.0% AP) and without sacrificing the efficiency both in training and inference. Moreover, our best model (Res2Net-101-DCN) can achieve a single-model single-scale AP of 51.4%. Linmao Zhou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan |
IEEE Trans. Image Process. | 3 |
| 2021 | BiCnet-TKS: Learning Efficient Spatial-Temporal Representation for Video Person Re-IdentificationabstractIn this paper, we present an efficient spatial-temporal representation for video person re-identification (reID). Firstly, we propose a Bilateral Complementary Network (BiCnet) for spatial complementarity modeling. Specifically, BiCnet contains two branches. Detail Branch processes frames at original resolution to preserve the detailed visual clues, and Context Branch with a down-sampling strategy is employed to capture long-range contexts. On each branch, BiCnet appends multiple parallel and diverse attention modules to discover divergent body parts for consecutive frames, so as to obtain an integral characteristic of target identity. Furthermore, a Temporal Kernel Selection (TKS) block is designed to capture short-term as well as long-term temporal relations by an adaptive mode. TKS can be inserted into BiCnet at any depth to construct BiCnet-TKS for spatial-temporal modeling. Experimental results on multiple benchmarks show that BiCnet-TKS outperforms state-of-the-arts with about 50% less computations. The source code is available at https://github.com/blue-blue272/BiCnet-TKS. Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Rui Huang 0001, Shiguang Shan |
CVPR | 3 |
| 2021 | Enhancing Latent Features for Unsupervised Video Anomaly Detection
Linmao Zhou, Hong Chang 0001, Nan Kang, Xiangjun Zhao, Bingpeng Ma |
PRCV (2) | 5 |
| 2021 | Learning efficient text-to-image synthesis via interstage cross-sample similarity distillation
Fengling Mao, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
Sci. China Inf. Sci. | 2 |
| 2021 | Cross-Modal Knowledge Adaptation for Language-Based Person SearchabstractIn this paper, we present a method named Cross-Modal Knowledge Adaptation (CMKA) for language-based person search. We argue that the image and text information are not equally important in determining a person's identity. In other words, image carries image-specific information such as lighting condition and background, while text contains more modal agnostic information that is more beneficial to cross-modal matching. Based on this consideration, we propose CMKA to adapt the knowledge of image to the knowledge of text. Specially, text-to-image guidance is obtained at different levels: individuals, lists, and classes. By combining these levels of knowledge adaptation, the image-specific information is suppressed, and the common space of image and text is better constructed. We conduct experiments on the CUHK-PEDES dataset. The experimental results show that the proposed CMKA outperforms the state-of-the-art methods. Rui Huang 0001, Hong Chang 0001, Chuanqi Tan, Bingpeng Ma |
IEEE Trans. Image Process. | 6 |
| 2021 | Location Sensitive Network for Human Instance SegmentationabstractLocation is an important distinguishing information for instance segmentation. In this paper, we propose a novel model, called Location Sensitive Network (LSNet), for human instance segmentation. LSNet integrates instance-specific location information into one-stage segmentation framework. Specifically, in the segmentation branch, Pose Attention Module (PAM) encodes the location information into the attention regions through coordinates encoding. Based on the location information provided by PAM, the segmentation branch is able to effectively distinguish instances in feature-level. Moreover, we propose a combination operation named Keypoints Sensitive Combination (KSCom) to utilize the location information from multiple sampling points. These sampling points construct the points representation for instances via human keypoints and random points. Human keypoints provide the spatial locations and semantic information of the instances, and random points expand the receptive fields. Based on the points representation for each instance, KSCom effectively reduces the mis-classified pixels. Our method is validated by the experiments on public datasets. LSNet-5 achieves 56.2 mAP at 18.5 FPS on COCOPersons. Besides, the proposed method is significantly superior to its peers in the case of severe occlusion. Xiangzhou Zhang, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | IAUnet: Global Context-Aware Feature Learning for Person ReidentificationabstractPerson reidentification (reID) by convolutional neural network (CNN)-based networks has achieved favorable performance in recent years. However, most of existing CNN-based methods do not take full advantage of spatial-temporal context modeling. In fact, the global spatial-temporal context can greatly clarify local distractions to enhance the target feature representation. To comprehensively leverage the spatial-temporal context information, in this work, we present a novel block, interaction-aggregation-update (IAU), for high-performance person reID. First, the spatial-temporal IAU (STIAU) module is introduced. STIAU jointly incorporates two types of contextual interactions into a CNN framework for target feature learning. Here, the spatial interactions learn to compute the contextual dependencies between different body parts of a single frame, while the temporal interactions are used to capture the contextual dependencies between the same body parts across all frames. Furthermore, a channel IAU (CIAU) module is designed to model the semantic contextual interactions between channel features to enhance the feature representation, especially for small-scale visual cues and body parts. Therefore, the IAU block enables the feature to incorporate the globally spatial, temporal, and channel context. It is lightweight, end-to-end trainable, and can be easily plugged into existing CNNs to form IAUnet. The experiments show that IAUnet performs favorably against state of the art on both image and video reID tasks and achieves compelling results on a general object categorization task. The source code is available at https://github.com/blue-blue272/ImgReID-IAnet. Ruibing Hou, Bingpeng Ma, Hong Chang 0001, Xinqian Gu, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | TCTS: A Task-Consistent Two-Stage Framework for Person SearchabstractThe state of the art person search methods separate person search into detection and re-ID stages, but ignore the consistency between these two stages. The general person detector has no special attention on the query target; The re-ID model is trained on hand-drawn bounding boxes which are not available in person search. To address the consistency problem, we introduce a Task-Consist Two-Stage (TCTS) person search framework, includes an identity-guided query (IDGQ) detector and a Detection Results Adapted (DRA) re-ID model. In the detection stage, the IDGQ detector learns an auxiliary identity branch to compute query similarity scores for proposals. With consideration of the query similarity scores and foreground score, IDGQ produces query-like bounding boxes for the re-ID stage. In the re-ID stage, we predict identity labels of detected bounding boxes, and use these examples to construct a more practical mixed train set for the DRA model. Training on the mixed train set improves the robustness of the re-ID stage to inaccurate detection. We evaluate our method on two benchmark datasets, CUHK-SYSU and PRW. Our framework achieves 93.9% of mAP and 95.1% of rank1 accuracy on CUHK-SYSU, outperforming the previous state of the art methods. Cheng Wang 0043, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2020 | Appearance-Preserving 3D Convolution for Video-Based Person Re-identification
Xinqian Gu, Hong Chang 0001, Bingpeng Ma, Xilin Chen 0001 |
ECCV (2) | 3 |
| 2020 | Temporal Complementary Learning for Video Person Re-identification
Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
ECCV (25) | 3 |
| 2020 | Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training
Hong Chang 0001, Bingpeng Ma, Naiyan Wang, Xilin Chen 0001 |
ECCV (15) | 3 |
| 2020 | Part alignment network for vehicle re-identification
Bingpeng Ma, Hong Chang 0001 |
Neurocomputing | 2 |
| 2020 | Isosceles Constraints for Person Re-IdentificationabstractIn the existing works of person re-identification (ReID), batch hard triplet loss has achieved great success. However, it only cares about the hardest samples within the batch. For any probe, there are massive mismatched samples (crucial samples) outside the batch which are closer than the matched samples. To reduce the disruptive influence of crucial samples, we propose a novel isosceles contraint for triplet. Theoretically, we show that if a matched pair has equal distance to any one of mismatched sample, the matched pair should be infinitely close. Motivated by this, the isosceles constraint is designed for the two mismatched pairs of each triplet, to restrict some matched pairs with equal distance to different mismatched samples. Meanwhile, to ensure that the distance of mismatched pairs are larger than the matched pairs, margin constraints are necessary. Minimizing the isosceles and margin constraints with respect to the feature extraction network makes the matched pairs closer and the mismatched pairs farther away than the matched ones. By this way, crucial samples are effectively reduced and the performance on ReID is improved greatly. Likewise, our isosceles contraint can be applied to quadruplet as well. Comprehensive experimental evaluations on Market-1501, DukeMTMC-reID and CUHK03 datasets demonstrate the advantages of our isosceles constraint over the related state-of-the-art approaches. Furong Xu, Bingpeng Ma, Hong Chang 0001, Shiguang Shan |
IEEE Trans. Image Process. | 2 |
| 2019 | MS-GAN: Text to Image Synthesis with Attention-Modulated Generators and Similarity-aware Discriminators
Fengling Mao, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
BMVC | 2 |
| 2019 | Cascade RetinaNet: Maintaining Consistency for Single-Stage Object Detection
Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
BMVC | 3 |
| 2019 | Relation-aware Multiple Attention Siamese Networks for Robust Visual Tracking
Fangyi Zhang, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
BMVC | 2 |
| 2019 | VRSTC: Occlusion-Free Video Person Re-IdentificationabstractVideo person re-identification (re-ID) plays an important role in surveillance video analysis. However, the performance of video re-ID degenerates severely under partial occlusion. In this paper, we propose a novel network, called Spatio-Temporal Completion network (STCnet), to explicitly handle partial occlusion problem. Different from most previous works that discard the occluded frames, STCnet can recover the appearance of the occluded parts. For one thing, the spatial structure of a pedestrian frame can be used to predict the occluded body parts from the unoccluded body parts of this frame. For another, the temporal patterns of pedestrian sequence provide important clues to generate the contents of occluded parts. With the spatio-temporal information, STCnet can recover the appearance for the occluded parts, which could be leveraged with those unoccluded parts for more accurate video re-ID. By combining a re-ID network with STCnet, a video re-ID framework robust to partial occlusion (VRSTC) is proposed. Experiments on three challenging video re-ID databases demonstrate that the proposed approach outperforms the state-of-the-arts. Ruibing Hou, Bingpeng Ma, Hong Chang 0001, Xinqian Gu, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2019 | Interaction-And-Aggregation Network for Person Re-IdentificationabstractPerson re-identification (reID) benefits greatly from deep convolutional neural networks (CNNs) which learn robust feature embeddings. However, CNNs are inherently limited in modeling the large variations in person pose and scale due to their fixed geometric structures. In this paper, we propose a novel network structure, Interaction-and-Aggregation (IA), to enhance the feature representation capability of CNNs. Firstly, Spatial IA (SIA) module is introduced. It models the interdependencies between spatial features and then aggregates the correlated features corresponding to the same body parts. Unlike CNNs which extract features from fixed rectangle regions, SIA can adaptively determine the receptive fields according to the input person pose and scale. Secondly, we introduce Channel IA (CIA) module which selectively aggregates channel features to enhance the feature representation, especially for small-scale visual cues. Further, IA network can be constructed by inserting IA blocks into CNNs at any depth. We validate the effectiveness of our model for person reID by demonstrating its superiority over state-of-the-art methods on three benchmark datasets. Ruibing Hou, Bingpeng Ma, Hong Chang 0001, Xinqian Gu, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2019 | Video Prediction with Bidirectional Constraint NetworkabstractFuture frame prediction in videos is promising avenue for unsupervised video representation learning. However video prediction has the huge solution space since the high-dimensionality and inherent uncertainty of the future video frames. Existing approaches impose weak constraints on the predictions, which results in motion confusion. To alleviate this problem, we propose a novel model named Bidirectional Constraint Network (BCnet). BCnet consists of forward prediction module and backward prediction module. The forward prediction module learns to predict the future sequence from the present sequence, while the backward prediction module learns to invert the task. The closed loop of the two modules allows that the backward prediction module generates informative feedback signals. The feedback signals clamp down the solution space of forward prediction module. Therefore, our approach can effectively alleviate the motion confusion. We further evaluate BCnet by fine-tuning it for a supervised learning problem: human action recognition on the UCF-101 dataset. We show that the representation help improve classification accuracy. Extensive experiments on several challenging public datasets show that our approach significantly outperforms state-of-the-art approaches, which demonstrates the effectiveness and generalization ability of our approach. Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Xilin Chen 0001 |
FG | 3 |
| 2019 | Temporal Knowledge Propagation for Image-to-Video Person Re-IdentificationabstractIn many scenarios of Person Re-identification (Re-ID), the gallery set consists of lots of surveillance videos and the query is just an image, thus Re-ID has to be conducted between image and videos. Compared with videos, still person images lack temporal information. Besides, the information asymmetry between image and video features increases the difficulty in matching images and videos. To solve this problem, we propose a novel Temporal Knowledge Propagation (TKP) method which propagates the temporal knowledge learned by the video representation network to the image representation network. Specifically, given the input videos, we enforce the image representation network to fit the outputs of video representation network in a shared feature space. With back propagation, temporal knowledge can be transferred to enhance the image features and the information asymmetry problem can be alleviated. With additional classification and integrated triplet losses, our model can learn expressive and discriminative image and video features for image-to-video re-identification. Extensive experiments demonstrate the effectiveness of our method and the overall results on two widely used datasets surpass the state-of-the-art methods by a large margin. Xinqian Gu, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2019 | Attribute-Aware Pedestrian Image Editing
Xiaoyi Yin, Xinqian Gu, Hong Chang 0001, Bingpeng Ma, Xilin Chen 0001 |
ICIG (1) | 4 |
| 2019 | Cross Attention Network for Few-shot ClassificationabstractFew-shot classification aims to recognize unlabeled samples from unseen classes given only few labeled samples. The unseen classes and low-data problem make few-shot classification very challenging. Many existing approaches extracted features from labeled and unlabeled samples independently, as a result, the features are not discriminative enough. In this work, we propose a novel Cross Attention Network to address the challenging problems in few-shot classification. Firstly, Cross Attention Module is introduced to deal with the problem of unseen classes. The module generates cross attention maps for each pair of class feature and query sample feature so as to highlight the target object regions, making the extracted feature more discriminative. Secondly, a transductive inference algorithm is proposed to alleviate the low-data problem, which iteratively utilizes the unlabeled query set to augment the support set, thereby making the class features more representative. Extensive experiments on two benchmarks show our method is a simple, effective and computationally efficient framework and outperforms the state-of-the-arts. Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
NeurIPS | 3 |
| 2019 | Unifying Visual Attribute Learning with Object Recognition in a Multiplicative FrameworkabstractAttributes are mid-level semantic properties of objects. Recent research has shown that visual attributes can benefit many typical learning problems in computer vision community. However, attribute learning is still a challenging problem as the attributes may not always be predictable directly from input images and the variation of visual attributes is sometimes large across categories. In this paper, we propose a unified multiplicative framework for attribute learning, which tackles the key problems. Specifically, images and category information are jointly projected into a shared feature space, where the latent factors are disentangled and multiplied to fulfil attribute prediction. The resulting attribute classifier is category-specific instead of being shared by all categories. Moreover, our model can leverage auxiliary data to enhance the predictive ability of attribute classifiers, which can reduce the effort of instance-level attribute annotation to some extent. By integrated into an existing deep learning framework, our model can both accurately predict attributes and learn efficient image representations. Experimental results show that our method achieves superior performance on both instance-level and category-level attribute prediction. For zero-shot learning based on visual attributes and human-object interaction recognition, our method can improve the state-of-the-art performance on several widely used datasets. Kongming Liang, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Style Transfer with Adversarial Learning for Cross-Dataset Person Re-identification
Furong Xu, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (6) | 2 |
| 2018 | Continuity-Discrimination Convolutional Neural Network for Visual Object TrackingabstractThis paper proposes a novel model, named Continuity-Discrimination Convolutional Neural Network (CD-CNN), for visual object tracking. Existing state-of-the-art tracking methods do not deal with temporal relationship in video sequences, which leads to imperfect feature representations. To address this problem, CD-CNN models temporal appearance continuity based on the idea of temporal slowness. Mathematically, we prove that, by introducing temporal appearance continuity into tracking, the upper bound of target appearance representation error can be sufficiently small with high probability. Further, in order to alleviate inaccurate target localization and drifting, we propose a novel notion, object-centroid, to characterize not only objectness but also the relative position of the target within a given patch. Both temporal appearance continuity and object-centroid are jointly learned during offline training and then transferred for online tracking. We evaluate our tracker through extensive experiments on two challenging benchmarks and show its competitive tracking performance compared with state-of-the-art trackers. Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
ICME | 2 |
| 2018 | A Benchmark for Full Rotation Head TrackingabstractThis paper introduces a new benchmark for 360-degree rotation head tracking, named Full Rotation Head Tracking (FRHT). The benchmark consists of 50 color sequences containing diverse human activities with complicated head motions. Specially, FRHT covers the most challenges of head tracking and focuses on the appearance variations of heads during the 360-degree rotation. It also pays attention to the clutters from the heads of nearby people. Further, we propose a baseline tracker. It guides a selective adaption updating by verifying strategies, thus alleviates error accumulation. Extensive experiments validate the advantages of FRHT in head rotation and similar object clutter. Bingpeng Ma, Hong Chong, Xilin Chen 0001 |
ICPR | 2 |
| 2018 | Multi-label double-layer learning for cross-modal retrieval
Bingpeng Ma, Shuhui Wang, Yugui Liu, Qingming Huang |
Neurocomputing | 2 |
| 2018 | Generalized Semi-supervised and Structured Subspace Learning for Cross-Modal RetrievalabstractMotivated by the fact that unlabeled data can be easily collected and help to exploit the correlations among different modalities, this paper proposes a novel method named generalized semi-supervised structured subspace learning (GSS-SL) for the task of cross-modal retrieval. First, to predict more relevant class labels for unlabeled data, we propose a label graph constraint that ensures the intrinsic geometric structures of different feature spaces consistent with that of label space. Second, considering that class labels directly reveal the semantic information of multimedia data, GSS-SL takes the label space as a linkage to model the correlations among different modalities. Concretely, the label graph constraint, label-linked loss function, and regularization are integrated into a joint minimization formulation to learn a discriminative common subspace. Finally, an efficient optimization algorithm is designed to alternately optimize multiple linear transformations for different modalities and update the class indicator matrices for unlabeled data. Furthermore, an arbitrary number of modalities can be solved in the proposed framework. Extensive experiments on three standard benchmark datasets demonstrate that GSS-SL outperforms previous methods on exploiting the correlations among different modalities. Bingpeng Ma, Guorong Li, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Siamese recurrent architecture for visual trackingabstractTreating visual tracking as a matching problem, siamese architecture has drawn increasing interest recently. In this paper, we propose a novel siamese recurrent architecture that can enhance the similarity matching by leveraging contextual information. Specifically, the multi-directional Recurrent Neural Network (RNN) is employed to memorize the long-range contextual dependencies of object parts and learn the self-structure information of the object. We test the proposed method on a challenging benchmark, and it gain promising results compared with the existing tracking algorithms. Xiaqing Xu, Bingpeng Ma, Hong Chang 0001, Xilin Chen 0001 |
ICIP | 2 |
| 2017 | Metric based on multi-order spaces for cross-modal retrievalabstractThis paper proposes a novel method for cross-modal retrieval. Different from vector (text)-to-vector (image) framework of the traditional cross-modal methods, we adopt a vector (text)-to-matrix (image) framework. We assume that compared with vectors, matrices can directly represent images and characterize the structure of feature space. Furthermore, we propose a Metric based on Multi-order spaces (MMs). Multi-order statistic features are used to represent images for enriching the semantic information, and metrics among the multi-spaces are jointly learned to measure the similarity between two different modalities. Specifically, there are three steps for MMs. First, we jointly use the bags of visual features (zero-order), mean (first-order) and covariance (second-order) to characterize each image. Second, considering that covariance matrices and vectors lie on a Riemannian manifold and an Euclidean space respectively, we embed multi-order spaces into their corresponding Hilbert spaces to reduce the heterogeneity among the original spaces. Finally, the similarity between two different modalities can be measured by learning multiple transformations from the different Hilbert spaces to a common subspace. The performance of the proposed method over the state-of-the-art has been demonstrated through the experiments on two public datasets. Bingpeng Ma, Guorong Li, Qingming Huang |
ICME | 2 |
| 2017 | Adaptively Unified Semi-supervised Learning for Cross-Modal RetrievalabstractMotivated by the fact that both relevancy of class labels and unlabeled data can help to strengthen multi-modal correlation, this paper proposes a novel method for cross-modal retrieval. To make each sample moving to the direction of its relevant label while far away from that of its irrelevant ones, a novel dragging technique is fused into a unified linear regression model. By this way, not only the relation between embedded features and relevant class labels but also the relation between embedded features and irrelevant class labels can be exploited. Moreover, considering that some unlabeled data contain specific semantic information, a weighted regression model is designed to adaptively enlarge their contribution while weaken that of the unlabeled data with non-specific semantic information. Hence, unlabeled data can supply semantic information to enhance discriminant ability of classifier. Finally, we integrate the constraints into a joint minimization formulation and develop an efficient optimization algorithm to learn a discriminative common subspace for different modalities. Experimental results on Wiki, Pascal and NUS-WIDE datasets show that the proposed method outperforms the state-of-the-art methods even when we set 20% samples without class labels. Bingpeng Ma, Guorong Li, Qingming Huang, Qi Tian 0001 |
IJCAI | 2 |
| 2017 | Multi-Networks Joint Learning for Large-Scale Cross-Modal RetrievalabstractThis paper proposes a novel deep framework of multi-networks joint learning for large-scale cross-modal retrieval. For most existing cross-modal methods, the processes of training and testing don't care about the problem of memory requirement. Hence, they are generally implemented on small-scale data. Moreover, they take feature learning and latent space embedding as two separate steps which cannot generate specific features to accord with the cross-modal task. To alleviate the problems, we first disintegrate the multiplication and inverse of some big matrices, usually involved in existing methods, into that of many sub-matrices. Each sub-matrix is targeted to dispose one pair of image-sentence, for which we further design a novel sampling strategy to select the most representative samples to construct the cross-modal ranking loss and within-modal discriminant loss functions. By this way, the proposed model consumes less memory each time such that it can scale to large-scale data. Furthermore, we apply the proposed discriminative ranking loss to effectively unify two heterogenous networks, deep residual network for images and long short-term memory for sentences, into an end-to-end deep learning architecture. Finally, we can simultaneously achieve specific features adapting to cross-modal task and learn a shared latent space for images and sentences. Extensive evaluations on two large-scale cross-modal datasets show that the proposed method brings substantial improvements over other state-of-the-art ranking methods. Bingpeng Ma, Guorong Li, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2017 | Updating initial labels from spectral graph by manifold regularization for saliency detection
Jiazhong Chen, Bingpeng Ma, Hua Cao, Jie Chen 0058, Yebin Fan |
Neurocomputing | 2 |
| 2017 | Attention region detection based on closure prior in layered bit Planes
Jiazhong Chen, Bingpeng Ma, Hua Cao, Jie Chen 0058, Yebin Fan |
Neurocomputing | 2 |
| 2017 | Video-Based Pedestrian Re-Identification by Adaptive Spatio-Temporal Appearance ModelabstractPedestrian re-identification is a difficult problem due to the large variations in a person's appearance caused by different poses and viewpoints, illumination changes, and occlusions. Spatial alignment is commonly used to address these issues by treating the appearance of different body parts independently. However, a body part can also appear differently during different phases of an action. In this paper, we consider the temporal alignment problem, in addition to the spatial one, and propose a new approach that takes the video of a walking person as input and builds a spatiotemporal appearance representation for pedestrian re-identification. Particularly, given a video sequence, we exploit the periodicity exhibited by a walking person to generate a spatiotemporal body-action model, which consists of a series of body-action units corresponding to certain action primitives of certain body parts. Fisher vectors are learned and extracted from individual body-action units and concatenated into the final representation of the walking person. Unlike previous spatiotemporal features that only take into account local dynamic appearance information, our representation aligns the spatiotemporal appearance of a pedestrian globally. Extensive experiments on public data sets show the effectiveness of our approach compared with the state of the art. Wei Zhang 0021, Bingpeng Ma, Kan Liu 0001, Rui Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Cross-Modal Retrieval Using Multiordered Discriminative Structured Subspace LearningabstractThis paper proposes a novel method for cross-modal retrieval. In addition to the traditional vector (text)-to-vector (image) framework, we adopt a matrix (text)-to-matrix (image) framework to faithfully characterize the structures of different feature spaces. Moreover, we propose a novel metric learning framework to learn a discriminative structured subspace, in which the underlying data distribution is preserved for ensuring a desirable metric. Concretely, there are three steps for the proposed method. First, the multiorder statistics are used to represent images and texts for enriching the feature information. We jointly use the covariance (second-order), mean (first-order), and bags of visual (textual) features (zeroth-order) to characterize each image and text. Second, considering that the heterogeneous covariance matrices lie on the different Riemannian manifolds and the other features on the different Euclidean spaces, respectively, we propose a unified metric learning framework integrating multiple distance metrics, one for each order statistical feature. This framework preserves the underlying data distribution and exploits complementary information for better matching heterogeneous data. Finally, the similarity between the different modalities can be measured by transforming the multiorder statistical features to the common subspace. The performance of the proposed method over the previous methods has been demonstrated through the experiments on two public datasets. Bingpeng Ma, Guorong Li, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2016 | Deep Second-Order Siamese Network for Pedestrian Re-identification
Xuesong Deng, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (2) | 2 |
| 2016 | Cross-modal Retrieval by Real Label Partial Least SquaresabstractThis paper proposes a novel method named Real Label Partial Least Squares (RL-PLS) for the task of cross-modal retrieval. Pervious works just take the texts and images as two modalities in PLS. But in RL-PLS, considering that the class label is more related to the semantics directly, we take the class label as the assistant modality. Specially, we build two KPLS models and project both images and texts into the label space. Then, the similarity of images and texts can be measured more accurately in the label space. Furthermore, we do not restrict the label indicator values as the binary values as the traditional methods. By contraries, in RL-PLS, the label indicator values are set to the real values. Specially, the label indicator values are comprised by two parts: positive or negative represents the sample class while the absolute value represents the local structure in the class. By this way, the discriminate ability of RL-PLS is improved greatly. To show the effectiveness of RL-PLS, the experiments are conducted on two cross-modal retrieval tasks (Wiki and Pascal Voc2007), on which the competitive results are obtained. Bingpeng Ma, Shuhui Wang, Yugui Liu, Qingming Huang |
ACM Multimedia | 2 |
| 2016 | PL-ranking: A Novel Ranking Method for Cross-Modal RetrievalabstractThis paper proposes a novel method for cross-modal retrieval named Pairwise-Listwise \textbf{ranking} (PL-ranking) based on the low-rank optimization framework. Motivated by the fact that optimizing the top of ranking is more applicable in practice, we focus on improving the precision at the top of ranked list for a given sample and learning a low-dimensional common subspace for multi-modal data. Concretely, there are three constraints in PL-ranking. First, we use a pairwise ranking loss constraint to optimize the top of ranking. Then, considering that the pairwise ranking loss constraint ignores class information, we further adopt a listwise constraint to minimize the intra-neighbors variance and maximize the inter-neighbors separability. By this way, class information is preserved while the number of iterations is reduced. Finally, low-rank based regularization is applied to exploit the correlations between features and labels so that the relevance between the different modalities can be enhanced after mapping them into the common subspace. We design an efficient low-rank stochastic subgradient descent method to solve the proposed optimization problem. The experimental results show that the average MAP scores of PL-ranking are improved 5.1%, 9.2%, 4.7% and 4.8% than those of the state-of-the-art methods on the Wiki, Flickr, Pascal and NUS-WIDE datasets, respectively. Bingpeng Ma, Guorong Li, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2016 | Beyond appearance model: Learning appearance variations for object tracking
Guorong Li, Bingpeng Ma, Jun Huang 0003, Qingming Huang, Weigang Zhang |
Neurocomputing | 2 |
| 2015 | A Spatio-Temporal Appearance Representation for Viceo-Based Pedestrian Re-IdentificationabstractPedestrian re-identification is a difficult problem due to the large variations in a person's appearance caused by different poses and viewpoints, illumination changes, and occlusions. Spatial alignment is commonly used to address these issues by treating the appearance of different body parts independently. However, a body part can also appear differently during different phases of an action. In this paper we consider the temporal alignment problem, in addition to the spatial one, and propose a new approach that takes the video of a walking person as input and builds a spatio-temporal appearance representation for pedestrian re-identification. Particularly, given a video sequence we exploit the periodicity exhibited by a walking person to generate a spatio-temporal body-action model, which consists of a series of body-action units corresponding to certain action primitives of certain body parts. Fisher vectors are learned and extracted from individual body-action units and concatenated into the final representation of the walking person. Unlike previous spatio-temporal features that only take into account local dynamic appearance information, our representation aligns the spatio-temporal appearance of a pedestrian globally. Extensive experiments on public datasets show the effectiveness of our approach compared with the state of the art. Kan Liu 0001, Bingpeng Ma, Wei Zhang 0021, Rui Huang 0001 |
ICCV | 2 |
| 2015 | Set-label modeling and deep metric learning on person re-identification
Hao Liu 0019, Bingpeng Ma, Junbiao Pang, Chunjie Zhang 0001, Qingming Huang |
Neurocomputing | 2 |
| 2015 | VoD: A novel image representation for head yaw estimation
Bingpeng Ma, Rui Huang 0001 |
Neurocomputing | 1 |
| 2014 | Person Search in a Scene by Jointly Modeling People Commonness and Person UniquenessabstractThis paper presents a novel framework for a multimedia search task: searching a person in a scene using human body appearance. Existing works mostly focus on two independent problems related to this task, i.e., people detection and person re-identification. However, a sequential combination of these two components does not solve the person search problem seamlessly for two reasons: 1) the errors in people detection are carried into person re-identification unavoidably; 2) the setting of person re-identification is different from that of person search which is essentially a verification problem. To bridge this gap, we propose a unified framework which jointly models the commonness of people (for detection) and the uniqueness of a person (for identification). We demonstrate superior performance of our approach on public benchmarks compared with the sequential combination of the state-of-the-art detection and identification algorithms. Yuanlu Xu, Bingpeng Ma, Rui Huang 0001, Liang Lin 0004 |
ACM Multimedia | 2 |
| 2014 | Joint sparse representation for video-based face recognition
Zhen Cui 0001, Hong Chang 0001, Shiguang Shan, Bingpeng Ma, Xilin Chen 0001 |
Neurocomputing | 4 |
| 2014 | CovGa: A novel descriptor based on symmetry of regions for head pose estimation
Bingpeng Ma, Annan Li, Xiujuan Chai, Shiguang Shan |
Neurocomputing | 1 |
| 2014 | Covariance descriptor based on bio-inspired features for person re-identification and face verification
Bingpeng Ma, Yu Su 0009, Frédéric Jurie |
Image Vis. Comput. | 1 |
| 2013 | A novel feature descriptor based on biologically inspired feature for head pose estimation
Bingpeng Ma, Xiujuan Chai, Tianjiang Wang |
Neurocomputing | 1 |
| 2013 | Accelerated implementation of adaptive directional lifting-based discrete wavelet transform on GPU
Jiazhong Chen, Zengwei Ju, Hua Cao, Bingpeng Ma, Changnian Chen, Leihua Qin |
Signal Process. Image Commun. | 4 |
| 2012 | BiCov: a novel image representation for person re-identification and face verificationabstractInternational audience Bingpeng Ma, Yu Su 0009, Frédéric Jurie |
BMVC | 1 |
| 2010 | Visual Selection and Attention Shifting Based on FitzHugh-Nagumo Equations
Yuanhua Qiao, Lijuan Duan, Faming Fang, Bingpeng Ma |
ISNN (2) | 6 |
| 2008 | Discriminant analysis for perceptionally comparable classesabstractTraditional discriminate analysis treats all the involved classes equally in the computation of the between-class scatter matrix. However, we find that for many vision tasks, the classes to be processed are not equal in perception, i.e. a distance metric can be defined between the classes. Typical examples include head pose classification and age estimation. Aiming at this category of classification problem, this paper proposes a novel discriminant analysis method, called Class Distance based Discriminant Analysis (CDDA). In CDDA, the perceptional distance between two classes is exploited to weight the outer product in the between-class scatter computation, to concentrate more on the classes difficult to separate. Another novelty of CDDA is that to preserve the within-class local structure of multimodal labeled data, the within-class scatter is re-defined by complementing the similarity of the samples pairs in the nearby classes. The method is then applied to head pose classification and age estimation problem, and experimental results demonstrate the effectiveness of CDDA. Bingpeng Ma, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
FG | 1 |
| 2008 | Head Yaw Estimation From Asymmetry of Facial AppearanceabstractThis paper proposes a novel method to estimate the head yaw rotations based on the asymmetry of 2-D facial appearance. In traditional appearance-based pose estimation methods, features are typically extracted holistically by subspace analysis such as principal component analysis, linear discriminant analysis (LDA), etc., which are not designed to directly model the pose variations. In this paper, we argue and reveal that the asymmetry in the intensities of each row of the face image is closely relevant to the yaw rotation of the head and, at the same time, evidently insensitive to the identity of the input face. Specifically, to extract the asymmetry information, 1-D Gabor filters and Fourier transform are exploited. LDA is further applied to the asymmetry features to enhance the discrimination ability. By using the simple nearest centroid classifier, experimental results on two multipose databases show that the proposed features outperform other features. In particular, the generalization of the proposed asymmetry features is verified by the impressive performance when the training and the testing data sets are heterogeneous. Bingpeng Ma, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2006 | Study on the extensibility of video conferencing with speech mixing
Shutang Yang, Bingpeng Ma |
Comput. Commun. | 2 |