EDBT 2026 Demo / reviewers in the wild / expert
Mingtao Pei
dblp:77/7398
· DBLP profile ↗
70ranked-venue papers
4as first author
24since 2021 · last 2025
0000-0003-4949-7997ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 2 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PrimHOI: Compositional Human-Object Interaction via Reusable Primitives
Tengyu Liu, Yixin Zhu 0001, Mingtao Pei, Siyuan Huang 0001 |
ICCV | 4 |
| 2025 | Trace3D: Consistent Segmentation Lifting via Gaussian Instance Tracing
Hongyu Shen, Junfeng Ni, Weishuo Li, Mingtao Pei |
ICCV | 5 |
| 2025 | Retrieval from Dynamic Phrases: Generating Radiograph Reports with Phrase-Level Template and Dynamic Memory BankabstractMost retrieval-based report generation methods rely on sentence-level templates, which often introduce ambiguities due to similar semantics across different sentences. To overcome this, we propose a phrase-level framework, comprising automatic phrase template extraction and report generation based on retrieval. In the first stage, we introduce a phrase scoring mechanism to evaluate the semantics and importance of phrases, enabling efficient template extraction. In the second stage, we retrieve relevant templates and fuse their features with visual features from the radiograph through a Retrieval-Aggregation strategy. The dynamic update of the template bank during training improves template representations. Experiments on IU X-Ray and MIMIC-CXR datasets demonstrate the effectiveness of our method in generating accurate radiology reports. Haoquan Chen, Hongyu Shen, Mingtao Pei |
IJCNN | 4 |
| 2025 | Align Modalities: Advancing Medical Report Generation with Unified Encoder and Inter-Case Contrastive LearningabstractMedical reports play a pivotal role in achieving accurate diagnoses. This technology not only reduces the burden on radiologists but also fosters consistency in treatment approaches. The crux of generating high-caliber medical reports lies in the model’s ability to interpret and integrate both visual and textual data. However, the inherent distributional disparities between modalities pose a significant challenge to this process. Hence, we propose UEMA framework, which extracts features from both modalities through a Unified Encoder and consists of an ICCL (Inter-Case Contrastive Learning) module to facilitate multimodal alignment. The ICCL module leverages multi-label contrastive learning across different cases to align visual and textual features. Extensive experiments have been conducted on the publicly available IU X-Ray and MIMIC-CXR datasets with additional case studies and visual analysis, demonstrating the effectiveness of our designed module and that our model outperforms state-of-the-art methods across a wide range of metrics. Haoquan Chen, Mingtao Pei, Zhengang Nie |
SMC | 2 |
| 2025 | Let storytelling tell vivid stories: A multi-modal-agent-based unified storytelling framework
Jiji Tang, Chuanqi Zang, Mingtao Pei, Wei Liang 0008, Zeng Zhao |
Neurocomputing | 4 |
| 2025 | A method of embedding a high-resolution image into a large field-of-view image
Yanmei Dong, Mingtao Pei, Yunde Jia |
Multim. Tools Appl. | 2 |
| 2025 | Integrating clinical knowledge and imaging for medical report generation
Juncai Liu, Hongyu Shen, Mingtao Pei |
Pattern Recognit. Lett. | 5 |
| 2024 | Automatic Radiology Reports Generation via Memory Alignment NetworkabstractThe automatic generation of radiology reports is of great significance, which can reduce the workload of doctors and improve the accuracy and reliability of medical diagnosis and treatment, and has attracted wide attention in recent years. Cross-modal mapping between images and text, a key component of generating high-quality reports, is challenging due to the lack of corresponding annotations. Despite its importance, previous studies have often overlooked it or lacked adequate designs for this crucial component. In this paper, we propose a method with memory alignment embedding to assist the model in aligning visual and textual features to generate a coherent and informative report. Specifically, we first get the memory alignment embedding by querying the memory matrix, where the query is derived from a combination of the visual features and their corresponding positional embeddings. Then the alignment between the visual and textual features can be guided by the memory alignment embedding during the generation process. The comparison experiments with other alignment methods show that the proposed alignment method is less costly and more effective. The proposed approach achieves better performance than state-of-the-art approaches on two public datasets IU X-Ray and MIMIC-CXR, which further demonstrates the effectiveness of the proposed alignment method. Hongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing Tian |
AAAI | 2 |
| 2024 | Foreign Object Classification for Coal Conveyor Belts Based on Deep Learning
Siyu Chen 0025, Mingtao Pei |
PRCV (2) | 2 |
| 2024 | Spatio-Temporal Contrastive Learning for Compositional Action Recognition
Yezi Gong, Mingtao Pei |
PRCV (7) | 2 |
| 2024 | Self-trained multi-cues model for video anomaly detection
Zhengang Nie, Wei Liang 0008, Mingtao Pei |
Multim. Tools Appl. | 4 |
| 2024 | Token labeling-guided multi-scale medical image classification
Fangyuan Yan, Wei Liang 0008, Mingtao Pei |
Pattern Recognit. Lett. | 4 |
| 2024 | Task and Environment-Aware Virtual Scene Rearrangement for Enhanced Safety in Virtual RealityabstractEmerging VR applications have revolutionized user experiences by immersing individuals in digitally crafted environments. However, fully immersive experiences introduce new challenges, notably the risk of physical hazards when users are unaware of their surroundings. Existing solutions, including guardian spaces and locomotion systems, present trade-offs that either disrupt the immersive experience or risk inducing motion sickness. To address these challenges, we propose a novel approach that dynamically rearranges VR scenes according to users' physical spaces, seamlessly embedding physical constraints and interaction tasks into the virtual environment. We design a computational model to optimize the rearranged scene through a cost function, ensuring collision-free interactions while maintaining visual fidelity and the goal of interaction tasks. The experiments demonstrate improvements in user experience and safety, presenting an innovative solution to harmonize physical and virtual environments in VR applications. Bing Ning, Mingtao Pei |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Discovering the Real Association: Multimodal Causal Reasoning in Video Question AnsweringabstractVideo Question Answering (VideoQA) is challenging as it requires capturing accurate correlations between modalities from redundant information. Recent methods focus on the explicit challenges of the task, e.g. multimodal feature extraction, video-text alignment and fusion. Their frameworks reason the answer relying on statistical evidence causes, which ignores potential bias in the multimodal data. In our work, we investigate relational structure from a causal representation perspective on multimodal data and propose a novel inference framework. For visual data, question-irrelevant objects may establish simple matching associations with the answer. For textual data, the model prefers the local phrase semantics which may deviate from the global semantics in long sentences. Therefore, to enhance the generalization of the model, we discover the real association by explicitly capturing visual features that are causally related to the question semantics and weakening the impact of local language semantics on question answering. The experimental results on two large causal VideoQA datasets verify that our proposed framework 1) improves the accuracy of the existing VideoQA backbone, 2) demonstrates robustness on complex scenes and questions. The code will be released at https://github.com/Chuanqi-Zang/Discovering-the-Real-Association. Chuanqi Zang, Hanqing Wang 0001, Mingtao Pei, Wei Liang 0008 |
CVPR | 3 |
| 2023 | Face Anti-spoofing Based on Client Identity Information and Depth Map
Mingtao Pei, Zhengang Nie, Xinmu Qi |
ICIG (1) | 2 |
| 2023 | Dual Transformer Encoder Model for Medical Image ClassificationabstractCompared with convolutional neural networks, vision transformer with powerful global modeling abilities has achieved promising results in natural image classification and has been applied in the field of medical image analysis. Vision transformer divides the input image into a token sequence of fixed hidden size and keeps the hidden size constant during training. However, a fixed size is unsuitable for all medical images. To address the above issue, we propose a new dual transformer encoder model which consists of two transformer encoders with different hidden sizes so that the model can be trained with two token sequences with different sizes. In addition, the vision transformer only considers the class token output by the last layer in the encoders when predicting the category, ignoring the information of other layers. We use a Layer-wise Class token Attention (LCA) classification module that leverages class tokens from all layers of encoders to predict categories. Extensive experiments show that our proposed model obtains better performance than other transformer-based methods, which proves the effectiveness of our model. Fangyuan Yan, Mingtao Pei |
ICIP | 3 |
| 2023 | Person-Specific Face Spoofing Detection Based on a Siamese Network
Mingtao Pei, Huiling Hao |
Pattern Recognit. | 1 |
| 2022 | Clinical-BERT: Vision-Language Pre-training for Radiograph Diagnosis and Reports GenerationabstractIn this paper, we propose a vision-language pre-training model, Clinical-BERT, for the medical domain, and devise three domain-specific tasks: Clinical Diagnosis (CD), Masked MeSH Modeling (MMM), Image-MeSH Matching (IMM), together with one general pre-training task: Masked Language Modeling (MLM), to pre-train the model. The CD task helps the model to learn medical domain knowledge by predicting disease from radiographs. Medical Subject Headings (MeSH) words are important semantic components in radiograph reports, and the MMM task helps the model focus on the prediction of MeSH words. The IMM task helps the model learn the alignment of MeSH words with radiographs by matching scores obtained by a two-level sparse attention: region sparse attention and word sparse attention. Region sparse attention generates corresponding visual features for each word, and word sparse attention enhances the contribution of images-MeSH matching to the matching scores. To the best of our knowledge, this is the first attempt to learn domain knowledge during pre-training for the medical domain. We evaluate the pre-training model on Radiograph Diagnosis and Reports Generation tasks across four challenging datasets: MIMIC-CXR, IU X-Ray, COV-CTR, and NIH, and achieve state-of-the-art results for all the tasks, which demonstrates the effectiveness of our pre-training model. Mingtao Pei |
AAAI | 2 |
| 2022 | Do You Live a Healthy Life? Analyzing Lifestyle by Visual Life LoggingabstractIn this work, we investigate the problem of lifestyle analysis and build a visual lifelogging dataset for lifestyle analysis (VLDLA). The VLDLA contains images captured by a wearable camera every 3 seconds from 8:00 am to 6:00 pm for seven days. In contrast to current lifelogging/egocentric datasets, our dataset is suitable for lifestyle analysis as images are taken with short intervals to capture activities of short duration; moreover, images are taken continuously from morning to evening to record all the activities performed by a user. Based on our dataset, we classify the user activities in each frame and use three latent fluents of the user, which change over time and are associated with activities, to measure the healthy degree of the user’s lifestyle. Experimental results show that our method can be used to analyze the healthiness of users’ lifestyles. Mingtao Pei, Hongyu Shen |
ICASSP | 2 |
| 2022 | Predicting Human Motion Using Key SubsequencesabstractHuman motion prediction is an important task in computer vision, and has a wide range of applications, such as autonomous driving and human-robot interaction. Usually, human motion tends to repeat itself and follows patterns that are well-represented by a few short key subsequences. Based on the above observations, we propose an attention-based feed-forward network, which is explicitly guided by the key subsequences, for human motion prediction. Specifically, we obtain the key subsequences by clustering, extract motion attention by the similarity between the observed poses and the motion context of corresponding key subsequences, and aggregate the relevant key subsequences by a graph convolutional network to predict human motion. Experimental results on public human motion datasets show that our method achieves better performance over state-of-the-art methods in motion prediction. Mingtao Pei, Wei Liang 0008 |
ICASSP | 2 |
| 2022 | Few-shot human motion prediction using deformable spatio-temporal CNN with parameter generation
Chuanqi Zang, Mingtao Pei |
Neurocomputing | 3 |
| 2022 | Stitching images from a conventional camera and a fisheye camera based on nonrigid warping
Yanmei Dong, Mingtao Pei, Yuwei Wu 0001, Yunde Jia |
Multim. Tools Appl. | 2 |
| 2022 | Prior Guided Transformer for Accurate Radiology Reports GenerationabstractIn this paper, we propose a prior guided transformer for accurate radiology reports generation. In the encoder part, a radiograph is firstly represented by a set of patch features, which is obtained through a convolutional neural network and a traditional transformer encoder. Then an Additive Gaussian model is applied to represent the prior knowledge based on unsupervised clustering and sparse attention. In the decoder part, prior embeddings are acquired by probabilistically sampling from the radiograph prior. Then the visual features, language embeddings, and prior embeddings are fused by our proposed Prior Guided Attention to generate accurate radiology reports. Experiment results show that our method achieves better performance than state-of-the-art methods on two public radiology datasets, which proves the effectiveness of our prior guided transformer. Mingtao Pei, Caifeng Shan, Zhaoxing Tian |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Laryngoscope8: Laryngeal image dataset and classification of laryngeal disease based on attention mechanism
Mingtao Pei, Jinrang Li, Mukun Wu |
Pattern Recognit. Lett. | 3 |
| 2020 | Few-shot Human Motion Prediction via Learning Novel Motion DynamicsabstractHuman motion prediction is a task where we anticipate future motion based on past observation. Previous approaches rely on the access to large datasets of skeleton data, and thus are difficult to be generalized to novel motion dynamics with limited training data. In our work, we propose a novel approach named Motion Prediction Network (MoPredNet) for few-short human motion prediction. MoPredNet can be adapted to predicting new motion dynamics using limited data, and it elegantly captures long-term dependency in motion dynamics. Specifically, MoPredNet dynamically selects the most informative poses in the streaming motion data as masked poses. In addition, MoPredNet improves its encoding capability of motion dynamics by adaptively learning spatio-temporal structure from the observed poses and masked poses. We also propose to adapt MoPredNet to novel motion dynamics based on accumulated motion experiences and limited novel motion dynamics data. Experimental results show that our method achieves better performance over state-of-the-art methods in motion prediction. Chuanqi Zang, Mingtao Pei, Yu Kong 0001 |
IJCAI | 2 |
| 2020 | Visual-Semantic Graph Matching for Visual GroundingabstractVisual Grounding is the task of associating entities in a natural language sentence with objects in an image. In this paper, we formulate visual grounding as a graph matching problem to find node correspondences between a visual scene graph and a language scene graph. These two graphs are heterogeneous, representing structure layouts of the sentence and image, respectively. We learn unified contextual node representations of the two graphs by using a cross-modal graph convolutional network to reduce their discrepancy. The graph matching is thus relaxed as a linear assignment problem because the learned node representations characterize both node information and structure information. A permutation loss and a semantic cycle-consistency loss are further introduced to solve the linear assignment problem with or without ground-truth correspondences. Experimental results on two visual grounding tasks, i.e., referring expression comprehension and phrase localization, demonstrate the effectiveness of our method. Chenchen Jing, Yuwei Wu 0001, Mingtao Pei, Yao Hu 0002, Yunde Jia, Qi Wu 0001 |
ACM Multimedia | 3 |
| 2020 | Online maximum a posteriori tracking of multiple objects using sequential trajectory prior
Min Yang 0003, Mingtao Pei, Yunde Jia |
Image Vis. Comput. | 2 |
| 2019 | Face Liveness Detection Based on Client Identity Using Siamese Network
Huiling Hao, Mingtao Pei |
PRCV (1) | 2 |
| 2019 | Diffusion-based kernel matrix model for face liveness detection
Changyong Yu, Chengtang Yao, Mingtao Pei, Yunde Jia |
Image Vis. Comput. | 3 |
| 2019 | Heterogeneous Hashing Network for Face Retrieval Across Image and Video DomainsabstractIn this paper, we present a heterogeneous hashing network to generate effective and compact hash representations of both face images and face videos for face retrieval across image and video domains. The network contains an image branch and a video branch to project face images and videos into a common space, respectively. Then, the non-linear hash functions are learned in the common space to obtain the corresponding binary hash representations. The network is trained with three loss functions: 1) the Fisher loss; 2) the softmax loss; and 3) the triplet ranking loss. The Fisher loss uses the difference form of within-class and between-class scatter and is appropriate for the mini-batch-based optimization method. The Fisher loss together with the softmax loss is exploited to enhance the discriminative power of the common space. The triplet ranking loss is enforced on the final binary hash representations to improve retrieval performance. Experiments on a large-scale face video dataset and two challenging TV-series datasets demonstrate the effectiveness of the proposed method. Chenchen Jing, Zhen Dong 0002, Mingtao Pei, Yunde Jia |
IEEE Trans. Multim. | 3 |
| 2018 | Vehicle Re-Identification by Deep Feature Fusion Based on Joint Bayesian CriterionabstractVehicle re-identification is a challenging task as the differences between vehicles of the same model are extremely small. In this paper, we propose to fuse deep features extracted by two different CNNs for vehicle re-identification. CNNs can extract discriminative features for classification tasks. Features extracted by different CNNs describe different aspects of the input image, and are complementary to each other. We propose a new loss function called the Joint Bayesian loss to fuse the different deep features. The proposed Joint Bayesian loss can minimize the intra-class variations and simultaneously maximize the inter-class variations of the fused features, and it is very fit for the vehicle re-identification. Experiments on a large-scale vehicle dataset demonstrate the effectiveness of the proposed method. Mingtao Pei, Leyi Zhu |
ICPR | 2 |
| 2018 | Deep CNN based binary hash video representations for face retrieval
Zhen Dong 0002, Chenchen Jing, Mingtao Pei, Yunde Jia |
Pattern Recognit. | 3 |
| 2017 | Deep Manifold Learning of Symmetric Positive Definite Matrices with Application to Face RecognitionabstractIn this paper, we aim to construct a deep neural network which embeds high dimensional symmetric positive definite (SPD) matrices into a more discriminative low dimensional SPD manifold. To this end, we develop two types of basic layers: a 2D fully connected layer which reduces the dimensionality of the SPD matrices, and a symmetrically clean layer which achieves non-linear mapping. Specifically, we extend the classical fully connected layer such that it is suitable for SPD matrices, and we further show that SPD matrices with symmetric pair elements setting zero operations are still symmetric positive definite. Finally, we complete the construction of the deep neural network for SPD manifold learning by stacking the two layers. Experiments on several face datasets demonstrate the effectiveness of the proposed method. Zhen Dong 0002, Su Jia, Chi Zhang 0063, Mingtao Pei, Yuwei Wu 0001 |
AAAI | 4 |
| 2016 | Face Video Retrieval via Deep Learning of Binary Hash RepresentationsabstractRetrieving faces from large mess of videos is an attractive research topic with wide range of applications. Its challenging problems are large intra-class variations, and tremendous time and space complexity. In this paper, we develop a new deep convolutional neural network (deep CNN) to learn discriminative and compact binary representations of faces for face video retrieval. The network integrates feature extraction and hash learning into a unified optimization framework for the optimal compatibility of feature extractor and hash functions. In order to better initialize the network, the low-rank discriminative binary hashing is proposed to pre-learn hash functions during the training procedure. Our method achieves excellent performances on two challenging TV-Series datasets. Zhen Dong 0002, Su Jia, Tianfu Wu 0001, Mingtao Pei |
AAAI | 4 |
| 2016 | Visual tracking with sparse correlation filtersabstractCorrelation filters have recently made significant improvements in visual object tracking on both efficiency and accuracy. In this paper, we propose a sparse correlation filter, which combines the effectiveness of sparse representation and the computational efficiency of correlation filters. The sparse representation is achieved through solving an ℓ0regularized least squares problem. The obtained sparse correlation filters are able to represent the essential information of the tracked target while being insensitive to noise. During tracking, the appearance of the target is modeled by a sparse correlation filter, and the filter is re-trained after tracking on each frame to adapt to the appearance changes of the target. The experimental results on the CVPR2013 Online Object Tracking Benchmark (OOTB) show the effectiveness of our sparse correlation filter-based tracker. Yanmei Dong, Min Yang 0003, Mingtao Pei |
ICIP | 3 |
| 2016 | 3D head pose estimation with convolutional neural network trained on synthetic imagesabstractIn this paper, we propose a method to estimate head pose with convolutional neural network, which is trained on synthetic head images. We formulate head pose estimation as a regression problem. A convolutional neural network is trained to learn head features and solve the regression problem. To provide annotated head poses in the training process, we generate a realistic head pose dataset by rendering techniques, in which we consider the variation of gender, age, race and expression. Our dataset includes 74000 head poses rendered from 37 head models. For each head pose, RGB image and annotated pose parameters are given. We evaluate our method on both synthetic and real data. The experiments show that our method improves the accuracy of head pose estimation. Xiabing Liu, Wei Liang 0008, Mingtao Pei |
ICIP | 5 |
| 2016 | Pose-indexed based multi-view method for face alignmentabstractThis paper presents a novel pose-indexed based multi-view (PIMV) face alignment framework. Most of the current cascaded regression face alignment methods generally start with a mean shape. However, when the initial shape is far from the ground truth, the performance significantly deteriorates. Our approach aims to obtain a preferable initial shape from a pose-indexed shape searching space. This space is established by a series of pose-shape pairs which are generally treated as mappings from poses to face shapes. Each shape in this space corresponds to one view which is used as an index of the shape. Subsequently, the index shape is employed as the initial shape for the following iterative stages. The powerful shape-initialization method effectively prevents the local optima problem caused by poor initialization in prediction. Experiments demonstrate that our approach outperforms previous methods on challenging datasets with large pose variations, occlusions and illuminations. Qingjie Zhao, Xiongpeng Wang, Mingtao Pei |
ICIP | 4 |
| 2016 | Attention Estimation for Input Switch in Scalable Multi-display Environments
Xingyuan Bu, Mingtao Pei, Yunde Jia |
ICONIP (4) | 2 |
| 2016 | Pedestrian Detection Using Deep Channel Features in Monocular Image Sequences
Yang He 0004, Hongyan Gu, Mingtao Pei |
ICONIP (3) | 6 |
| 2016 | Driver Face Detection Based on Aggregate Channel Features and Deformable Part-Based Model in Traffic Camera
Xiaoma Xu, Mingtao Pei |
ICONIP (2) | 3 |
| 2016 | A low-cost tele-presence wheelchair systemabstractThis paper presents the architecture and implementation of a tele-presence wheelchair system based on tele-presence robot, intelligent wheelchair, and touch screen technologies. The tele-presence wheelchair system consists of a commercial electric wheelchair, an add-on tele-presence interaction module, and a touchable live video image based user interface (called TIUI). The tele-presence interaction module is used to provide video-chatting for an elderly or disabled person with the family members or caregivers, and also captures the live video of an environment for tele-operation and semi-autonomous navigation. The user interface developed in our lab allows an operator to access the system anywhere and directly touch the live video image of the wheelchair to push it as if he/she did it in the presence. This paper also discusses the evaluation of the user experience. Bin Xu 0019, Mingtao Pei, Yunde Jia |
IROS | 3 |
| 2016 | Nonnegative correlation coding for image classification
Zhen Dong 0002, Wei Liang 0008, Yuwei Wu 0001, Mingtao Pei, Yunde Jia |
Sci. China Inf. Sci. | 4 |
| 2016 | Orthonormal dictionary learning and its application to face recognition
Zhen Dong 0002, Mingtao Pei, Yunde Jia |
Image Vis. Comput. | 2 |
| 2016 | Online Discriminative Tracking With Active Example SelectionabstractMost existing discriminative tracking algorithms use a sampling-and-labeling strategy to collect examples and treat the training example collection as a task that is independent of classifier learning. However, the examples collected directly by sampling are neither necessarily informative nor intended to be useful for classifier learning. Updating the classifier with these examples might introduce ambiguity to the tracker. In this paper, we present a novel online discriminative tracking framework that explicitly couples the objectives of example collection and classifier learning. Our method uses Laplacian regularized least squares (LapRLS) to learn a robust classifier that can sufficiently exploit unlabeled data and preserve the local geometrical structure of the feature space. To ensure the high classification confidence of the classifier, we propose an active example selection approach to automatically select the most informative examples for LapRLS. Part of the selected examples that satisfy strict constraints are labeled to enhance the adaptivity of our tracker, which actually provides robust supervisory information to guide semisupervised learning. With active example selection, we are able to avoid the ambiguity introduced by an independent example collection strategy and to alleviate the drift problem caused by misaligned examples. Comparison with the state-of-the-art trackers on the comprehensive benchmark demonstrates that our tracking algorithm is more effective and accurate. Min Yang 0003, Yuwei Wu 0001, Mingtao Pei, Bo Ma 0001, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Fusion of Skeletal and STIP-Based Features for Action Recognition with RGB-D Devices
Mingtao Pei |
ICIG (2) | 2 |
| 2015 | Discriminative Neighborhood Preserving Dictionary Learning for Image Classification
Shiye Zhang, Zhen Dong 0002, Yuwei Wu 0001, Mingtao Pei |
ICIG (2) | 4 |
| 2015 | Discriminative Orthonormal Dictionary Learning for Fast Low-Rank Representation
Zhen Dong 0002, Mingtao Pei, Yunde Jia |
ICONIP (1) | 2 |
| 2015 | Vehicle Detection Using Appearance and Shape Constrained Active Basis Model
Sai Liu, Mingtao Pei |
ICONIP (3) | 2 |
| 2015 | Robust Online Multi-object Tracking by Maximum a Posteriori Estimation with Sequential Trajectory Prior
Min Yang 0003, Mingtao Pei, Yunde Jia |
ICONIP (1) | 2 |
| 2015 | Non-linear Metric Learning Using Metric Tensor
Liangying Yin, Mingtao Pei |
ICONIP (1) | 2 |
| 2015 | Learning online structural appearance model for robust object tracking
Min Yang 0003, Mingtao Pei, Yuwei Wu 0001, Yunde Jia |
Sci. China Inf. Sci. | 2 |
| 2015 | Online visual tracking by integrating spatio-temporal cuesabstractThe performance of online visual trackers has improved significantly, but designing an effective appearance‐adaptive model is still a challenging task because of the accumulation of errors during the model updating with newly obtained results, which will cause tracker drift. In this study, the authors propose a novel online tracking algorithm by integrating spatio‐temporal cues to alleviate the drift problem. The authors' goal is to develop a more robust way of updating an adaptive appearance model. The model consists of multiple modules called temporal cues, and these modules are updated in an alternate way which can keep both the historical and current information of the tracked object to handle drastic appearance change. Each module is represented by several fragments called spatial cues. In order to incorporate all the spatial and temporal cues, the authors develop an efficient cue quality evaluation criterion that combines appearance and motion information. Then the tracking results are obtained by a two‐stage dynamic integration mechanism. Both qualitative and quantitative evaluations on challenging video sequences demonstrate that the proposed algorithm performs more favourably against the state‐of‐the‐art methods. Yang He 0004, Mingtao Pei, Min Yang 0003, Yuwei Wu 0001, Yunde Jia |
IET Comput. Vis. | 2 |
| 2015 | Robust Discriminative Tracking via Landmark-Based Label PropagationabstractThe appearance of an object could be continuously changing during tracking, thereby being not independent identically distributed. A good discriminative tracker often needs a large number of training samples to fit the underlying data distribution, which is impractical for visual tracking. In this paper, we present a new discriminative tracker via landmark-based label propagation (LLP) that is nonparametric and makes no specific assumption about the sample distribution. With an undirected graph representation of samples, the LLP locally approximates the soft label of each sample by a linear combination of labels on its nearby landmarks. It is able to effectively propagate a limited amount of initial labels to a large amount of unlabeled samples. To this end, we introduce a local landmarks approximation method to compute the cross-similarity matrix between the whole data and landmarks. Moreover, a soft label prediction function incorporating the graph Laplacian regularizer is used to diffuse the known labels to all the unlabeled vertices in the graph, which explicitly considers the local geometrical structure of all samples. Tracking is then carried out within a Bayesian inference framework, where the soft label prediction value is used to construct the observation model. Both qualitative and quantitative evaluations on the benchmark data set containing 51 challenging image sequences demonstrate that the proposed algorithm outperforms the state-of-the-art methods. Yuwei Wu 0001, Mingtao Pei, Min Yang 0003, Junsong Yuan 0001, Yunde Jia |
IEEE Trans. Image Process. | 2 |
| 2015 | Vehicle Type Classification Using a Semisupervised Convolutional Neural NetworkabstractIn this paper, we propose a vehicle type classification method using a semisupervised convolutional neural network from vehicle frontal-view images. In order to capture rich and discriminative information of vehicles, we introduce sparse Laplacian filter learning to obtain the filters of the network with large amounts of unlabeled data. Serving as the output layer of the network, the softmax classifier is trained by multitask learning with small amounts of labeled data. For a given vehicle image, the network can provide the probability of each type to which the vehicle belongs. Unlike traditional methods by using handcrafted visual features, our method is able to automatically learn good features for the classification task. The learned features are discriminative enough to work well in complex scenes. We build the challenging BIT-Vehicle dataset, including 9850 high-resolution vehicle frontal-view images. Experimental results on our own dataset and a public dataset demonstrate the effectiveness of the proposed method. Zhen Dong 0002, Yuwei Wu 0001, Mingtao Pei, Yunde Jia |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2014 | Landmark-Based Inductive Model for Robust Discriminative Tracking
Yuwei Wu 0001, Mingtao Pei, Min Yang 0003, Yang He 0004, Yunde Jia |
ACCV (5) | 2 |
| 2014 | Coupling Semi-supervised Learning and Example Selection for Online Object Tracking
Min Yang 0003, Yuwei Wu 0001, Mingtao Pei, Bo Ma 0001, Yunde Jia |
ACCV (4) | 3 |
| 2014 | Stereovision-Only Based Interactive Mobile Robot for Human-Robot Face-to-Face InteractionabstractIn this paper, we present a stereovision-only based interactive mobile robot for supporting human-robot face-to-face interaction in the real world. A three-level architecture, which consists of sensor level, perception level and behavior level, is designed for the robot in order to perceive, understand and react to the human activity during interaction based only on visual information. A high performance stand-alone stereovision system (RGBD imager), developed in our lab, is applied to obtain the composite of color (RGB) images and dense disparity (D) maps at video rate. The RGBD imager allows the robot a human-like 3-D visual perception ability to (1) autonomously detect the human of interest whom the robot could interact with using the offline learning approaches, and (2) focus exclusively on the target human while both the human and the robot are moving during interaction using on-line learning approaches. We demonstrate and evaluate the performance of our interactive mobile robot in an office environment. The experimental results show that a reliable and dynamic face-to-face interaction is achieved, so that the target human face is always kept in the field of view and at a suitable social distance from the robot. Lei Chen 0020, Zhen Dong 0002, Baofeng Yuan, Mingtao Pei |
ICPR | 5 |
| 2014 | Vehicle Type Classification Using Unsupervised Convolutional Neural NetworkabstractIn this paper, we propose an appearance-based vehicle type classification method from vehicle frontal view images. Unlike other methods using hand-crafted visual features, our method is able to automatically learn good features for vehicle type classification by using a convolutional neural network. In order to capture rich and discriminative information of vehicles, the network is pre-trained by the sparse filtering which is an unsupervised learning method. Besides, the network is with layer-skipping to ensure that final features contain both high-level global and low-level local features. After the final features are obtained, the soft max regression is used to classify vehicle types. We build a challenging vehicle dataset called BIT-Vehicle dataset to evaluate the performance of our method. Experimental results on a public dataset and our own dataset demonstrate that our method is quite effective in classifying vehicle types. Zhen Dong 0002, Mingtao Pei, Yang He 0004, Yanmei Dong, Yunde Jia |
ICPR | 2 |
| 2014 | Visual Tracking Using Multi-stage Random Simple FeaturesabstractIn recent years, deep models offer a promising solution to extract powerful features. Motivated by the effectiveness of the Convolutional Networks (ConvNets) model in image classification and object detection, we present a visual tracking algorithm using the ConvNets model to extract multistage features. The key point of this paper is to show that the multi-stage features extracted by the ConvNets are very proper for visual tracking. In addition, we design a general procedure to generate simple rectangle filters with different complexity, and employ the rectangle filters to construct the ConvNets. The computational cost is reduced by using integral images. The filters are kept constant, thus the update of our tracker would not cost much time. The tracking is formulated as a binary classification problem, and we use an online naive Bayes classifier to build our tracker. The experimental results demonstrate that our tracker achieves comparable results against several state-of-the-art methods. Yang He 0004, Zhen Dong 0002, Min Yang 0003, Lei Chen 0020, Mingtao Pei, Yunde Jia |
ICPR | 5 |
| 2014 | Learning a discriminative mid-level feature for action recognition
Cuiwei Liu, Mingtao Pei, Xinxiao Wu, Yu Kong 0001, Yunde Jia |
Sci. China Inf. Sci. | 2 |
| 2014 | Coupling-and-decoupling: A hierarchical model for occlusion-free object detection
Bo Li 0031, Tianfu Wu 0001, Wenze Hu, Mingtao Pei |
Pattern Recognit. | 5 |
| 2013 | Event recognition based-on social roles in continuous videoabstractIn this paper, we present a new method for video event recognition based on social roles of agents, which are inferred from their daily activities in continuous video. This is motivated from the observation that people have their social roles, and the information of social roles in certain scene provides useful cues for recognizing video events. First, events are represented by an And-Or Graph (AOG), which can represent both the hierarchical decompositions from events, sub-events and atomic actions and the contexts for temporal relations. Then, a model of social roles is proposed to infer the roles of the agents in continuous video. Finally, an improved event parsing algorithm based on social roles context is adopted to recognize events. Experimental results show that our method is effective in performing inference tasks of social roles and can improve performance of event recognition. Mingtao Pei, Zhen Dong 0002 |
ICME | 1 |
| 2013 | Learning and parsing video events with goal and intent prediction
Mingtao Pei, Zhangzhang Si, Benjamin Z. Yao, Song-Chun Zhu |
Comput. Vis. Image Underst. | 1 |
| 2012 | Coupling-and-Decoupling: A Hierarchical Model for Occlusion-Free Car Detection
Bo Li 0031, Tianfu Wu 0001, Wenze Hu, Mingtao Pei |
ACCV (1) | 4 |
| 2012 | Tracking Pedestrian with Multi-component Online Deformable Part-Based Model
Yi Xie 0006, Mingtao Pei, Tianfu Wu 0001 |
ACCV (3) | 2 |
| 2012 | Probabilistic depth map fusion for real-time multi-view stereo
Mingtao Pei, Yunde Jia |
ICPR | 2 |
| 2012 | Robust tracking by accounting for hard negatives explicitly
Tianfu Wu 0004, Mingtao Pei, Anlong Ming, Zhenyu Yao |
ICPR | 3 |
| 2011 | Parsing video events with goal inference and intent predictionabstractIn this paper, we present an event parsing algorithm based on Stochastic Context Sensitive Grammar (SCSG) for understanding events, inferring the goal of agents, and predicting their plausible intended actions. The SCSG represents the hierarchical compositions of events and the temporal relations between the sub-events. The alphabets of the SCSG are atomic actions which are defined by the poses of agents and their interactions with objects in the scene. The temporal relations are used to distinguish events with similar structures, interpolate missing portions of events, and are learned from the training data. In comparison with existing methods, our paper makes the following contributions. i) We define atomic actions by a set of relations based on the fluents of agents and their interactions with objects in the scene. ii) Our algorithm handles events insertion and multi-agent events, keeps all possible interpretations of the video to preserve the ambiguities, and achieves the globally optimal parsing solution in a Bayesian framework; iii) The algorithm infers the goal of the agents and predicts their intents by a top-down process; iv) The algorithm improves the detection of atomic actions by event contexts. We show satisfactory results of event recognition and atomic action detection on the data set we captured which contains 12 event categories in both indoor and outdoor videos. Mingtao Pei, Yunde Jia, Song-Chun Zhu |
ICCV | 1 |
| 2011 | Unsupervised learning of event AND-OR grammar and semantics from videoabstractWe study the problem of automatically learning event AND-OR grammar from videos of a certain environment, e.g. an office where students conduct daily activities. We propose to learn the event grammar under the information projection and minimum description length principles in a coherent probabilistic framework, without manual supervision about what events happen and when they happen. Firstly a predefined set of unary and binary relations are detected for each video frame: e.g. agent's position, pose and interaction with environment. Then their co-occurrences are clustered into a dictionary of simple and transient atomic actions. Recursively these actions are grouped into longer and complexer events, resulting in a stochastic event grammar. By modeling time constraints of successive events, the learned grammar becomes context-sensitive. We introduce a new dataset of surveillance-style video in office, and present a prototype system for video analysis integrating bottom-up detection, grammatical learning and parsing. On this dataset, the learning algorithm is able to automatically discover important events and construct a stochastic grammar, which can be used to accurately parse newly observed video. The learned grammar can be used as a prior to improve the noisy bottom-up detection of atomic actions. It can also be used to infer semantics of the scene. In general, the event grammar is an efficient way for common knowledge acquisition from video. Zhangzhang Si, Mingtao Pei, Benjamin Z. Yao, Song-Chun Zhu |
ICCV | 2 |
| 2011 | Tracking pedestrians with incremental learned intensity and contour templates for PTZ camera visual surveillanceabstractThis paper presents a novel particle-based pedestrian tracking algorithm for PTZ visual surveillance. Most of the state-of-art particle-based tracking algorithms are challenged due to lacking of a reliable moving object detection and drastic scale along with perspective shift of the target. Therefore, pure intensity based algorithms usually miss the target gradually without other features for correcting target location. Our method learns and maintains a contour template of the target besides intensity. Taking into account both the evolution and sudden change of the pedestrian contour, the proposed tracking algorithm maintains several sets of profiles from different perspectives and evolves them incrementally. The effectiveness of our tracking algorithm with extra contour measurement is tested over several surveillance records captured from PTZ camera and estimates the location more robustly than other cutting edge tracking algorithms compared in our experiments. Yi Xie 0006, Mingtao Pei, Guanqun Yu, Yunde Jia |
ICME | 2 |