VLDB 2026 Research / reviewers in the wild / expert
Xinmiao Ding
dblp:123/2897
· DBLP profile ↗
15ranked-venue papers
8as first author
6since 2021 · last 2025
0000-0001-7699-4329ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancement-suppression driven lightweight fine-grained micro-expression recognition
Xinmiao Ding |
J. Vis. Commun. Image Represent. | 1 |
| 2025 | iESTA: Instance-Enhanced Spatial-Temporal Alignment for Video Copy LocalizationabstractVideo copy Segment Localization (VSL) requires the identification of the temporal segments within a pair of videos that contain copied content. Current methods primarily focus on global temporal modeling, overlooking the complementarity of global semantic and local fine-grained features, which limits their effectiveness. Some related methods attempt to incorporate local spatial information but often disrupt spatial semantic structures, resulting in less accurate matching. To address these issues, we propose the Instance-Enhanced Spatial-Temporal Alignment Framework (iESTA), based on a proper representation granularity that integrates instance-level local features and semantic global features. Specifically, the Instance-relation Graph (IRG) is constructed to capture instance-level features and fine-grained interactions, preserving local information integrity and better representing the video feature space in a proper granularity. An instance-GNN structure is designed to refine these graph representations. For global features, we enhance the representation of semantic information, capturing temporal relationships within videos using a Transformer framework. Additionally, we design a Complementarity-perception Alignment Module (CAM) to effectively process and integrate complementary spatial-temporal information, producing accurate frame-to-frame alignment maps. Our approach also incorporates a differentiable Dynamic Time Warping (DTW) method to utilize latent temporal alignments as weak supervisory signals, improving the accuracy of the matching process. Experimental results indicate that our proposed iESTA outperforms state-of-the-art methods on both the small-scale dataset VCDB and the large-scale dataset VCSL. Xinmiao Ding, Jinming Lou, Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Noise-aware progressive multi-scale deepfake detection
Xinmiao Ding, Shuai Pang, Wen Guo 0003 |
Multim. Tools Appl. | 1 |
| 2022 | Multi-cue multi-hypothesis tracking with re-identification for multi-object tracking
Wen Guo 0003, Yuelong Jin, Bin Shan, Xinmiao Ding |
Multim. Syst. | 4 |
| 2021 | Web Objectionable Video Recognition Based on Deep Multi-Instance Learning With Representative Prototypes SelectionabstractTo protect underage people from accessing objectionable videos in the Internet, an effective objectionable video recognition algorithm is necessary for web filtering. Recently, the multi-instance learning has been introduced for objectionable video recognition and achieves impressive results. However, hand-crafted features as well as redundant and noisy frames in objectionable videos become an intractable problem that inevitably degrades the recognition performance. In this paper, we propose a novel representative prototype selection algorithm embedding deep multi-instance representation learning. In the proposed method, an improved convolutional neural network is designed for multimodal multi-instance feature learning and a self-expressive dictionary learning model based on sparse and low rank constraint is designed to select the representative prototypes from each subspace of instances. Then the bag-level feature is constructed via mapping the bag to the selected prototypes. Experiments on three objectionable video sets show the effectiveness of our method for objectionable video recognition. Xinmiao Ding, Bing Li 0001, Yangxi Li, Weihua Xiong, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Multi-Scale Low-Discriminative Feature Reactivation for Weakly Supervised Object LocalizationabstractFor weakly supervised object localization (WSOL), how to avoid the network focusing only on some small discriminative parts is a main challenge needed to solve. The widely-used Class Activation Mapping (CAM) based paradigm usually employs Adversarial Learning (AL) strategy to search more object parts by constantly hiding discovered object features, but the adversarial process is difficult to control. In this paper, we propose a novel CAM-based framework with Multi-scale Low-Discriminative Feature Reactivation (mLDFR) for WSOL. The mLDFR framework reactivates the low-discriminative object parts via bottom-up continuous feature maps recalibration and multi-scale object category mapping. Compared with the AL-based methods, our method fully improves the localization power of the network without damaging the classification power and can perform multi-instance localization, which are hard to achieve under the AL-based framework. Moreover, the mLDFR framework is flexible, and can be built on the top of various classical CNN backbones. Experimental results demonstrate the superiority of our method. With VGG16 as backbone, we achieve 46.96% Cls-Loc top1 err and 66.12% CorLoc on ILSVRC2014, 38.07% Cls-Loc top1 err and 75.04% CorLoc on CUB200-2011, surpassing the state-of-the-arts by a large margin. Bo Wang 0147, Chunfeng Yuan, Bing Li 0001, Xinmiao Ding, Zeya Li, Ying Wu 0001, Weiming Hu 0004 |
IEEE Trans. Image Process. | 4 |
| 2017 | Multi-View Multi-Instance Learning Based on Joint Sparse Representation and Multi-View Dictionary LearningabstractIn multi-instance learning (MIL), the relations among instances in a bag convey important contextual information in many applications. Previous studies on MIL either ignore such relations or simply model them with a fixed graph structure so that the overall performance inevitably degrades in complex environments. To address this problem, this paper proposes a novel multi-view multi-instance learning algorithm (MIL) that combines multiple context structures in a bag into a unified framework. The novel aspects are: (i) we propose a sparse -graph model that can generate different graphs with different parameters to represent various context relations in a bag, (ii) we propose a multi-view joint sparse representation that integrates these graphs into a unified framework for bag classification, and (iii) we propose a multi-view dictionary learning algorithm to obtain a multi-view graph dictionary that considers cues from all views simultaneously to improve the discrimination of the MIL. Experiments and analyses in many practical applications prove the effectiveness of the M IL. Bing Li 0001, Chunfeng Yuan, Weihua Xiong, Weiming Hu 0004, Houwen Peng, Xinmiao Ding, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2016 | A Novel Emotional Saliency Map to Model Emotional Attention Mechanism
Xinmiao Ding, Lulu Huang, Bing Li 0001, Congyan Lang, Zhen Hua |
MMM (2) | 1 |
| 2016 | A multi-instance multi-label learning algorithm based on instance correlations
Tongtong Chen, Xinmiao Ding, Hailin Zou |
Multim. Tools Appl. | 3 |
| 2016 | Multi-Instance Multi-Label Learning Combining Hierarchical Context and its Application to Image AnnotationabstractIn image annotation, one image is often modeled as a bag of regions (“instances”) associated with multiple labels, which is a typical application of multi-instance multi-label learning (MIML). Although lots of research has shown that the interplay embedded among instances and labels can largely boost the image annotation accuracy, most existing MIML methods consider none or partial context cues. In this paper, we propose a novel context-aware MIML model to integrate the instance context and label context into a general framework. Specially, the instance context is constructed with multiple graphs, while the label context is built up through a linear combination of several common latent conceptions that link low level features and high level semantic labels. Comparison with other leading methods on several benchmark datasets in terms of image annotation shows that our proposed method can get better performance than the state-of-the-art approaches. Xinmiao Ding, Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Bo Wang 0147 |
IEEE Trans. Multim. | 1 |
| 2016 | Multi-Perspective Cost-Sensitive Context-Aware Multi-Instance Sparse Coding and Its Application to Sensitive Video RecognitionabstractWith the development of video-sharing websites, P2P, micro-blog, mobile WAP websites, and so on, sensitive videos can be more easily accessed. Effective sensitive video recognition is necessary for web content security. Among web sensitive videos, this paper focuses on violent and horror videos. Based on color emotion and color harmony theories, we extract visual emotional features from videos. A video is viewed as a bag and each shot in the video is represented by a key frame which is treated as an instance in the bag. Then, we combine multi-instance learning (MIL) with sparse coding to recognize violent and horror videos. The resulting MIL-based model can be updated online to adapt to changing web environments. We propose a cost-sensitive context-aware multi- instance sparse coding (MI-SC) method, in which the contextual structure of the key frames is modeled using a graph, and fusion between audio and visual features is carried out by extending the classic sparse coding into cost-sensitive sparse coding. We then propose a multi-perspective multi- instance joint sparse coding (MI-J-SC) method that handles each bag of instances from an independent perspective, a contextual perspective, and a holistic perspective. The experiments demonstrate that the features with an emotional meaning are effective for violent and horror video recognition, and our cost-sensitive context-aware MI-SC and multi-perspective MI-J-SC methods outperform the traditional MIL methods and the traditional SVM and KNN-based methods. Weiming Hu 0004, Xinmiao Ding, Bing Li 0001, Fangshi Wang, Stephen J. Maybank |
IEEE Trans. Multim. | 2 |
| 2014 | A Hierarchical Model Based on Latent Dirichlet Allocation for Action RecognitionabstractInspired by the recent success of hierarchical representation, we propose a new hierarchical variant of latent Dirichlet allocation (h-LDA) for action recognition. The model consists of an appearance group and a motion group, and we introduce a new hierarchical structure including two-layer topics in each group to learn the spatial temporal patterns (STPs) of human actions. The basic idea is that the two-layer topics are used to model the global STPs and the local STPs of the actions respectively. Two groups of discrete words are generated from two complementary kinds of features for each group. Each topic learned in these two groups is used to describe a particular aspect of the actions. Specifically, the mid-level topics are learned to describe the local STPs by including the geometric structure information in the lower-level words. The top-level topics are learned from the mid-level topics and are the mixture distribution of the local STPs, which makes the top-level topics appropriate to represent the global STPs. In addition, we give the learning and inference process by Gibbs sampling with reasonable assumptions. Finally, each sample is discriminatively represented as the probabilistic distribution over the global STPs learned by the proposed h-LDA. Experimental results on two datasets demonstrate the effectiveness of our approach for action recognition. Chunfeng Yuan, Xinmiao Ding |
ICPR | 4 |
| 2012 | Horror Video Scene Recognition Based on Multi-view Multi-instance Learning
Xinmiao Ding, Bing Li 0001, Weiming Hu 0004, Weihua Xiong, Zhenchong Wang |
ACCV (3) | 1 |
| 2012 | Context-aware horror video scene recognition via cost-sensitive sparse coding
Xinmiao Ding, Bing Li 0001, Weiming Hu 0004, Weihua Xiong, Zhenchong Wang |
ICPR | 1 |
| 2012 | Context-aware affective images classification based on bilayer sparse representationabstractIn image understanding, the automatic recognition of emotion in an image is becoming important from an applicative viewpoint. Considering the fact that the emotion evoked by an image is not only from its global appearance but also interplays among local regions, we propose a novel context-aware classification model based on bilayer sparse representation (BSR) that simultaneously takes the local context and global-local context into account. The BSR model contains two layers: global sparse representation (GSR) and local sparse representation (LSR). The GSR is to define global similarities between a test image and all training images; while the LSR is to define similarities of local regions' appearances and their co-occurrence between a test image and all training images. The experiments on two data sets demonstrate that our method is effective on affective images classification. Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Xinmiao Ding |
ACM Multimedia | 4 |