EDBT 2026 Demo / reviewers in the wild / expert
Jingjing Meng
dblp:96/1653
· DBLP profile ↗
35ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scalable High-Fidelity 3D Hand Shape Reconstruction via Graph-Image Frequency Mapping and Graph Frequency DecompositionabstractDespite the impressive performance obtained by recent single-image hand modeling techniques, they lack the capability to capture sufficient details of the 3D hand mesh. This deficiency greatly limits their applications when high-fidelity hand modeling is required, e.g., personalized hand modeling. To address this problem, we design a frequency split network to generate 3D hand meshes using different frequency bands in a coarse-to-fine manner. To capture high-frequency personalized details, we transform the 3D mesh into the frequency domain, and proposed a novel frequency decomposition loss to supervise each frequency component. By leveraging such a coarse-to-fine scheme, hand details that correspond to the higher frequency domain can be preserved. In addition, the proposed network is scalable, and can stop the inference at any resolution level to accommodate different hardware with varying computational powers. To feed the scalable frequency network with frequency split image features, we proposed an image-graph ring feature mapping strategy. To train our network with per-vertex supervision, we use a bidirectional registration strategy to generate a topology-fixed ground-truth. To quantitatively evaluate the performance of our method in terms of recovering personalized shape details, we introduce a new evaluation metric named Mean-frequency Signal-to-Noise Ratio (MSNR) to measure the mean signal-to-noise ratio of mesh signal on each frequency component. Extensive experiments demonstrate that our approach generates fine-grained details for high-fidelity 3D hand reconstruction, and our evaluation metric is more effective than traditional metrics for measuring mesh details. Tianyu Luan, Yuanhao Zhai 0001, Jingjing Meng, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Interaction-Centric Spatio-Temporal Context Reasoning for Multi-person Video HOI Recognition
Yisong Wang 0005, Nan Xi, Jingjing Meng, Junsong Yuan 0001 |
ECCV (32) | 3 |
| 2024 | Cross-Scenario Unknown-Aware Face Anti-Spoofing With Evidential Semantic Consistency LearningabstractIn recent years, domain adaptation techniques have been widely used to adapt face anti-spoofing models to a cross-scenario target domain. Most previous methods assume that the Presentation Attack Instruments (PAIs) in such cross-scenario target domain are same as in the source domain. However, as the malicious users are free to use any form of unknown PAIs to attack the system, this assumption does not always hold in practical applications of face anti-spoofing. Thus, unknown PAIs would inevitably lead to significant performance degradation, since samples of known and unknown PAIs usually have large differences. In this paper, we propose an Evidential Semantic Consistency Learning (ESCL) framework to address this problem. Specifically, a regularized evidential deep learning strategy with a two-way balance of class probability and uncertainty is leveraged to produce uncertainty scores for unknown PAI detection. Meanwhile, entropy optimization-based semantic consistency learning strategy is also employed to encourage features of live and known PAIs to be gathered in the label-conditioned clusters across the source and target domains, while make the features of unknown PAIs to be self-clustered according to intrinsic semantic information. In addition, a new evaluation metric, KUHAR, is proposed to comprehensively evaluate the error rate of known classes and unknown PAIs. Extensive experimental results on six public datasets demonstrate the effectiveness of our method in generalizing face anti-spoofing models to both known classes and unknown PAIs with different types and quantities in a cross-scenario testing domain. Our method achieves state-of-the-art performance on eight different protocols. Fangling Jiang, Yunfan Liu 0001, Haolin Si, Jingjing Meng, Qi Li 0005 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion AnalysisabstractFashion representation learning involves the analysis and understanding of various visual elements at different granularities and the interactions among them. Existing works often learn fine-grained fashion representations at the attribute level without considering their relationships and inter-dependencies across different classes. In this work, we propose to learn an attribute and class-specific fashion representation duet to better model such attribute relationships and inter-dependencies by leveraging prior knowledge about the taxonomy of fashion attributes and classes. Through two sub-networks for the attributes and classes, respectively, our proposed an embedding network progressively learns and refines the visual representation of a fashion image to improve its robustness for fashion retrieval. A multi-granularity loss consisting of attribute-level and class-level losses is proposed to introduce appropriate inductive bias to learn across different granularities of the fashion representations. Experimental results on three benchmark datasets demonstrate the effectiveness of our method, which outperforms the state-of-the-art methods by a large margin. Yan Gao 0029, Jingjing Meng |
CVPR | 3 |
| 2023 | High Fidelity 3D Hand Shape Reconstruction via Scalable Graph Frequency DecompositionabstractDespite the impressive performance obtained by recent single-image hand modeling techniques, they lack the capability to capture sufficient details of the 3D hand mesh. This deficiency greatly limits their applications when high-fidelity hand modeling is required, e.g., personalized hand modeling. To address this problem, we design a frequency split network to generate 3D hand mesh using different frequency bands in a coarse-to-fine manner. To capture high-frequency personalized details, we transform the 3D mesh into the frequency domain, and propose a novel frequency decomposition loss to supervise each frequency component. By leveraging such a coarse-to-fine scheme, hand details that correspond to the higher frequency domain can be preserved. In addition, the proposed network is scalable, and can stop the inference at any resolution level to accommodate different hardware with varying computational powers. To quantitatively evaluate the performance of our method in terms of recovering personalized shape details, we introduce a new evaluation metric named Mean Signal-to-Noise Ratio (MSNR) to measure the signal-to-noise ratio of each mesh frequency component. Extensive experiments demonstrate that our approach generates fine-grained details for high-fidelity 3D hand reconstruction, and our evaluation metric is more effective for measuring mesh details compared with traditional metrics. The code is available at https://github.com/tyluann/FreqHand. Tianyu Luan, Yuanhao Zhai 0001, Jingjing Meng, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001 |
CVPR | 3 |
| 2023 | Open Set Video HOI detection from Action-centric Chain-of-Look PromptingabstractHuman-Object Interaction (HOI) detection is essential for understanding and modeling real-world events. Existing works on HOI detection mainly focus on static images and a closed setting, where all HOI classes are provided in the training set. In comparison, detecting HOIs in videos in open set scenarios is more challenging. First, under open set circumstances, HOI detectors are expected to hold strong generalizability to recognize unseen HOIs not included in the training data. Second, accurately capturing temporal contextual information from videos is difficult, but it is crucial for detecting temporal-related actions such as open, close, pull, push. To this end, we propose ACoLP, a model of Action-centric Chain-of-Look Prompting for open set video HOI detection. ACoLP regards actions as the carrier of semantics in videos, which captures the essential semantic information across frames. To make the model generalizable on unseen classes, inspired by the chain-of-thought prompting in natural language processing, we introduce the chain-of-look prompting scheme that decomposes prompt generation from large-scale vision-language model into a series of intermediate visual reasoning steps. Consequently, our model captures complex visual reasoning processes underlying the HOI events in videos, providing essential guidance for detecting unseen classes. Extensive experiments on two video HOI datasets, VidHOI and CAD120, demonstrate that ACoLP achieves competitive performance compared with the state-of-the-art methods in the conventional closed setting, and outperforms existing methods by a large margin in the open set setting. Our code is avaliable at https://github.com/southnx/ACoLP. Nan Xi, Jingjing Meng, Junsong Yuan 0001 |
ICCV | 2 |
| 2023 | Chain-of-Look Prompting for Verb-centric Surgical Triplet Recognition in Endoscopic VideosabstractSurgical triplet recognition aims to recognize surgical activities as triplets (i.e., ), which provides fine-grained information essential for surgical scene understanding. Existing methods for surgical triplet recognition rely on compositional methods that recognize the instrument, verb, and target simultaneously. In contrast, our method, called chain-of-look prompting, casts the problem of surgical triplet recognition as visual prompt generation from large-scale vision-language (VL) models, and explicitly decomposes the task into a series of video reasoning processes. Chain-of-Look prompting is inspired by: (1) the chain-of-thought prompting in natural language processing, which divides a problem into a sequence of intermediate reasoning steps; (2) the inter-dependency between motion and visual appearance in the human vision system. Since surgical activities are conveyed by the actions of physicians, we regard the verbs as the carrier of semantics in surgical endoscopic videos. Additionally, we utilize the BioMed large language model to calibrate the generated visual prompt features for surgical scenarios. Our approach captures the visual reasoning processes underlying surgical activities and achieves better performance compared to the state-of-the-art methods on the largest surgical triplet recognition dataset, CholecT50. The code is available at https://github.com/southnx/CoLSurgical. Nan Xi, Jingjing Meng, Junsong Yuan 0001 |
ACM Multimedia | 2 |
| 2022 | Forest Graph Convolutional Network for Surgical Action Triplet Recognition in Endoscopic VideosabstractRecognizing surgical activities in endoscopic videos is of vital importance for developing context-aware decision support in the operating room. In this work, we model each surgical activity as an action triplet, consisting of the surgical instrument, the action, and the target organ that the instrument is interacting with. The goal is to recognize these action triplets from endoscopic videos. However, correctly recognizing fine-grained activity triplets is challenging because of the long-tail distribution of the triplet classes and the complex associations between triplets as well as within each triplet. In addition, multiple triplets may appear in a given video frame. To address these challenges, we propose a new model for surgical action triplet recognition based on a classification forest and Graph Convolutional Network (GCN), which we call Forest GCN. The classification forest is employed to calibrate fine-grained triplet classifiers by the upstream parent classifiers to suppress noisy logits of the triplet classes in the long tail. And stacked GCNs are designed to model the dependencies between triplet classes while leveraging the language embedding. Experiments on the endoscopic video dataset, CholecT50, demonstrate that our proposed method outperforms current state-of-the-art methods on surgical action triplet recognition. Nan Xi, Jingjing Meng, Junsong Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | 3D Object Representation Learning: A Set-to-Set Matching PerspectiveabstractIn this paper, we tackle the 3D object representation learning from the perspective of set-to-set matching. Given two 3D objects, calculating their similarity is formulated as the problem of set-to-set similarity measurement between two set of local patches. As local convolutional features from convolutional feature maps are natural representations of local patches, the set-to-set matching between sets of local patches is further converted into a local features pooling problem. To highlight good matchings and suppress the bad ones, we exploit two pooling methods: 1) bilinear pooling and 2) VLAD pooling. We analyze their effectiveness in enhancing the set-to-set matching and meanwhile establish their connection. Moreover, to balance different components inherent in a bilinear-pooled feature, we propose the harmonized bilinear pooling operation, which follows the spirits of intra-normalization used in VLAD pooling. To achieve an end-to-end trainable framework, we implement the proposed harmonized bilinear pooling and intra-normalized VLAD as two layers to construct two types of neural network, multi-view harmonized bilinear network (MHBN) and multi-view VLAD network (MVLADN). Systematic experiments conducted on two public benchmark datasets demonstrate the efficacy of the proposed MHBN and MVLADN in 3D object recognition. Jingjing Meng, Ming Yang 0007, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | S3F: A Multi-View Slow-Fast Network For Alzheimer's Disease DiagnosisabstractAlzheimer's disease (AD) is the most common form of dementia in the elderly. As early detection and diagnosis is imperative for the intervention and prevention of its progression into more detrimental stages, pioneering works have been proposed that use the resting-state functional MRI (rs-fMRI) to identify early mild cognitive impairment (EMCI) based on various convolutional neural networks (CNNs). However the accuracy is not satisfactory. In this paper, we propose a multi-view model based on the SlowFast network, a recently proposed model for video recognition. The rs-fMRI data are treated as videos from three perspectives (i.e. coronal, horizontal and sagittal, corresponding to three anatomical planes in human body) and the jointly learned hierarchical representations are fused in the fully connected layer. We examine our model on a publicly accessible Alzheimer's Disease Neuroimaging Initiative (ADNI) database. Our method significantly outperforms other competing methods and achieves state-of-the-art accuracy. Besides, we also provide a baseline on the classification task over all clinical phases of AD. Ziqiao Weng, Jingjing Meng, Zhaohua Ding, Junsong Yuan 0001 |
ICME | 2 |
| 2020 | Dynamic Graph CNN for Event-Camera Based Gesture RecognitionabstractEvent camera is a kind of bio-inspired sensor which is able to capture the motion in asynchronous events stream. An event is triggered when the pixel has a brightness change. In spatio-temporal space, those events will form an event cloud, which has specific 3D geometry to capture the dynamic scene. To analyze the event cloud, previous works usually convert event streams into frame-based images which did not fully utilize its 3D geometry in the spatio-temporal event space. In this work, we propose to recognize the spatio-temporal 3D event clouds for gesture recognition using Dynamic Graph CNN (DGCNN) which directly takes 3D points as input and is successfully used for 3D object recognition. We adapt DGCNN to perform action recognition by recognizing 3D geometry features in spatio-temporal space of the event data. We achieve state-of-the-art accuracy of 98.56% on the IBM DVS128 Gesture dataset and 95.94% on the DHP19 dataset. Jingjing Meng, Xinchao Wang, Junsong Yuan 0001 |
ISCAS | 2 |
| 2020 | HOT-Net: Non-Autoregressive Transformer for 3D Hand-Object Pose EstimationabstractAs we use our hands frequently in daily activities, the analysis of hand-object interactions plays a critical role to many multimedia understanding and interaction applications. Different from conventional 3D hand-only and object-only pose estimation, estimating 3D hand-object pose is more challenging due to the mutual occlusions between hand and object, as well as the physical constraints between them. To overcome these issues, we propose to fully utilize the structural correlations among hand joints and object corners in order to obtain more reliable poses. Our work is inspired by structured output learning models in sequence transduction field like Transformer encoder-decoder framework. Besides modeling inherent dependencies from extracted 2D hand-object pose, our proposed Hand-Object Transformer Network (HOT-Net) also captures the structural correlations among 3D hand joints and object corners. Similar to Transformer's autoregressive decoder, by considering structured output patterns, this helps better constrain the output space and leads to more robust pose estimation. However, different from Transformer's sequential modeling mechanism, HOT-Net adopts a novel non-autoregressive decoding strategy for 3D hand-object pose estimation. Specifically, our model removes the Transformer's dependence on previously generated results and explicitly feeds a reference 3D hand-object pose into the decoding process to provide equivalent target pose patterns for parallely localizing each 3D keypoint. To further improve physical validity of estimated hand pose, besides anatomical constraints, we propose a cooperative pose constraint, aiming to enable the hand pose to cooperate with hand shape, to generate hand mesh. We demonstrate real-time speed and state-of-the-art performance on benchmark hand-object datasets for both 3D hand and object poses. Lin Huang 0004, Jianchao Tan, Jingjing Meng, Ji Liu 0002, Junsong Yuan 0001 |
ACM Multimedia | 3 |
| 2020 | Product Quantization Network for Fast Visual Search
Jingjing Meng, Hailin Jin, Junsong Yuan 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | Detecting spatiotemporal irregularities in videos via a 3D convolutional autoencoder
Mengjia Yan 0003, Jingjing Meng, Chunluan Zhou, Zhigang Tu 0001, Yap-Peng Tan, Junsong Yuan 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | Asymmetric Mapping Quantization for Nearest Neighbor SearchabstractNearest neighbor search is a fundamental problem in computer vision and machine learning. The straightforward solution, linear scan, is both computationally and memory intensive in large scale high-dimensional cases, hence is not preferable in practice. Therefore, there have been a lot of interests in algorithms that perform approximate nearest neighbor (ANN) search. In this paper, we propose a novel addition-based vector quantization algorithm, Asymmetric Mapping Quantization (AMQ), to efficiently conduct ANN search. Unlike existing addition-based quantization methods that suffer from handling the problem caused by the norm of database vector, we map the query vector and database vector using different mapping functions to transform the computation of L-2 distance to inner product similarity, thus do not need to evaluate the norm of database vector. Moreover, we further propose Distributed Asymmetric Mapping Quantization (DAMQ) to enable AMQ to work on very large dataset by distributed learning. Extensive experiments on approximate nearest neighbor search and image retrieval validate the merits of the proposed AMQ and DAMQ. Weixiang Hong 0001, Xueyan Tang, Jingjing Meng, Junsong Yuan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Joint Representative Selection and Feature Learning: A Semi-Supervised ApproachabstractIn this paper, we propose a semi-supervised approach for representative selection, which finds a small set of representatives that can well summarize a large data collection. Given labeled source data and big unlabeled target data, we aim to find representatives in the target data, which can not only represent and associate data points belonging to each labeled category, but also discover novel categories in the target data, if any. To leverage labeled source data, we guide representative selection from labeled source to unlabeled target. We propose a joint optimization framework which alternately optimizes (1) representative selection in the target data and (2) discriminative feature learning from both the source and the target for better representative selection. Experiments on image and video datasets demonstrate that our proposed approach not only finds better representatives, but also can discover novel categories in the target data that are not in the source. Suchen Wang, Jingjing Meng, Junsong Yuan 0001, Yap-Peng Tan |
CVPR | 2 |
| 2019 | Boosting Positive and Unlabeled Learning for Anomaly Detection With Multi-FeaturesabstractOne of the key challenges of machine learning-based anomaly detection relies on the difficulty of obtaining anomaly data for training, which is usually rare, diversely distributed, and difficult to collect. To address this challenge, we formulate anomaly detection as a Positive and Unlabeled (PU) learning problem where only labeled positive (normal) data and unlabeled (normal and anomaly) data are required for learning an anomaly detector. As a semi-supervised learning method, it does not require providing labeled anomaly data for the training, thus it is easily deployed to various applications. As the unlabeled data can be extremely unbalanced, we introduce a novel PU learning method, which can tackle the situation where an unlabeled data set is mostly composed of positive instances. We start by using a linear model to extract the most reliable negative instances followed by a self-learning process to add reliable negative and positive instances with different speeds based on the estimated positive class prior. Furthermore, when feedback is available, we adopt boosting in the self-learning process to advantageously exploit the instability characteristic of PU learning. The classifiers in the self-learning process are weighted combined based on the estimated error rate to build the final classifier. Extensive experiments on six real datasets and one synthetic dataset show that our methods have better results under different conditions compared to existing methods. Jingjing Meng, Yap-Peng Tan, Junsong Yuan 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Distributed Composite QuantizationabstractApproximate nearest neighbor (ANN) search is a fundamental problem in computer vision, machine learning and information retrieval. Recently, quantization-based methods have drawn a lot of attention due to their superior accuracy and comparable efficiency compared with traditional hashing techniques. However, despite the prosperity of quantization techniques, they are all designed for the centralized setting, i.e., quantization is performed on the data on a single machine. This makes it difficult to scale these techniques to large-scale datasets. Built upon the Composite Quantization, we propose a novel quantization algorithm for data dis- tributed across different nodes of an arbitrary network. The proposed Distributed Composite Quantization (DCQ) decom-poses Composite Quantization into a set of decentralized sub-problems such that each node solves its own sub-problem on its local data, meanwhile is still able to attain consistent quantizers thanks to the consensus constraint. Since there is no exchange of training data across the nodes in the learning process, the communication cost of our method is low. Ex- tensive experiments on ANN search and image retrieval tasks validate that the proposed DCQ significantly improves Composite Quantization in both efficiency and scale, while still maintaining competitive accuracy. Weixiang Hong 0001, Jingjing Meng, Junsong Yuan 0001 |
AAAI | 2 |
| 2018 | Tensorized Projection for High-Dimensional Binary EmbeddingabstractEmbedding high-dimensional visual features (d-dimensional) to binary codes (b-dimensional) has shown advantages in various vision tasks such as object recognition and image retrieval. Meanwhile, recent works have demonstrated that to fully utilize the representation power of high-dimensional features, it is critical to encode them into long binary codes rather than short ones, i.e., b ~ O(d). However, generating long binary codes involves large projection matrix and high-dimensional matrix-vector multiplication, thus is memory and computationally intensive. To tackle these problems, we propose Tensorized Projection (TP) to decompose the projection matrix using Tensor-Train (TT) format, which is a chain-like representation that allows to operate tensor in an efficient manner. As a result, TP can drastically reduce the computational complexity and memory cost. Moreover, by using the TT-format, TP can regulate the projection matrix against the risk of over-fitting, consequently, lead to better performance than using either dense projection matrix (like ITQ) or sparse projection matrix. Experimental comparisons with state-of-the-art methods over various visual tasks demonstrate both the efficiency and performance ad- vantages of our proposed TP, especially when generating high dimensional binary codes, e.g., when b ≥ d. Weixiang Hong 0001, Jingjing Meng, Junsong Yuan 0001 |
AAAI | 2 |
| 2018 | Multi-View Harmonized Bilinear Network for 3D Object RecognitionabstractView-based methods have achieved considerable success in 3D object recognition tasks. Different from existing view-based methods pooling the view-wise features, we tackle this problem from the perspective of patches-to-patches similarity measurement. By exploiting the relationship between polynomial kernel and bilinear pooling, we obtain an effective 3D object representation by aggregating local convolutional features through bilinear pooling. Meanwhile, we harmonize different components inherited in the bilinear feature to obtain a more discriminative representation. To achieve an end-to-end trainable framework, we incorporate the harmonized bilinear pooling as a layer of a network, constituting the proposed Multi-view Harmonized Bilinear Network (MHBN). Systematic experiments conducted on two public benchmark datasets demonstrate the efficacy of the proposed methods in 3D object recognition. Jingjing Meng, Junsong Yuan 0001 |
CVPR | 2 |
| 2018 | Video Summarization Via Multiview Representative SelectionabstractVideo contents are inherently heterogeneous. To exploit different feature modalities in a diverse video collection for video summarization, we propose to formulate the task as a multiview representative selection problem. The goal is to select visual elements that are representative of a video consistently across different views (i.e., feature modalities). We present in this paper the multiview sparse dictionary selection with centroid co-regularization method, which optimizes the representative selection in each view, and enforces that the view-specific selections to be similar by regularizing them towards a consensus selection. We also introduce a diversity regularizer to favor a selection of diverse representatives. The problem can be efficiently solved by an alternating minimizing optimization with the fast iterative shrinkage thresholding algorithm. Experiments on synthetic data and benchmark video datasets validate the effectiveness of the proposed approach for video summarization, in comparison with other video summarization methods and representative selection methods such as K-medoids, sparse dictionary selection, and multiview clustering. Jingjing Meng, Suchen Wang, Hongxing Wang 0001, Junsong Yuan 0001, Yap-Peng Tan |
IEEE Trans. Image Process. | 1 |
| 2018 | Query Adaptive Multiview Object Instance Search and Localization Using SketchesabstractSketch-based object search is a challenging problem mainly due to three difficulties: 1) how to match the primary sketch query with the colorful image; 2) how to locate the small object in a big image that is similar to the sketch query; and 3) given the large image database, how to ensure an efficient search scheme that is reasonably scalable. To address the above challenges, we propose leveraging object proposals for object search and localization. However, instead of purely relying on sketch features, we propose fully utilizing the appearance features of object proposals to resolve the ambiguities between the matching sketch query and object proposals. Our proposed query adaptive search is formulated as a subgraph selection problem, which can be solved by the maximum flow algorithm. By performing query expansion, it can accurately locate the small target objects in a cluttered background or densely drawn deformation-intensive cartoon (Manga like) images. To improve the computing efficiency of matching proposal candidates, the proposed Multi View Spatially Constrained Proposal Selection encodes each identified object proposal in terms of a small local basis of anchor objects. The results on benchmark datasets validate the advantages of utilizing both the sketch and appearance features for sketch-based search, while ensuring sufficient scalability at the same time. Sreyasee Das Bhattacharjee, Junsong Yuan 0001, Jingjing Meng, Ling-Yu Duan |
IEEE Trans. Multim. | 4 |
| 2017 | Is My Object in This Video? Reconstruction-based Object Search in VideosabstractThis paper addresses the problem of video-level object instance search, which aims to retrieve the videos in the database that contain a given query object instance. Without prior knowledge about "when" and "where" an object of interest may appear in a video, determining "whether" a video contains the target object is computationally prohibitive, as it requires exhaustively matching the query against all possible spatial-temporal locations in each video that an object may appear. To alleviate the computational and memory cost, we propose the Reconstruction-based Object SEarch (ROSE) method.It characterizes a huge corpus of features of possible spatial-temporal locations in the video into the parameters of the reconstruction model. Since the memory cost of storing reconstruction model is much less than that of storing features of possible spatial-temporal locations in the video, the efficiency of the search is significantly boosted. Comprehensive experiments on three benchmark datasets demonstrate the promising performance of the proposed ROSE method. Jingjing Meng, Junsong Yuan 0001 |
IJCAI | 2 |
| 2016 | From Keyframes to Key Objects: Video Summarization by Representative Object Proposal SelectionabstractWe propose to summarize a video into a few key objects by selecting representative object proposals generated from video frames. This representative selection problem is formulated as a sparse dictionary selection problem, i.e., choosing a few representatives object proposals to reconstruct the whole proposal pool. Compared with existing sparse dictionary selection based representative selection methods, our new formulation can incorporate object proposal priors and locality prior in the feature space when selecting representatives. Consequently it can better locate key objects and suppress outlier proposals. We convert the optimization problem into a proximal gradient problem and solve it by the fast iterative shrinkage thresholding algorithm (FISTA). Experiments on synthetic data and real benchmark datasets show promising results of our key object summarization approach in video content mining and search. Comparisons with existing representative selection approaches such as K-mediod, sparse dictionary selection and density based selection validate that our formulation can better capture the key video objects despite appearance variations, cluttered backgrounds and camera motions. Jingjing Meng, Hongxing Wang 0001, Junsong Yuan 0001, Yap-Peng Tan |
CVPR | 1 |
| 2016 | Object Instance Search in Videos via Spatio-Temporal Trajectory DiscoveryabstractGiven a specific object as query, object instance search aims to not only retrieve the images or frames that contain the query, but also locate all its occurrences. In this work, we explore the use of spatio-temporal cues to improve the quality of object instance search from videos. To this end, we formulate this problem as the spatio-temporal trajectory search problem, where a trajectory is a sequence of bounding boxes that locate the object instance in each frame. The goal is to find the top- K trajectories that are likely to contain the target object. Despite the large number of trajectory candidates, we build on a recent spatio- temporal search algorithm for event detection to efficiently find the optimal spatio- temporal trajectories in large video volumes , with complexity linear to the video volume size. We solve the key bottleneck in applying this approach to object instance search by leveraging a randomized approach to enable fast scoring of any bounding boxes in the video volume. In addition , we present a new dataset for video object instance search. Experimental results on a 73-hour video dataset demonstrate that our approach improves the performance of video object instance search and localization over the state-of-the-art search and tracking methods. Jingjing Meng, Junsong Yuan 0001, Gang Wang 0012, Yap-Peng Tan |
IEEE Trans. Multim. | 1 |
| 2015 | Fast object instance search in videos from one exampleabstractWe present an efficient approach to search for and locate all occurrences of a specific object in large video volumes, given a single query example. Locations of object occurrences are returned as spatio-temporal trajectories in the 3D video volume. Despite much work on object instance search in image datasets, these methods locate the object independently in each image, therefore do not preserve the spatio-temporal consistency in consecutive video frames. This results in sub-optimal performance if directly applied to videos, as will be shown in our experiments. We propose to locate the object jointly across video frames using spatio-temporal search. The efficiency and effectiveness of the proposed approach is demonstrated on a consumer video dataset consisting of crawled YouTube videos and mobile captured consumer clips. Our method significantly improves the localized search accuracy over the baseline, which treats each frame independently. Moreover, it is able to find the top 100 object trajectories in the 5.5-hour dataset within 30 seconds. Jingjing Meng, Junsong Yuan 0001, Yap-Peng Tan, Gang Wang 0012 |
ICIP | 1 |
| 2015 | Randomized Spatial Context for Object SearchabstractSearching visual objects in large image or video data sets is a challenging problem, because it requires efficient matching and accurate localization of query objects that often occupy a small part of an image. Although spatial context has been shown to help produce more reliable detection than methods that match local features individually, how to extract appropriate spatial context remains an open problem. Instead of using fixed-scale spatial context, we propose a randomized approach to deriving spatial context, in the form of spatial random partition. The effect of spatial context is achieved by averaging the matching scores over multiple random patches. Our approach offers three benefits: 1) the aggregation of the matching scores over multiple random patches provides robust local matching; 2) the matched objects can be directly identified on the pixelwise confidence map, which results in efficient object localization; and 3) our algorithm lends itself to easy parallelization and also allows a flexible tradeoff between accuracy and speed through adjusting the number of partition times. Both theoretical studies and experimental comparisons with the state-of-the-art methods validate the advantages of our approach. Yuning Jiang 0001, Jingjing Meng, Junsong Yuan 0001, Jiebo Luo 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Robust Part-Based Hand Gesture Recognition Using Kinect SensorabstractThe recently developed depth sensors, e.g., the Kinect sensor, have provided new opportunities for human-computer interaction (HCI). Although great progress has been made by leveraging the Kinect sensor, e.g., in human body tracking, face recognition and human action recognition, robust hand gesture recognition remains an open problem. Compared to the entire human body, the hand is a smaller object with more complex articulations and more easily affected by segmentation errors. It is thus a very challenging problem to recognize hand gestures. This paper focuses on building a robust part-based hand gesture recognition system using Kinect sensor. To handle the noisy hand shapes obtained from the Kinect sensor, we propose a novel distance metric, Finger-Earth Mover's Distance (FEMD), to measure the dissimilarity between hand shapes. As it only matches the finger parts while not the whole hand, it can better distinguish the hand gestures of slight differences. The extensive experiments demonstrate that our hand gesture recognition system is accurate (a 93.2% mean accuracy on a challenging 10-gesture dataset), efficient (average 0.0750 s per frame), robust to hand articulations, distortions and orientation or scale changes, and can work in uncontrolled environments (cluttered backgrounds and lighting conditions). The superiority of our system is further demonstrated in two real-life HCI applications. Zhou Ren, Junsong Yuan 0001, Jingjing Meng, Zhengyou Zhang |
IEEE Trans. Multim. | 3 |
| 2012 | Randomized visual phrases for object searchabstractAccurate matching of local features plays an essential role in visual object search. Instead of matching individual features separately, using the spatial context, e.g., bundling a group of co-located features into a visual phrase, has shown to enable more discriminative matching. Despite previous work, it remains a challenging problem to extract appropriate spatial context for matching. We propose a randomized approach to deriving visual phrase, in the form of spatial random partition. By averaging the matching scores over multiple randomized visual phrases, our approach offers three benefits: 1) the aggregation of the matching scores over a collection of visual phrases of varying sizes and shapes provides robust local matching; 2) object localization is achieved by simple thresholding on the voting map, which is more efficient than subimage search; 3) our algorithm lends itself to easy parallelization and also allows a flexible trade-off between accuracy and speed by adjusting the number of partition times. Both theoretical studies and experimental comparisons with the state-of-the-art methods validate the advantages of our approach. Yuning Jiang 0001, Jingjing Meng, Junsong Yuan 0001 |
CVPR | 2 |
| 2012 | Rapid object search engine for contextual advertisementabstractVisual object search, with the goal to find and locate the target object in large image or video collections, is of great interest for many applications and hence has received intensive attentions in recent years. In this demo, we present a spatial context-aware large-scale visual object search system, which is robust to cluttered backgrounds and can well handle scale variations of the objects. Different from the traditional image retrieval systems only matching individual points or fixed-scale spatial contexts, the proposed system considers spatial contexts of varying sizes and shapes, in the form of randomized spatial partition (RSP), and hence provides more accurate search results. Moreover, compared to the computational expensive RANSAC algorithm used in the state-of-the-art retrieval systems, the RSP framework lends our system to easy parallelization and significant speedup for object localization. Consequently, our system works accurately and efficiently. In addition, an Android application has been developed for mobile tasks, by which the user can take a photo of the object he/she wants and then search the same products and their selling information. Yuning Jiang 0001, Junsong Yuan 0001, Jingjing Meng |
ACM Multimedia | 3 |
| 2011 | Grid-based local feature bundling for efficient object search and localizationabstractWe propose a new grid-based image representation for discriminative visual object search, with the goal to efficiently locate the query object in a large image collection. After extracting local invariant features, we partition the image into non-overlapping rectangular grid cells. Each grid bundles the local features within it and is characterized by a histogram of visual words. Given both positive and negative queries, each grid is assigned a mutual information score to match and locate the query object. This new image representation offers two great benefits for efficient object search: 1) as the grid bundles local features, the spatial contextual information enhances the discriminative matching; and 2) it enables faster object localization by searching visual object in the grid-level image. To evaluate our approach, we perform experiments on a very challenging logo database BelgaLogos [1] of 10,000 images. The comparison with the state-of-the-art methods highlights the effectiveness of our approach in both accuracy and speed. Yuning Jiang 0001, Jingjing Meng, Junsong Yuan 0001 |
ICIP | 2 |
| 2011 | Robust hand gesture recognition with kinect sensorabstractHand gesture based Human-Computer-Interaction (HCI) is one of the most natural and intuitive ways to communicate between people and machines, since it closely mimics how human interact with each other. In this demo, we present a hand gesture recognition system with Kinect sensor, which operates robustly in uncontrolled environments and is insensitive to hand variations and distortions. Our system consists of two major modules, namely, hand detection and gesture recognition. Different from traditional vision-based hand gesture recognition methods that use color-markers for hand detection, our system uses both the depth and color information from Kinect sensor to detect the hand shape, which ensures the robustness in cluttered environments. Besides, to guarantee its robustness to input variations or the distortions caused by the low resolution of Kinect sensor, we apply a novel shape distance metric called Finger-Earth Mover's Distance (FEMD) for hand gesture recognition. Consequently, our system operates accurately and efficiently. In this demo, we demonstrate the performance of our system in two real-life applications, arithmetic computation and rock-paper-scissors game. Zhou Ren, Jingjing Meng, Junsong Yuan 0001, Zhengyou Zhang |
ACM Multimedia | 2 |
| 2010 | Interactive visual object search through mutual information maximizationabstractSearching for small objects (e.g., logos) in images is a critical yet challenging problem. It becomes more difficult when target objects differ significantly from the query object due to changes in scale, viewpoint or style, not to mention partial occlusion or cluttered backgrounds. With the goal to retrieve and accurately locate the small object in the images, we formulate the object search as the problem of finding subimages with the largest mutual information toward the query object. Each image is characterized by a collection of local features. Instead of only using the query object for matching, we propose a discriminative matching using both positive and negative queries to obtain the mutual information score. The user can verify the retrieved subimages and improve the search results incrementally. Our experiments on a challenging logo database of 10,000 images highlight the effectiveness of this approach. Jingjing Meng, Junsong Yuan 0001, Yuning Jiang 0001, Nitya Narasimhan, Venu Vasudevan, Ying Wu 0001 |
ACM Multimedia | 1 |
| 2008 | Mining Recurring Events Through Forest GrowingabstractRecurring events are short temporal patterns that consist of multiple instances in the target database. Without anya prioriknowledge of the recurring events, in terms of their lengths, temporal locations, the total number of such events, and possible variations, it is a challenging problem to discover them because of the enormous computational cost involved in analyzing huge databases and the difficulty in accommodating all the possible variations without even knowing the target. Junsong Yuan 0001, Jingjing Meng, Ying Wu 0001, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Mining repetitive clips through finding continuous pathsabstractAutomatically discovering repetitive clips from large video database is a challenging problem due to the enormous computational cost involved in exploring the huge solution space. Without any a priori knowledge of the contents, lengths and total number of the repetitive clips, we need to discover all of them in the video database. To address the large computational cost, we propose a novel method which translates repetitive clip mining to the continuous path finding problem in a matching trellis, where sequence matching can be accelerated by taking advantage of the temporal redundancies in the videos. By applying the locality sensitive hashing (LSH) for efficient similarity query and the proposed continuous path finding algorithm, our method is of only quadratic complexity of the database size. Experiments conducted on a 10.5-hour TRECVID news dataset have shown the effectiveness, which can discover repetitive clips of various lengths and contents in only 25 minutes, with features extracted off-line. Junsong Yuan 0001, Wei Wang 0040, Jingjing Meng, Ying Wu 0001, Dongge Li |
ACM Multimedia | 3 |