EDBT 2026 Demo / reviewers in the wild / expert
Wanlei Zhao
dblp:27/5322 · also Wan-Lei Zhao
· DBLP profile ↗
37ranked-venue papers
12as first author
16since 2021 · last 2026
0000-0002-7915-447XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards the Distributed Large-Scale $k$-NN Graph Construction by Graph MergeabstractIn order to support the real-time interaction with LLMs and the instant search or the instant recommendation on social media, it becomes an imminent problem to build a k-NN graph or an indexing graph for the massive number of vectorized multimedia data. In such scenarios, the scale of the data or the scale of the graph may exceed the processing capacity of a single machine. This paper aims to address the graph construction problem of such scale via efficient graph merge. For the graph construction on a single node, two generic and highly parallelizable algorithms, namely Two-way Merge and Multi-way Merge are proposed to merge subgraphs into one. For the graph construction across multiple nodes, a multi-node procedure based on Two-way Merge is presented. The procedure makes it feasible to construct a large-scale k-NN graph/indexing graph on either a single node or multiple nodes when the data size exceeds the memory capacity of one node. Extensive experiments are conducted on both large-scale k-NN graph and indexing graph construction. For the k-NN graph construction, the large-scale and high-quality k-NN graphs are constructed by graph merge in parallel. Typically, a billion-scale k-NN graph can be built in approximately 17h when only three nodes are employed. For the indexing graph construction, similar NN search performance as the original indexing graph is achieved with the merged indexing graphs while requiring much less time of construction. Wanlei Zhao, Shihai Xiao, Jiajie Yao, Xuecang Zhang |
ICDE | 2 |
| 2025 | Highly Efficient Disk-based Nearest Neighbor Search on Extended Neighborhood GraphabstractNearest neighbor search (NN search) plays a fundamental role in many disciplines. According to recent studies, graph-based search methods show superior performance over other types of methods. In order to accommodate the high dimensionality as well as the growing data-scale, the disk-based NN search in which the index graph and the full-precision vectors are kept in SSD has become a promising direction. This paper optimizes the disk-based NN search from three perspectives. Firstly, an eXtended Neighborhood Graph (XN-Graph) structure is proposed. In contrast to the existing index graphs, the out-edges of the graph neighborhood are collected from much wider coverage of the data space. It therefore reduces the number of hops during NN search, which in turn reduces the search latency. Additionally, a dataset partitioning method called Boundary-adaptive Balanced Partition is proposed to facilitate the graph construction in cases where the system cannot handle large datasets in a single round. Moreover, an efficient hybrid NN search method called In-Memory First Search is proposed. Compared to the existing methods, it considerably reduces the CPU idle times. With the support of XN-Graph, it shows 1.5-3 times lower search latency than SOTA methods. On billion-scale datasets, its QPS is still above 4000 when Recall@10 is as high as 0.9. Jianzhi Wang, Wanlei Zhao, Shihai Xiao |
SIGIR | 3 |
| 2025 | Dynamic NN-Descent: An Efficient k-NN Graph Construction MethodabstractAs a classick-NN graph construction method, NN-Descent has been adopted in various applications for its simplicity, genericness, and efficiency. However, its memory consumption is high due to the employment of two extra supporting graph structures. In this paper, a novelk-NN graph construction method is proposed. Similar to NN-Descent, thek-NN graph is constructed by doing cross-matching continuously on the sampled neighbors on each neighborhood. Whereas different from NN-Descent, the cross-matching is undertaken directly on thek-NN graph under construction. It makes the extra graph structures adopted to support the cross-matching no longer necessary. Moreover, no synchronization between different threads is needed within one iteration. The high-quality graph is constructed at the high-speed efficiency and considerably better memory efficiency over NN-Descent on both the multi-thread CPU and the GPU. Jie-Feng Wang, Wanlei Zhao, Shihai Xiao, Jiajie Yao, Xuecang Zhang |
IEEE Trans. Big Data | 2 |
| 2023 | Decoupled Mutual Distillation for Incremental Object DetectionabstractVisual object detection aims to localize and classify objects in the images. Popular detectors perform well when all the object categories are pre-defined. However, they show poor performance when the new categories join in the training incrementally and the replay of old training samples is not allowed. This issue is widely known as catastrophic forgetting. In this paper, the detection head of classic Faster R-CNN is decoupled into the classification head and localization head. Based on the modified network, a dynamic mutual distillation framework is designed. The distillation is applied only to the classification head. On the one hand, this design allows the knowledge to be transferred from both the old model and the assistant model to the new. On the other hand, the degree of knowledge-transfer from either the old or assistant model is dynamically determined by their expertise on a certain category. Compared to the conventional teacher-student distillation, our design makes a better balance between the inheritance of knowledge from the old model and adapting to the new categories. Gao-Dong Liu, Wanlei Zhao |
ICME | 2 |
| 2023 | Shape-Aware Monocular 3D Object DetectionabstractThe 3D object detection is the key issue in the autonomous driving system. This issue is particularly challenging when the detection only relies on a single perspective camera. The anchor-free and keypoint-based models receive increasing attention recently due to their effectiveness and simplicity. However, most of these methods are vulnerable to the occlusion and truncation of objects. In this paper, a single-stage monocular 3D object detection model is proposed. An instance-segmentation head is integrated into the model training, which allows the model to be aware of the visible shape of a target object. Therefore, the detection largely avoids interference from irrelevant regions surrounding the target objects. In addition, we also reveal that the popular IoU-based evaluation metrics, which were originally designed for evaluating stereo or LiDAR-based detection methods, are insensitive to the improvement achieved by the monocular 3D object detection algorithms. A novel evaluation metric, namely average depth similarity (ADS) is proposed for the monocular 3D object detection models. Our method outperforms the comparison baseline in terms of both the popular and the proposed evaluation metrics while maintaining real-time efficiency. Wei Chen 0078, Wanlei Zhao, Song-Yuan Wu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Online Deep Metric Learning via Mutual DistillationabstractDeep metric learning aims to transform input data into an embedding space, where similar samples are close while dissimilar samples are far apart from each other. In practice, samples of new categories arrive incrementally, which requires the periodical augmentation of the learned model. The fine-tuning on the new categories usually leads to poor performance on the old, which is known as “catastrophic forgetting”. Existing solutions either retrain the model from scratch or require the replay of old samples during the training. In this paper, a complete online deep metric learning framework is proposed based on mutual distillation for both one-task and multi-task scenarios. Different from the teacher-student framework, the proposed approach treats the old and new learning tasks with equal importance. No preference over the old or new knowledge is caused. In addition, a novel virtual feature estimation approach is proposed to recover the features assumed to be extracted by the old models. It allows the distillation between the new and the old models without the replay of old training samples or the holding of old models during the training. A comprehensive study shows the superior performance of our approach with the support of different backbones. Gao-Dong Liu, Wanlei Zhao |
ICME | 2 |
| 2022 | On the Merge of k-NN Graphabstractk-nearest neighbor graph is a fundamental data structure in many disciplines such as information retrieval, data-mining, pattern recognition, and machine learning, etc. In the literature, considerable research has been focusing on how to efficiently build an approximatek-nearest neighbor graph (k-NN graph) for a fixed dataset. Unfortunately, a closely related issue of how to merge two existingk-NN graphs has been overlooked. In this paper, we address the issue ofk-NN graph merging in two different scenarios. In the first scenario, a symmetric merge algorithm is proposed to combine two approximatek-NN graphs. The algorithm facilitates large-scale processing by the efficient merging ofk-NN graphs that are produced in parallel. In the second scenario, a joint merge algorithm is proposed to expand an existingk-NN graph with a raw dataset. The algorithm enables the incremental construction of a hierarchical approximatek-NN graph. Superior performance is attained when leveraging the hierarchy for NN search of various data types, dimensionality, and distance measures. Wanlei Zhao, Peng-Cheng Lin, Chong-Wah Ngo |
IEEE Trans. Big Data | 1 |
| 2022 | Approximate k-NN Graph Construction: A Generic Online ApproachabstractNearest neighbor search andk-nearest neighbor graph construction are two fundamental issues that arise from many disciplines such as multimedia information retrieval, data-mining, and machine learning. They become more and more imminent given the big data emerge in various fields in recent years. In this paper, a simple but effective solution both for approximatek-nearest neighbor search and approximatek-nearest neighbor graph construction is presented. These two issues are addressed jointly in our solution. On one hand, the approximatek-nearest neighbor graph construction is treated as a search task. Each sample along with itsk-nearest neighbors is joined into thek-nearest neighbor graph by performing the nearest neighbor search sequentially on the graph under construction. On the other hand, the builtk-nearest neighbor graph is used to supportk-nearest neighbor search. Since the graph is built online, the dynamic update on the graph, which is not possible for most of the existing solutions, is supported. This solution is feasible for various distance measures. Its effectiveness both ask-nearest neighbor construction andk-nearest neighbor search approaches is verified across different types of data in different scales, various dimensions, and under different metrics. Wanlei Zhao, Chong-Wah Ngo |
IEEE Trans. Multim. | 1 |
| 2022 | Deeply Activated Salient Region for Instance SearchabstractThe performance of instance search relies heavily on the ability to locate and describe a wide variety of object instances in a video/image collection. Due to the lack of a proper mechanism for locating instances and deriving feature representation, instance search is generally only effective when the instances are from known object categories. In this article, a simple but effective instance-level feature representation approach is presented. Different from the existing approaches, the issues of class-agnostic instance localization and distinctive feature representation are considered. The former is achieved by detecting salient instance regions from an image by a layer-wise back-propagation process. The back-propagation starts from the last convolution layer of a pre-trained CNNs that is originally used for classification. The back-propagation proceeds layer by layer until it reaches the input layer. This allows the salient instance regions in the input image from both known and unknown categories to be activated. Each activated salient region covers the full or, more usually, a major range of an instance. The distinctive feature representation is produced by average-pooling on the feature map of a certain layer with the detected instance region. Experiments show that this kind of feature representation demonstrates considerably better performance than most of the existing approaches. Hui-Chu Xiao, Wanlei Zhao, Yi-Geng Hong, Chong-Wah Ngo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Fast k-NN Graph Construction by GPU based NN-DescentabstractNN-Descent is a classic k-NN graph construction approach. It is still widely employed in machine learning, computer vision, and information retrieval tasks due to its efficiency and genericness. However, the current design only works well on CPU. In this paper, NN-Descent has been redesigned to adapt to the GPU architecture. A new graph update strategy called selective update is proposed. It reduces the data exchange between GPU cores and GPU global memory significantly, which is the processing bottleneck under GPU computation architecture. This redesign leads to full exploitation of the parallelism of the GPU hardware. In the meantime, the genericness, as well as the simplicity of NN-Descent, are well-preserved. Moreover, a procedure that allows to k-NN graph to be merged efficiently on GPU is proposed. It makes the construction of high-quality k-NN graphs for out-of-GPU-memory datasets tractable. Our approach is 100-250× faster than the single-thread NN-Descent and is 2.5-5× faster than the existing GPU-based approaches as we tested on million as well as billion scale datasets. Wanlei Zhao, Xiangxiang Zeng, Jianye Yang 0001 |
CIKM | 2 |
| 2021 | k-sums Clustering: A Stochastic Optimization Approach
Wanlei Zhao, Shi-Ying Lan, Run-Qing Chen, Chong-Wah Ngo |
CIKM | 1 |
| 2021 | Towards Accurate Localization by Instance SearchabstractVisual object localization is the key step in a series of object detection tasks. In the literature, high localization accuracy is achieved with the mainstream strongly supervised frameworks. However, such methods require object-level annotations and are unable to detect objects of unknown categories. Weakly supervised methods face similar difficulties. In this paper, a self-paced learning framework is proposed to achieve accurate object localization on the rank list returned by instance search. The proposed framework mines the target instance gradually from the queries and their corresponding top-ranked search results. Since a common instance is shared between the query and the images in the rank list, the target visual instance can be accurately localized even without knowing what the object category is. In addition to performing localization on instance search, the issue of few-shot object detection is also addressed under the same framework. Superior performance over state-of-the-art methods is observed on both tasks. Yi-Geng Hong, Hui-Chu Xiao, Wanlei Zhao |
ACM Multimedia | 3 |
| 2021 | A joint model for IT operation series prediction and anomaly detection
Run-Qing Chen, Guang-Hui Shi, Wanlei Zhao, Chang-Hui Liang |
Neurocomputing | 3 |
| 2021 | Instance search based on weakly supervised feature learning
Jie Lin 0015, Wanlei Zhao |
Neurocomputing | 3 |
| 2021 | Instance search via instance level segmentation and feature representation
Wanlei Zhao |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Dynamic sampling for deep metric learning
Chang-Hui Liang, Wanlei Zhao, Run-Qing Chen |
Pattern Recognit. Lett. | 2 |
| 2019 | Saliency-aware inter-image color transfer for image manipulation
Zhi Liu 0003, Qihan Jiao, Olivier Le Meur, Wanlei Zhao |
Multim. Tools Appl. | 5 |
| 2018 | Fast k-Means Based on k-NN GraphabstractIn the big data era, k-means clustering has been widely adopted as a basic processing tool in various contexts. However, its computational cost could be prohibitively high when the data size and the cluster number are large. The processing bottleneck of k-means lies in the operation of seeking the closest centroid in each iteration. In this paper, a novel solution towards the scalability issue of k-means is presented. In the proposal, k-means is supported by an approximate k-nearest neighbors graph. In the k-means iteration, each data sample is only compared to clusters that its nearest neighbors reside. Since the number of nearest neighbors we consider is much less than k, the processing cost in this step becomes minor and irrelevant to k. The processing bottleneck is therefore broken. The most interesting thing is that k-nearest neighbor graph is constructed by calling the fast k-means itself. Compared with existing fast k-means variants, the proposed algorithm achieves hundreds to thousands times speed-up while maintaining high clustering quality. As it is tested on 10 million 512-dimensional data, it takes only 5.2 hours to produce 1 million clusters. In contrast, it would take 3 years for traditional k-means to fulfill the same scale of clustering. Cheng-Hao Deng, Wanlei Zhao |
ICDE | 2 |
| 2018 | k-means: A revisit
Wanlei Zhao, Cheng-Hao Deng, Chong-Wah Ngo |
Neurocomputing | 1 |
| 2018 | Object Discovery via Cohesion MeasurementabstractColor and intensity are two important components in an image. Usually, groups of image pixels, which are similar in color or intensity, are an informative representation for an object. They are therefore particularly suitable for computer vision tasks, such as saliency detection and object proposal generation. However, image pixels, which share a similar real-world color, may be quite different since colors are often distorted by intensity. In this paper, we reinvestigate the affinity matrices originally used in image segmentation methods based on spectral clustering. A new affinity matrix, which is robust to color distortions, is formulated for object discovery. Moreover, a cohesion measurement (CM) for object regions is also derived based on the formulated affinity matrix. Based on the new CM, a novel object discovery method is proposed to discover objects latent in an image by utilizing the eigenvectors of the affinity matrix. Then we apply the proposed method to both saliency detection and object proposal generation. Experimental results on several evaluation benchmarks demonstrate that the proposed CM-based method has achieved promising performance for these two tasks. Guanjun Guo, Hanzi Wang, Wanlei Zhao, Yan Yan 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2017 | Motion Segmentation Via a Sparsity ConstraintabstractMotion segmentation is an important task for intelligent transportation systems. In this paper, inspired by the fact that a feature point trajectory can be sparsely represented as a combination of several feature point trajectories that share coherent transformations, an efficient and effective motion segmentation method with a sparsity constraint is proposed. Specifically, we first propose an accumulated scheme to efficiently integrate motion information from all the frames of a video sequence to construct a correlation matrix. Then, a sparse affinity matrix is built on the correlation matrix by using information-theoretic principles, where the nonzero elements in the same row of the sparse affinity matrix correspond to the feature point trajectories more likely belonging to the same motion. Thereafter, a segment and merge procedure is proposed to effectively estimate the number of motions via the sparse affinity matrix. Finally, by applying spectral clustering on the sparse affinity matrix, different motions in the video sequence are accurately segmented based on the estimated number of motions. Experimental results on theHopkins 155and the62-clipdatasets demonstrate that the proposed method achieves superior performance compared with several state-of-the-art methods. Taotao Lai, Hanzi Wang, Yan Yan 0001, Tat-Jun Chin, Wanlei Zhao |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2016 | Fast Covariant VLAD for Image SearchabstractVector of locally aggregated descriptor (VLAD) is a popular image encoding approach for its simplicity and better scalability over conventional bag-of-visual-word approach. In order to enhance its distinctiveness and geometric invariance, covariant VLAD (CVLAD) is proposed to pool local features based on their dominant orientations/characteristic scales, which leads to a geometric-aware representation. This representation achieves rotation/scale invariance when being associated with circular matching. However, the circular matching induces several times of computation overhead, which makes CVLAD hardly suitable for large-scale retrieval tasks. In this paper, the issue of computation overhead is alleviated by performing the circular matching in CVLAD's frequency domain. In addition, by operating PCA on CVLAD in its frequency domain, much better scalability is achieved than when it is undertaken in the original feature space. Furthermore, the high-dimensional CVLAD subvectors are converted to dozens of very low-dimensional subvectors, which is possible when transforming the feature into its frequency domain. Nearest neighbor search is therefore undertaken on very low-dimensional subspaces, which becomes easily tractable. The effectiveness of our approach is demonstrated in the retrieval scenario on popular benchmarks comprising up to 1 million database images. Wanlei Zhao, Chong-Wah Ngo, Hanzi Wang |
IEEE Trans. Multim. | 1 |
| 2015 | Affine hull based target representation for visual tracking
Jun Wang 0131, Hanzi Wang, Wanlei Zhao |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | K-means based histogram using multiresolution feature vectors for color texture database retrieval
Cong Bai, Jinglin Zhang 0001, Zhi Liu 0003, Wanlei Zhao |
Multim. Tools Appl. | 4 |
| 2015 | Co-Saliency Detection via Co-Salient Object Discovery and RecoveryabstractThis letter proposes a novel co-saliency model to effectively discover and highlight co-salient objects in a set of images. Based on the gross similarity which combines color features and SIFT descriptors, some co-salient object regions are first discovered in each image as exemplars, which are exploited to generate the exemplar saliency maps with the use of single-image saliency model. Then both local recovery and global recovery of co-salient object regions are performed by propagating the exemplar saliency to the matched regions, and border connectivity is further exploited to generate the region-level co-saliency maps. Finally, the foci of attention area based pixel-level saliency derivation is used to generate the pixel-level co-saliency maps with even better quality. Experimental results on two benchmark datasets demonstrate that the proposed co-saliency model outperforms the state-of-the-art co-saliency models. Linwei Ye, Zhi Liu 0003, Wanlei Zhao, Liquan Shen |
IEEE Signal Process. Lett. | 4 |
| 2014 | A Hamming Embedding Kernel with Informative Bag-of-Visual Words for Video Semantic IndexingabstractIn this article, we propose a novel Hamming embedding kernel with informative bag-of-visual words to address two main problems existing in traditional BoW approaches for video semantic indexing. First, Hamming embedding is employed to alleviate the information loss caused by SIFT quantization. The Hamming distances between keypoints in the same cell are calculated and integrated into the SVM kernel to better discriminate different image samples. Second, to highlight the concept-specific visual information, we propose to weight the visual words according to their informativeness for detecting specific concepts. We show that our proposed kernels can significantly improve the performance of concept detection. Feng Wang 0036, Wanlei Zhao, Chong-Wah Ngo, Bernard Mérialdo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Oriented pooling for dense and non-dense rotation-invariant featuresabstractInternational audience Wanlei Zhao, Guillaume Gravier, Hervé Jégou |
BMVC | 1 |
| 2013 | Sim-min-hash: an efficient matching technique for linking large image collectionsabstractOne of the most successful method to link all similar images within a large collection is min-Hash, which is a way to significantly speed-up the comparison of images when the underlying image representation is bag-of-words. However, the quantization step of min-Hash introduces important information loss. In this paper, we propose a generalization of min-Hash, called Sim-min-Hash, to compare sets of real-valued vectors. We demonstrate the effectiveness of our approach when combined with the Hamming embedding similarity. Experiments on large-scale popular benchmarks demonstrate that Sim-min-Hash is more accurate and faster than min-Hash for similar image search. Linking a collection of one million images described by 2 billion local descriptors is done in 7 minutes on a single core machine. Wanlei Zhao, Hervé Jégou, Guillaume Gravier |
ACM Multimedia | 1 |
| 2013 | Flip-Invariant SIFT for Copy and Object DetectionabstractScale-invariant feature transform (SIFT) feature has been widely accepted as an effective local keypoint descriptor for its invariance to rotation, scale, and lighting changes in images. However, it is also well known that SIFT, which is derived from directionally sensitive gradient fields, is not flip invariant. In real-world applications, flip or flip-like transformations are commonly observed in images due to artificial flipping, opposite capturing viewpoint, or symmetric patterns of objects. This paper proposes a new descriptor, named flip-invariant SIFT (or F-SIFT), that preserves the original properties of SIFT while being tolerant to flips. F-SIFT starts by estimating the dominant curl of a local patch and then geometrically normalizes the patch by flipping before the computation of SIFT. We demonstrate the power of F-SIFT on three tasks: large-scale video copy detection, object recognition, and detection. In copy detection, a framework, which smartly indices the flip properties of F-SIFT for rapid filtering and weak geometric checking, is proposed. F-SIFT not only significantly improves the detection accuracy of SIFT, but also leads to a more than 50% savings in computational cost. In object recognition, we demonstrate the superiority of F-SIFT in dealing with flip transformation by comparing it to seven other descriptors. In object detection, we further show the ability of F-SIFT in describing symmetric objects. Consistent improvement across different kinds of keypoint detectors is observed for F-SIFT over the original SIFT. Wanlei Zhao, Chong-Wah Ngo |
IEEE Trans. Image Process. | 1 |
| 2010 | On the Annotation of Web Videos by Efficient Near-Duplicate SearchabstractWith the proliferation of Web 2.0 applications, user-supplied social tags are commonly available in social media as a means to bridge the semantic gap. On the other hand, the explosive expansion of social web makes an overwhelming number of web videos available, among which there exists a large number of near-duplicate videos. In this paper, we investigate techniques which allow effective annotation of web videos from a data-driven perspective. A novel classifier-free video annotation framework is proposed by first retrieving visual duplicates and then suggesting representative tags. The significance of this paper lies in the addressing of two timely issues for annotating query videos. First, we provide a novel solution for fast near-duplicate video retrieval. Second, based on the outcome of near-duplicate search, we explore the potential that the data-driven annotation could be successful when huge volume of tagged web videos is freely accessible online. Experiments on cross sources (annotating Google videos and Yahoo! videos using YouTube videos) and cross time periods (annotating YouTube videos using historical data) show the effectiveness and efficiency of the proposed classifier-free approach for web video tag annotation. Wanlei Zhao, Xiao Wu 0001, Chong-Wah Ngo |
IEEE Trans. Multim. | 1 |
| 2009 | Large-scale near-duplicate web video search: Challenge and opportunityabstractThe massive amount of near-duplicate and duplicate Web videos has presented both challenge and opportunity to multimedia computing. On one hand, browsing videos on Internet becomes highly inefficient for the need to repeatedly fast-forward videos of similar content. On the other hand, the tremendous amount of somewhat duplicate content also makes some traditionally difficult vision tasks become simple and easy. For example, annotating pictures can be as simple as recycling the tags of Internet images retrieved from image search engines. Such tasks, of either to eliminate or to recycle near-duplicates, can usually be achieved by the nearest neighbor search of videos from Internet. The fundamental problem lies on the scalability of a search technique, in face of the intractable volume of videos which keep rolling on the Web. In this paper, we investigate scalability of several well-known features including color signature and visual keywords for Web-based retrieval. Indexing these features based on embedding technique for scalable retrieval is also presented. On an Internet video dataset of more than 700 hours collected during years 2006 to 2008, we show some preliminary insights to the challenge of scalable retrieval. Wanlei Zhao, Song Tan, Chong-Wah Ngo |
ICME | 1 |
| 2009 | Towards google challenge: combining contextual and social information for web video categorizationabstractWeb video categorization is a fundamental task for web video search. In this paper, we explore the Google challenge from a new perspective by combing contextual and social information under the scenario of social web. The semantic meaning of text (title and tags), video relevance from related videos, and user interest induced from user videos, are integrated to robustly determine the video category. Experiments on YouTube videos demonstrate the effectiveness of the proposed solution. The performance reaches 60% improvement compared to the traditional text based classifiers. Xiao Wu 0001, Wanlei Zhao, Chong-Wah Ngo |
ACM Multimedia | 2 |
| 2009 | Scale-Rotation Invariant Pattern Entropy for Keypoint-Based Near-Duplicate DetectionabstractNear-duplicate (ND) detection appears as a timely issue recently, being regarded as a powerful tool for various emerging applications. In the Web 2.0 environment particularly, the identification of near-duplicates enables the tasks such as copyright enforcement, news topic tracking, image and video search. In this paper, we describe an algorithm, namely Scale-Rotation invariant Pattern Entropy (SR-PE), for the detection of near-duplicates in large-scale video corpus. SR-PE is a novel pattern evaluation technique capable of measuring the spatial regularity of matching patterns formed by local keypoints. More importantly, the coherency of patterns and the perception of visual similarity, under the scenario that there could be multiple ND regions undergone arbitrary transformations, respectively, are carefully addressed through entropy measure. To demonstrate our work in large-scale dataset, a practical framework composed of three components: bag-of-words representation, local keypoint matching and SR-PE evaluation, is also proposed for the rapid detection of near-duplicates. Wanlei Zhao, Chong-Wah Ngo |
IEEE Trans. Image Process. | 1 |
| 2008 | Accelerating near-duplicate video matching by combining visual similarity and alignment distortionabstractIn this paper, we investigate a novel approach to accelerate the matching of two video clips by exploiting the temporal coherence property inherent in the keyframe sequence of a video. Motivated by the fact that keyframe correspondences between near-duplicate videos typically follow certain spatial arrangements, such property could be employed to guide the alignment of two keyframe sequences. We set the alignment problem as an integer quadratic programming problem, where the cost function takes into account both the visual similarity of the corresponding keyframes as well as the alignment distortion among the set of correspondences. The set of keyframe-pairs found by our algorithm provides our proposal on the list of candidate keyframe-pairs for near-duplicate detection using local interest points. This eliminates the need for exhaustive keyframe-pair comparisons, which significantly accelerates the matching speed. Experiments on a dataset of 12,790 web videos demonstrate that the proposed method maintains a similar near-duplicate video retrieval performance as the hierarchical method proposed in [12] but with a significantly reduced number of keyframe-pair comparisons. Hung-Khoon Tan, Xiao Wu 0001, Chong-Wah Ngo, Wanlei Zhao |
ACM Multimedia | 4 |
| 2007 | Efficient Near-Duplicate Keyframe Retrieval with Visual Language ModelsabstractNear-duplicate keyframe retrieval is a critical task for video similarity measure, video threading and tracking. In this paper, instead of using expensive point-to-point matching on keypoints, we investigate the visual language models built on visual keywords to speed up the near-duplicate keyframe retrieval. The main idea is to estimate a visual language model on visual keywords for each keyframe and compare keyframes by the likelihood of their visual language models. Experiments on a subset of TRECVID-2004 video corpus show that visual language models built on visual keywords demonstrate promising performance for near-duplicate keyframe retrieval, which greatly speed up the retrieval speed although sacrifice a little performance compared to expensive point-to-point matching. Xiao Wu 0001, Wanlei Zhao, Chong-Wah Ngo |
ICME | 2 |
| 2007 | Near-Duplicate Keyframe Identification With Interest Point Matching and Pattern LearningabstractThis paper proposes a new approach for near-duplicate keyframe (NDK) identification by matching, filtering and learning of local interest points (LIPs) with PCA-SIFT descriptors. The issues in matching reliability, filtering efficiency and learning flexibility are novelly exploited to delve into the potential of LIP-based retrieval and detection. In matching, we propose a one-to-one symmetric matching (OOS) algorithm which is found to be highly reliable for NDK identification, due to its capability in excluding false LIP matches compared with other matching strategies. For rapid filtering, we address two issues: speed efficiency and search effectiveness, to support OOS with a new index structure called LIP-IS. By exploring the properties of PCA-SIFT, the filtering capability and speed of LIP-IS are asymptotically estimated and compared to locality sensitive hashing (LSH). Owing to the robustness consideration, the matching of LIPs across keyframes forms vivid patterns that are utilized for discriminative learning and detection with support vector machines. Experimental results on TRECVID-2003 corpus show that our proposed approach outperforms other popular methods including the techniques with LSH in terms of retrieval and detection effectiveness. In addition, the proposed LIP-IS successfully speeds up OOS for more than ten times and possesses several avorable properties compared to LSH. Wanlei Zhao, Chong-Wah Ngo, Hung-Khoon Tan, Xiao Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2006 | Fast tracking of near-duplicate keyframes in broadcast domain with transitivity propagationabstractThe identification of near-duplicate keyframe (NDK) pairs is a useful task for a variety of applications such as news story threading and content-based video search. In this paper, we propose a novel approach for the discovery and tracking of NDK pairs and threads in the broadcast domain. The detection of NDKs in a large data set is a challenging task due to the fact that when the data set increases linearly, the computational cost increases in a quadratic speed, and so does the number of false alarms. This paper explores the symmetric and transitive nature of near-duplicate for the effective detection and fast tracking of NDK pairs based upon the matching of local keypoints in frames. In the detection phase, we propose a robust measure, namely pattern entropy (PE), to measure the coherency of symmetric keypoint matching across the space of two keyframes. This measure is shown to be effective in discovering the NDK identity of a frame. In the tracking phase, the NDK pairs and threads are rapidly propagated and linked with sitivity without the need of detection. This step ends up a significant boost in speed efficiency. We evaluate proposed approach against a month of the 2004 broadcast videos. The experimental results indicate our approach outperforms other techniques in terms of recall and precision with a large margin. In addition, by considering the transitivity and the underlying distribution of NDK pairs along time span, a speed up of 3 to 5 times is achieved when keeping the performance close enough to the optimal one obtained by exhaustive evaluation. Chong-Wah Ngo, Wanlei Zhao, Yu-Gang Jiang 0001 |
ACM Multimedia | 2 |