VLDB 2026 Research / reviewers in the wild / expert
Zhanning Gao
dblp:153/2320
· DBLP profile ↗
26ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0003-2031-2805ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous DrivingabstractIn light of the dynamic nature of autonomous driving environments and stringent safety requirements, general MLLMs combined with CLIP alone often struggle to accurately represent driving-specific scenarios, particularly in complex interactions and long-tail cases. To address this, we propose the Hints of Prompt (HoP) framework, which introduces three key enhancements: Affinity hint to emphasize instance-level structure by strengthening token-wise connections, Semantic hint to incorporate high-level information relevant to driving-specific cases, such as complex interactions among vehicles and traffic signs, and Question hint to align visual features with the query context, focusing on question-relevant regions. These hints are fused through a Hint Fusion module, enriching visual representations by capturing driving-related representations with limited domain data, ensuring faster adaptation to driving scenarios. Extensive experiments confirm the effectiveness of the HoP framework, showing that it significantly outperforms previous state-of-the-art methods in all key metrics. Zhanning Gao, Maosheng Ye, Qifeng Chen 0001, Tongyi Cao, Honggang Qi |
ICCV | 2 |
| 2021 | Noise-Resistant Deep Metric Learning With Ranking-Based Instance SelectionabstractThe existence of noisy labels in real-world data negatively impacts the performance of deep learning models. Although much research effort has been devoted to improving robustness to noisy labels in classification tasks, the problem of noisy labels in deep metric learning (DML) remains open. In this paper, we propose a noise-resistant training technique for DML, which we name Probabilistic Ranking-based Instance Selection with Memory (PRISM). PRISM identifies noisy data in a minibatch using average similarity against image features extracted by several previous versions of the neural network. These features are stored in and retrieved from a memory bank. To alleviate the high computational cost brought by the memory bank, we introduce an acceleration method that replaces individual data points with the class centers. In extensive comparisons with 12 existing approaches under both synthetic and real-world label noise, PRISM demonstrates superior performance of up to 6.06% in Precision@1. Chang Liu 0040, Han Yu 0001, Boyang Li 0001, Zhiqi Shen 0001, Zhanning Gao, Peiran Ren, Xuansong Xie, Li-Zhen Cui 0001, Chunyan Miao |
CVPR | 5 |
| 2021 | Enhancing Viewing Experience of Generated Visual Storylines for Promotional VideosabstractVisual storyline generation is the problem of selecting and sequencing a set of visual materials (i.e. images and video clips) to produce a video to elicit certain cognitive or emotional responses from viewers. In this paper, we enhance the viewing experience of generated visual storylines with the Shot Composition, Selection and Plotting (ShotCSP) approach. Designed for generating promotional videos in ecommerce settings, ShotCSP considers three key film-making principles into the visual storyline generation pipeline: a) proximity-aware scene transition, b) sound logic flow, and c) graphic discontinuity. We propose two novel metrics to enhance viewing experience: 1) Semantic Distance, which measures how related a shot is to the product being promoted; and 2) Salient Region Ratio, which estimates attention to product details in a shot. Through large-scale user evaluation involving 1,748 pairwise comparisons against five state-of-the-art approaches, ShotCSP achieves significantly improved viewing experience. It is a promising approach to enable AI generated promotional videos to benefit e-commerce businesses. Chang Liu 0040, Han Yu 0001, Zhiqi Shen 0001, Ian Dixon, Yingxue Yu, Zhanning Gao, Pan Wang 0008, Peiran Ren, Xuansong Xie, Li-Zhen Cui 0001, Chunyan Miao |
ICME | 6 |
| 2021 | Intrinsic Temporal Regularization for High-resolution Human Video SynthesisabstractFashion video synthesis has attracted increasing attention due to its huge potential in immersive media, virtual reality and online retail applications, yet traditional 3D graphic pipelines often require extensive manual labor on data capture and model rigging. In this paper, we investigate an image-based approach to this problem that generates a fashion video clip from a still source image of the desired outfit, which is then rigged in a framewise fashion under the guidance of a driving video. A key challenge for this task lies in the modeling of feature transformation across source and driving frames, where fine-grained transform helps promote visual details at garment regions, but often at the expense of intensified temporal flickering. To resolve this dilemma, we propose a novel framework with 1) a multi-scale transform estimation and feature fusion module to preserve fine-grained garment details, and 2) an intrinsic regularization loss to enforce temporal consistency of learned transform between adjacent frames. Our solution is capable of generating 512\times512 fashion videos with rich garment details and smooth fabric movements beyond existing results. Extensive experiments over the FashionVideo benchmark dataset have demonstrated the superiority of the proposed framework over several competitive baselines. Lingbo Yang, Zhanning Gao, Siwei Ma 0001, Wen Gao 0001 |
ACM Multimedia | 2 |
| 2021 | Towards Fine-Grained Human Pose Transfer With Detail Replenishing NetworkabstractHuman pose transfer (HPT) is an emerging research topic with huge potential in fashion design, media production, online advertising and virtual reality. For these applications, the visual realism of fine-grained appearance details is crucial for production quality and user engagement. However, existing HPT methods often suffer from three fundamental issues: detail deficiency, content ambiguity and style inconsistency, which severely degrade the visual quality and realism of generated images. Aiming towards real-world applications, we develop a more challenging yet practical HPT setting, termed as Fine-grained Human Pose Transfer (FHPT), with a higher focus on semantic fidelity and detail replenishment. Concretely, we analyze the potential design flaws of existing methods via an illustrative example, and establish the core FHPT methodology by combing the idea of content synthesis and feature transfer together in a mutually-guided fashion. Thereafter, we substantiate the proposed methodology with a Detail Replenishing Network (DRN) and a corresponding coarse-to-fine model training scheme. Moreover, we build up a complete suite of fine-grained evaluation protocols to address the challenges of FHPT in a comprehensive manner, including semantic analysis, structural detection and perceptual quality assessment. Extensive experiments on the DeepFashion benchmark dataset have verified the power of proposed benchmark against start-of-the-art works, with 12%-14% gain on top-10 retrieval recall, 5% higher joint localization accuracy, and near 40% gain on face identity preservation. Our codes, models and evaluation tools will be released at https://github.com/Lotayou/RATE. Lingbo Yang, Pan Wang 0008, Chang Liu 0047, Zhanning Gao, Peiran Ren, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Xian-Sheng Hua 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Generating Engaging Promotional Videos for E-commerce Platforms (Student Abstract)abstractThere is an emerging trend for sellers to use videos to promote their products on e-commerce platforms such as Taobao.com. Current video production workflow includes the production of visual storyline by human directors. We propose a system to automatically generate visual storyline based on the input set of visual materials (e.g. video clips or still images) and then produce a promotional video. In particular, we propose an algorithm called Shot Composition, Selection and Plotting (ShotCSP), which generates visual storylines leveraging film-making principles to improve viewing experience and perceived persuasiveness. Chang Liu 0040, Han Yu 0001, Zhiqi Shen 0001, Yingxue Yu, Ian Dixon, Zhanning Gao, Pan Wang 0008, Peiran Ren, Xuansong Xie, Chunyan Miao |
AAAI | 7 |
| 2020 | Ladder Loss for Coherent Visual-Semantic EmbeddingabstractFor visual-semantic embedding, the existing methods normally treat the relevance between queries and candidates in a bipolar way – relevant or irrelevant, and all “irrelevant” candidates are uniformly pushed away from the query by an equal margin in the embedding space, regardless of their various proximity to the query. This practice disregards relatively discriminative information and could lead to suboptimal ranking in the retrieval results and poorer user experience, especially in the long-tail query scenario where a matching candidate may not necessarily exist. In this paper, we introduce a continuous variable to model the relevance degree between queries and multiple candidates, and propose to learn a coherent embedding space, where candidates with higher relevance degrees are mapped closer to the query than those with lower relevance degrees. In particular, the new ladder loss is proposed by extending the triplet loss inequality to a more general inequality chain, which implements variable push-away margins according to respective relevance degrees. In addition, a proper Coherent Score metric is proposed to better measure the ranking results including those “irrelevant” candidates. Extensive experiments on multiple datasets validate the efficacy of our proposed method, which achieves significant improvement over existing state-of-the-art methods. Zhenxing Niu, Le Wang 0003, Zhanning Gao, Qilin Zhang 0004, Gang Hua 0001 |
AAAI | 4 |
| 2020 | Region-Adaptive Texture Enhancement For Detailed Person Image SynthesisabstractThe ability to produce convincing textural details is essential for the fidelity of synthesized person images. Existing methods typically follow a “warping-based” strategy that propagates appearance features through the same pathway used for pose transfer. However, most fine-grained features would be lost during down-sampling, leading to over-smoothed clothes and missing details in the output images. In this paper we presents RATE-Net, a novel framework for synthesizing person images with sharp texture details. The proposed framework leverages an additional texture enhancing module to extract appearance information from the source image and estimate a fine-grained residual texture map, which helps to refine the coarse estimation from the pose transfer module. In addition, we design an effective alternate updating strategy to promote mutual guidance between two modules for better shape and appearance consistency. Experiments conducted on DeepFashion benchmark dataset have demonstrated the superiority of our framework compared with existing networks. Lingbo Yang, Pan Wang 0008, Xinfeng Zhang 0001, Shanshe Wang, Zhanning Gao, Peiran Ren, Xuansong Xie, Siwei Ma 0001, Wen Gao 0001 |
ICME | 5 |
| 2020 | An AI-empowered Visual Storyline GeneratorabstractVideo editing is currently a highly skill- and time-intensive process. One of the most important tasks in video editing is to compose the visual storyline. This paper outlines Visual Storyline Generator (VSG), an artificial intelligence (AI)-empowered system that automatically generates visual storylines based on a set of images and video footages provided by the user. It is designed to produce engaging and persuasive promotional videos with an easy-to-use interface. In addition, users can be involved in refining the AI-generated visual storylines. The editing results can be used as training data to further improve the AI algorithms in VSG. Chang Liu 0040, Zhao Yong Lim, Han Yu 0001, Zhiqi Shen 0001, Ian Dixon, Zhanning Gao, Pan Wang 0008, Peiran Ren, Xuansong Xie, Li-Zhen Cui 0001, Chunyan Miao |
IJCAI | 6 |
| 2020 | Action Co-localization in an Untrimmed Video by Graph Neural Networks
Changbo Zhai, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
MMM (1) | 4 |
| 2020 | EleAtt-RNN: Adding Attentiveness to Neurons in Recurrent Neural NetworksabstractRecurrent neural networks (RNNs) are capable of modeling temporal dependencies of complex sequential data. In general, current available structures of RNNs tend to concentrate on controlling the contributions of current and previous information. However, the exploration of different importance levels of different elements within an input vector is always ignored. We propose a simple yet effective Element-wise-Attention Gate (EleAttG), which can be easily added to an RNN block (e.g. all RNN neurons in an RNN layer), to empower the RNN neurons to have attentiveness capability. For an RNN block, an EleAttG is used for adaptively modulating the input by assigning different levels of importance, i.e., attention, to each element/dimension of the input. We refer to an RNN block equipped with an EleAttG as an EleAtt-RNN block. Instead of modulating the input as a whole, the EleAttG modulates the input at fine granularity, i.e., element-wise, and the modulation is content adaptive. The proposed EleAttG, as an additional fundamental unit, is general and can be applied to any RNN structures, e.g., standard RNN, Long Short-Term Memory (LSTM), or Gated Recurrent Unit (GRU). We demonstrate the effectiveness of the proposed EleAtt-RNN by applying it to different tasks including the action recognition, from both skeleton-based data and RGB videos, gesture recognition, and sequential MNIST classification. Experiments show that adding attentiveness through EleAttGs to RNN blocks significantly improves the power of RNNs. Pengfei Zhang 0005, Jianru Xue, Cuiling Lan, Wenjun Zeng 0001, Zhanning Gao, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Video Imprint Segmentation for Temporal Action Detection in Untrimmed VideosabstractWe propose a temporal action detection by spatial segmentation framework, which simultaneously categorize actions and temporally localize action instances in untrimmed videos. The core idea is the conversion of temporal detection task into a spatial semantic segmentation task. Firstly, the video imprint representation is employed to capture the spatial/temporal interdependences within/among frames and represent them as spatial proximity in a feature space. Subsequently, the obtained imprint representation is spatially segmented by a fully convolutional network. With such segmentation labels projected back to the video space, both temporal action boundary localization and per-frame spatial annotation can be obtained simultaneously. The proposed framework is robust to variable lengths of untrimmed videos, due to the underlying fixed-size imprint representations. The efficacy of the framework is validated in two public action detection datasets. Zhanning Gao, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 1 |
| 2019 | Object Affordances Graph Network for Action Recognition
Haoliang Tan, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Nanning Zheng 0001, Gang Hua 0001 |
BMVC | 4 |
| 2019 | Generating Persuasive Visual Storylines for Promotional VideosabstractVideo contents have become a critical tool for promoting products in E-commerce. However, the lack of automatic promotional video generation solutions makes large-scale video-based promotion campaigns infeasible. The first step of automatically producing promotional videos is to generate visual storylines, which is to select the building block footage and place them in an appropriate order. This task is related to the subjective viewing experience. It is hitherto performed by human experts and thus, hard to scale. To address this problem, we propose WundtBackpack, an algorithmic approach to generate storylines based on available visual materials, which can be video clips or images. It consists of two main parts, 1) the Learnable Wundt Curve to evaluate the perceived persuasiveness based on the stimulus intensity of a sequence of visual materials, which only requires a small volume of data to train; and 2) a clustering-based backpacking algorithm to generate persuasive sequences of visual materials while considering video length constraints. In this way, the proposed approach provides a dynamic structure to empower artificial intelligence (AI) to organize video footage in order to construct a sequence of visual stimuli with persuasive power. Extensive real-world experiments show that our approach achieves close to 10% higher perceived persuasiveness scores by human testers, and 12.5% higher expected revenue compared to the best performing state-of-the-art approach. Chang Liu 0040, Han Yu 0001, Zhiqi Shen 0001, Zhanning Gao, Pan Wang 0008, Changgong Zhang, Peiran Ren, Xuansong Xie, Li-Zhen Cui 0001, Chunyan Miao |
CIKM | 5 |
| 2019 | Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksabstractWeakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video tags as video-level labels. However, such coarse video-level supervision inevitably incurs confusions, especially in untrimmed videos containing multiple action instances. To address this challenge, we propose the Contrast-based Localization EvaluAtioN Network (CleanNet) with our new action proposal evaluator, which provides pseudo-supervision by leveraging the temporal contrast in snippet-level action classification predictions. Essentially, the new action proposal evaluator enforces an additional temporal contrast constraint so that high-evaluation-score action proposals are more likely to coincide with true action instances. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Experiments on THUMOS14 and ActivityNet datasets validate the efficacy of CleanNet against existing state-ofthe- art WS-TAL algorithms. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 4 |
| 2019 | Personalized Video Summarization with Idiom AdaptationabstractShort videos are becoming key for media consumers exploring the TV and Internet. The production of short videos however remains costly. In this paper, we present a domain specific video summarization application with idiom adaptation that leverages multimedia content analysis and insights from cinematic and persuasive domains. From the back-end, content curators can push raw materials and the pre-processing algorithms will automatically extract the features and encode them as editing idioms. Users can create personalized video summaries based on these idioms. We have validated the effectiveness of the demonstration on a TVC data-set with over 600 videos, enabling production of domain specific video summaries with combinations of editing idioms. This approach has been put into trial in the testbed at Alibaba Wood. Chang Liu 0040, Zhiqi Shen 0001, Han Yu 0001, Zhanning Gao, Pan Wang 0008, Changgong Zhang, Peiran Ren, Xuansong Xie |
ACM Multimedia | 5 |
| 2019 | Domain Specific and Idiom Adaptive Video SummarizationabstractAs short videos become an increasingly popular form of storytelling, there is a growing demand for video summarization to convey information concisely with a subset of video frames. Some criteria such as interestingness and diversity are used by existing efforts to pick appropriate segments of content. However, there lacks a mechanism to infuse insights from cinematography and persuasion into this process. As a result, the results of the video summarization sometimes deviate from the original. In addition, the exploration of the vast design space to create customized video summaries is costly for video producer. To address these challenges, we propose a domain specific and idiom adaptive video summarization approach. Specifically, our approach first segments the input video and extracts high-level information from each segment. Such labels are used to represent a collection of idioms and summarization metrics as submodular components which users can combine to create personalized summary styles in a variety of ways. In order to identify the importance of the idioms and metrics in different domains, we leverage max margin learning. Experimental results have validated the effectiveness of our approach. We also plan to release a dataset containing over 600 videos with expert annotations which can benefit further research in this area. Chang Liu 0040, Zhiqi Shen 0001, Zhanning Gao, Pan Wang 0008, Changgong Zhang, Peiran Ren, Xuansong Xie, Han Yu 0001, Qingming Huang |
MMAsia | 4 |
| 2019 | Video ImprintabstractA new unified video analytics framework (ER3) is proposed for complex event retrieval, recognition and recounting, based on the proposed video imprint representation, which exploits temporal correlations among image features across video frames. With the video imprint representation, it is convenient to reverse map back to both temporal and spatial locations in video frames, allowing for both key frame identification and key areas localization within each frame. In the proposed framework, a dedicated feature alignment module is incorporated for redundancy removal across frames to produce the tensor representation, i.e., the video imprint. Subsequently, the video imprint is individually fed into both a reasoning network and a feature aggregation module, for event recognition/recounting and event retrieval tasks, respectively. Thanks to its attention mechanism inspired by the memory networks used in language modeling, the proposed reasoning network is capable of simultaneous event category recognition and localization of the key pieces of evidence for event recounting. In addition, the latent structure in our reasoning network highlights the areas of the video imprint, which can be directly used for event recounting. With the event retrieval task, the compact video representation aggregated from the video imprint contributes to better retrieval results than existing state-of-the-art methods. Zhanning Gao, Le Wang 0003, Nebojsa Jojic, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Adding Attentiveness to the Neurons in Recurrent Neural Networks
Pengfei Zhang 0005, Jianru Xue, Cuiling Lan, Wenjun Zeng 0001, Zhanning Gao, Nanning Zheng 0001 |
ECCV (9) | 5 |
| 2018 | Large-scale vocabularies with local graph diffusion and mode seeking
Shanmin Pang, Jianru Xue, Zhanning Gao, Lihong Zheng, Li Zhu 0003 |
Signal Process. Image Commun. | 3 |
| 2017 | ER3: A Unified Framework for Event Retrieval, Recognition and RecountingabstractWe develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alignment procedure identifies and removes the feature redundancies across frames and outputs an intermediate tensor representation we call video imprint. The video imprint is then fed into a reasoning network, whose attention mechanism parallels that of memory networks used in language modeling. The reasoning network simultaneously recognizes the event category and locates the key pieces of evidence for event recounting. In event retrieval tasks, we show that the compact video representation aggregated from the video imprint achieves significantly better retrieval accuracy compared with existing methods. We also set new state of the art results in event recognition tasks with an additional benefit: The latent structure in our reasoning network highlights the areas of the video imprint and can be directly used for event recounting. As video imprint maps back to locations in the video frames, the network allows not only the identification of key frames but also specific areas inside each frame which are most influential to the decision process. Zhanning Gao, Gang Hua 0001, Dongqing Zhang, Nebojsa Jojic, Le Wang 0003, Jianru Xue, Nanning Zheng 0001 |
CVPR | 1 |
| 2016 | Democratic Diffusion Aggregation for Image RetrievalabstractContent-based image retrieval is an important research topic in the multimedia field. In large-scale image search using local features, image features are encoded and aggregated into a compact vector to avoid indexing each feature individually. In the aggregation step, sum-aggregation is wildly used in many existing works and demonstrates promising performance. However, it is based on a strong and implicit assumption that the local descriptors of an image are identically and independently distributed in descriptor space and image plane. To address this problem, we propose a new aggregation method named democratic diffusion aggregation (DDA) with weak spatial context embedded. The main idea of our aggregation method is to re-weight the embedded vectors before sum-aggregation by considering the relevance among local descriptors. Different from previous work, by conducting a diffusion process on the improved kernel matrix, we calculate the weighting coefficients more efficiently without any iterative optimization. Besides considering the relevance of local descriptors from different images, we also discuss an efficient query fusion strategy which uses the initial top-ranked image vectors to enhance the retrieval performance. Experimental results show that our aggregation method exhibits much higher efficiency (about × 14 faster) and better retrieval accuracy compared with previous methods, and the query fusion strategy consistently improves the retrieval quality. Zhanning Gao, Jianru Xue, Wengang Zhou 0001, Shanmin Pang, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2015 | Fast Democratic Aggregation and Query Fusion for Image SearchabstractIn image search using local features, to avoid indexing each feature individually, encoding methods are popularly adopted to embed and aggregate local features of an image into a compact vector. Democratic aggregation with triangulation embedding (T-embedding) exhibits significant retrieval accuracy improvement over previous works. However, it suffers high computational complexity. To address this problem and consistently improve the retrieval performance, we propose a new democratic method to accelerate aggregating step without accuracy lost. We also embed weak spatial context in the kernel construction to depress co-occurrence caused by local feature detector. Furthermore, we enhance the retrieval performance with an efficient query fusion strategy. The evaluation on public datasets shows that our democratic aggregation is an order of magnitude faster than the original democratic aggregation with comparable retrieval accuracy, and the query fusion achieves a significant accuracy improvement over previous works. Zhanning Gao, Jianru Xue, Wengang Zhou 0001, Shanmin Pang, Qi Tian 0001 |
ICMR | 1 |
| 2015 | Image re-ranking with an alternating optimization
Shanmin Pang, Jianru Xue, Zhanning Gao, Qi Tian 0001 |
Neurocomputing | 3 |
| 2014 | Image Re-ranking with an Alternating OptimizationabstractIn this work, we propose an efficient image re-ranking method, without additional memory cost compared with the baseline method~\cite{philbin2007object}, to re-rank all retrieved images. The motivation of the proposed method is that, there are usually many visual words in the query image that only give votes to irrelevant images. With this observation, we propose to only use visual words which can help to find relevant images to re-rank the retrieved images. To achieve the goal, we first find some similar images to the query by maximizing a quadratic function when given an initial ranking of the retrieved images. Then we select query visual words with an alternating optimization strategy: (1) at each iteration, select words based on the similar images that we have found and (2) in turn, update the similar images with the selected words. These two steps are repeated until convergence. Experimental results on standard benchmark datasets show that the proposed method outperforms spatial based re-ranking methods. Shanmin Pang, Jianru Xue, Zhanning Gao, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2014 | Joint Segmentation and Recognition of Categorized Objects From Noisy Web Image CollectionabstractThe segmentation of categorized objects addresses the problem of joint segmentation of a single category of object across a collection of images, where categorized objects are referred to objects in the same category. Most existing methods of segmentation of categorized objects made the assumption that all images in the given image collection contain the target object. In other words, the given image collection is noise free. Therefore, they may not work well when there are some noisy images which are not in the same category, such as those image collections gathered by a text query from modern image search engines. To overcome this limitation, we propose a method for automatic segmentation and recognition of categorized objects from noisy Web image collections. This is achieved by cotraining an automatic object segmentation algorithm that operates directly on a collection of images, and an object category recognition algorithm that identifies which images contain the target object. The object segmentation algorithm is trained on a subset of images from the given image collection which are recognized to contain the target object with high confidence, while training the object category recognition model is guided by the intermediate segmentation results obtained from the object segmentation algorithm. This way, our co-training algorithm automatically identifies the set of true positives in the noisy Web image collection, and simultaneously extracts the target objects from all the identified images. Extensive experiments validated the efficacy of our proposed approach on four datasets: 1) the Weizmann horse dataset, 2) the MSRC object category dataset, 3) the iCoseg dataset, and 4) a new 30-categories dataset including 15,634 Web images with both hand-annotated category labels and ground truth segmentation labels. It is shown that our method compares favorably with the state-of-the-art, and has the ability to deal with noisy image collections. Le Wang 0003, Gang Hua 0001, Jianru Xue, Zhanning Gao, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 4 |