Lin Chen 0021

dblp:13/3479-21 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
4since 2021 · last 2023
0000-0001-6426-6682ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Transfer learning and domain adaptation · 44% Image recognition and object detection · 35% Video understanding and tracking · 20%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 75% Information retrieval · 12% Web and social media mining · 12%
Human-computer interaction and pervasive computing
1 paper
Ubiquitous computing and smart environments · 100%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 24 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems
federated recommendation
0.712023
Wyze Rule: Federated Rule Dataset for Rule Recommendation Benchmarking · NeurIPS 2023
Ubiquitous computing and smart environments › smart home
home automation
0.712023
Wyze Rule: Federated Rule Dataset for Rule Recommendation Benchmarking · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.522018
Visual Recognition in RGB Images and Videos by Learning from RGB-D Data · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Recognizing RGB Images by Learning from RGB-D Data · CVPR 2014
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.512021
Sequential Instance Refinement for Cross-Domain Object Detection in Images · IEEE Trans. Image Process. 2021
Computer vision › Image recognition and object detection › object detection
domain adaptive object detection
0.512021
Sequential Instance Refinement for Cross-Domain Object Detection in Images · IEEE Trans. Image Process. 2021
Computer vision › Image recognition and object detection
object detection
0.512021
Sequential Instance Refinement for Cross-Domain Object Detection in Images · IEEE Trans. Image Process. 2021
Computer vision › Image recognition and object detection
visual recognition
0.522020
TransMatch: A Transfer-Learning Scheme for Semi-Supervised Few-Shot Learning · CVPR 2020
Event Recognition in Videos by Learning from Heterogeneous Web Sources · CVPR 2013
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot image classification
0.412020
TransMatch: A Transfer-Learning Scheme for Semi-Supervised Few-Shot Learning · CVPR 2020
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.412020
TransMatch: A Transfer-Learning Scheme for Semi-Supervised Few-Shot Learning · CVPR 2020
Machine learning › Transfer learning and domain adaptation › few-shot learning
semi-supervised few-shot learning
0.412020
TransMatch: A Transfer-Learning Scheme for Semi-Supervised Few-Shot Learning · CVPR 2020
Computer vision › Video understanding and tracking
action recognition
0.312018
Visual Recognition in RGB Images and Videos by Learning from RGB-D Data · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › Image recognition and object detection
object recognition
0.312018
Visual Recognition in RGB Images and Videos by Learning from RGB-D Data · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › Video understanding and tracking › action recognition › robust action recognition
view-invariant action recognition
0.312018
Visual Recognition in RGB Images and Videos by Learning from RGB-D Data · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › Video understanding and tracking
motion analysis
0.212015
Video Object Segmentation Via Dense Trajectories · IEEE Trans. Multim. 2015
Computer vision › Video understanding and tracking
video object segmentation
0.212015
Video Object Segmentation Via Dense Trajectories · IEEE Trans. Multim. 2015
Computer vision › Image recognition and object detection › object recognition › multimodal object recognition
RGB-D object recognition
0.212014
Recognizing RGB Images by Learning from RGB-D Data · CVPR 2014
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation
0.212013
Event Recognition in Videos by Learning from Heterogeneous Web Sources · CVPR 2013
Multimedia analysis and retrieval
image retrieval
0.112012
Tag-Based Image Retrieval Improved by Augmented Features and Group-Based Refinement · IEEE Trans. Multim. 2012
Multimedia analysis and retrieval › cross-modal retrieval › image-text retrieval
tag-based image retrieval
0.112012
Tag-Based Image Retrieval Improved by Augmented Features and Group-Based Refinement · IEEE Trans. Multim. 2012
Web and social media mining › social tagging
tag refinement
0.112010
Tag-based web photo retrieval improved by batch mode re-tagging · CVPR 2010
Information retrieval › image retrieval
web image search
0.112010
Tag-based web photo retrieval improved by batch mode re-tagging · CVPR 2010
Computer vision › Face, body and person analysis › facial attribute analysis
gender recognition
0.112014
Recognizing RGB Images by Learning from RGB-D Data · CVPR 2014
Computer vision › Video understanding and tracking
event recognition
0.012013
Event Recognition in Videos by Learning from Heterogeneous Web Sources · CVPR 2013
Multimedia analysis and retrieval › interactive retrieval
relevance feedback
0.012012
Tag-Based Image Retrieval Improved by Augmented Features and Group-Based Refinement · IEEE Trans. Multim. 2012

Methods — techniques the papers use, named apart from their topics

federated learning · 1.3centralized learning · 1.3SVM · 0.5reinforcement learning · 0.5feature alignment · 0.5semi-supervised learning · 0.4mixmatch · 0.4meta-learning · 0.4imprinting · 0.4subspace learning · 0.3projection matrix optimization · 0.3inverted file · 0.3graph-based segmentation · 0.2support vector machine · 0.1laplacian regularized least squares · 0.1SVM with augmented features · 0.1
YearPublicationVenuePosition
2023 Wyze Rule: Federated Rule Dataset for Rule Recommendation Benchmarking
abstract
In the rapidly evolving landscape of smart home automation, the potential of IoT devices is vast. In this realm, rules are the main tool utilized for this automation, which are predefined conditions or triggers that establish connections between devices, enabling seamless automation of specific processes. However, one significant challenge researchers face is the lack of comprehensive datasets to explore and advance the field of smart home rule recommendations. These datasets are essential for developing and evaluating intelligent algorithms that can effectively recommend rules for automating processes while preserving the privacy of the users, as it involves personal information about users' daily lives. To bridge this gap, we present the Wyze Rule Dataset, a large-scale dataset designed specifically for smart home rule recommendation research. Wyze Rule encompasses over 1 million rules gathered from a diverse user base of 300,000 individuals from Wyze Labs, offering an extensive and varied collection of real-world data. With a focus on federated learning, our dataset is tailored to address the unique challenges of a cross-device federated learning setting in the recommendation domain, featuring a large-scale number of clients with widely heterogeneous data. To establish a benchmark for comparison and evaluation, we have meticulously implemented multiple baselines in both centralized and federated settings. Researchers can leverage these baselines to gauge the performance and effectiveness of their rule recommendation systems, driving advancements in the domain. The Wyze Rule Dataset is publicly accessible through HuggingFace's dataset API.
Mohammad Mahdi Kamani, Yuhang Yao 0003, Hanjia Lyu, Zhongwei Cheng, Lin Chen 0021, Liangju Li, Carlee Joe-Wong, Jiebo Luo 0001
NeurIPS5
2023 Open Set Domain Adaptation With Soft Unknown-Class Rejection
abstract
The goal of domain adaptation (DA) is to train a good model for a target domain, with a large amount of labeled data in a source domain but only limited labeled data in the target domain. Conventional closed set domain adaptation (CSDA) assumes source and target label spaces are the same. However, this is not quite practical in real-world applications. In this work, we study the problem of open set domain adaptation (OSDA), which only requires the target label space to partially overlap with the source label space. Consequently, the solution to OSDA requires unknown classes detection and separation, which is normally achieved by introducing a threshold for the prediction of target unknown classes; however, the performance can be quite sensitive to that threshold. In this article, we tackle the above issues by proposing a novel OSDA method to perform soft rejection of unknown target classes and simultaneously match the source and target domains. Extensive experiments on three standard datasets validate the effectiveness of the proposed method over the state-of-the-art competitors.
Yiming Xu 0006, Lin Chen 0021, Lixin Duan, Ivor W. Tsang, Jiebo Luo 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 When Few-Shot Learning Meets Video Object Detection
abstract
Different from static images, videos contain additional temporal and spatial information for better object detection. However, it is costly to obtain a large number of videos with bounding box annotations that are required for supervised deep learning. Although humans can easily learn to recognize new objects by watching only a few video clips, deep learning usually suffers from overfitting. This leads to an important question: how to effectively learn a video object detector from only a few labeled video clips? In this paper, we study the new problem of few-shot learning for video object detection. We first define the few-shot setting and create a new benchmark dataset for few-shot video object detection derived from the widely used ImageNet VID dataset. We employ a transfer-learning framework to effectively train the video object detector on a large number of base-class objects and a few video clips of novel-class objects. By analyzing the results of two methods under this framework (Joint and Freeze) on our designed weak and strong base datasets, we reveal insufficiency and overfitting problems. A simple but effective method, called Thaw, is naturally developed to trade off the two problems and validate our analysis. Extensive experiments on our proposed benchmark datasets with different scenarios demonstrate the effectiveness of our novel analysis in this new few-shot video object detection problem.
Zhongjie Yu 0003, Gaoang Wang, Lin Chen 0021, Sebastian Raschka, Jiebo Luo 0001
ICPR3
2021 Sequential Instance Refinement for Cross-Domain Object Detection in Images
abstract
Cross-domain object detection in images has attracted increasing attention in the past few years, which aims at adapting the detection model learned from existing labeled images (source domain) to newly collected unlabeled ones (target domain). Existing methods usually deal with the cross-domain object detection problem through direct feature alignment between the source and target domains at the image level, the instance level (i.e., region proposals) or both. However, we have observed that directly aligning features of all object instances from the two domains often results in the problem of negative transfer, due to the existence of (1) outlier target instances that contain confusing objects not belonging to any category of the source domain and thus are hard to be captured by detectors and (2) low-relevance source instances that are considerably statistically different from target instances although their contained objects are from the same category. With this in mind, we propose a reinforcement learning based method, coined as sequential instance refinement, where two agents are learned to progressively refine both source and target instances by taking sequential actions to remove both outlier target instances and low-relevance source instances step by step. Extensive experiments on several benchmark datasets demonstrate the superior performance of our method over existing state-of-the-art baselines for cross-domain object detection.
Jin Chen 0009, Xinxiao Wu, Lixin Duan, Lin Chen 0021
IEEE Trans. Image Process.4
2020 TransMatch: A Transfer-Learning Scheme for Semi-Supervised Few-Shot Learning
abstract
The successful application of deep learning to many visual recognition tasks relies heavily on the availability of a large amount of labeled data which is usually expensive to obtain. The few-shot learning problem has attracted increasing attention from researchers for building a robust model upon only a few labeled samples. Most existing works tackle this problem under the meta-learning framework by mimicking the few-shot learning task with an episodic training strategy. In this paper, we propose a new transfer-learning framework for semi-supervised few-shot learning to fully utilize the auxiliary information from labeled base-class data and unlabeled novel-class data. The framework consists of three components: 1) pre-training a feature extractor on base-class data; 2) using the feature extractor to initialize the classifier weights for the novel classes; and 3) further updating the model with a semi-supervised learning method. Under the proposed framework, we develop a novel method for semi-supervised few-shot learning called TransMatch by instantiating the three components with imprinting and MixMatch. Extensive experiments on two popular benchmark datasets for few-shot learning, CUB-200-2011 and miniImageNet, demonstrate that our proposed method can effectively utilize the auxiliary information from labeled base-class data and unlabeled novel-class data to significantly improve the accuracy of few-shot learning task, and achieve new state-of-the-art results.
Zhongjie Yu 0003, Lin Chen 0021, Zhongwei Cheng, Jiebo Luo 0001
CVPR2
2020 DAIL: Dataset-Aware and Invariant Learning for Face Recognition
abstract
To achieve good performance in face recognition, a large scale training dataset is usually required. A simple yet effective way to improve the recognition performance is to use a dataset as large as possible by combining multiple datasets in the training. However, it is problematic and troublesome to naively combine different datasets due to two major issues. First, the same person can possibly appear in different datasets, leading to an identity overlapping issue between different datasets. Naively treating the same person as different classes in different datasets during training will affect back-propagation and generate nonrepresentative embeddings. On the other hand, manually cleaning labels may take formidable human efforts, especially when there are millions of images and thousands of identities. Second, different datasets are collected in different situations and thus will lead to different domain distributions. Naively combining datasets will make it difficult to learn domain invariant embeddings across different datasets. In this paper, we propose DAIL: Dataset-Aware and Invariant Learning to resolve the above-mentioned issues. To solve the first issue of identity overlapping, we propose a dataset-aware loss for multi-dataset training by reducing the penalty when the same person appears in multiple datasets. This can be readily achieved with a modified softmax loss with a dataset-aware term. To solve the second issue, domain adaptation with gradient reversal layers is employed for dataset invariant learning. The proposed approach not only achieves the state-of-the-art results on several commonly used face recognition validation sets, including LFW, CFP-FP, and AgeDB-30, but also shows great benefit for practical use.
Gaoang Wang, Lin Chen 0021, Tianqiang Liu, Mingwei He, Jiebo Luo 0001
ICPR2
2018 Visual Recognition in RGB Images and Videos by Learning from RGB-D Data
abstract
In this work, we propose a framework for recognizing RGB images or videos by learning from RGB-D training data that contains additional depth information. We formulate this task as a new unsupervised domain adaptation (UDA) problem, in which we aim to take advantage of the additional depth features in the source domain and also cope with the data distribution mismatch between the source and target domains. To handle the domain distribution mismatch, we propose to learn an optimal projection matrix to map the samples from both domains into a common subspace such that the domain distribution mismatch can be reduced. Such projection matrix can be effectively optimized by exploiting different strategies. Moreover, we also use different ways to utilize the additional depth features. To simultaneously cope with the above two issues, we formulate a unified learning framework called domain adaptation from multi-view to single-view (DAM2S). By defining various forms of regularizers in our DAM2S framework, different strategies can be readily incorporated to learn robust SVM classifiers for classifying the target samples, and three methods are developed under our DAM2S framework. We conduct comprehensive experiments for object recognition, cross-dataset and cross-view action recognition, which demonstrate the effectiveness of our proposed methods for recognizing RGB images and videos by learning from RGB-D data.
Wen Li 0001, Lin Chen 0021, Dong Xu 0001, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Action and Event Recognition in Videos by Learning From Heterogeneous Web Sources
abstract
In this paper, we propose new approaches for action and event recognition by leveraging a large number of freely available Web videos (e.g., from Flickr video search engine) and Web images (e.g., from Bing and Google image search engines). We address this problem by formulating it as a new multi-domain adaptation problem, in which heterogeneous Web sources are provided. Specifically, we are given different types of visual features (e.g., the DeCAF features from Bing/Google images and the trajectory-based features from Flickr videos) from heterogeneous source domains and all types of visual features from the target domain. Considering the target domain is more relevant to some source domains, we propose a new approach named multi-domain adaptation with heterogeneous sources (MDA-HS) to effectively make use of the heterogeneous sources. In MDA-HS, we simultaneously seek for the optimal weights of multiple source domains, infer the labels of target domain samples, and learn an optimal target classifier. Moreover, as textual descriptions are often available for both Web videos and images, we propose a novel approach called MDA-HS using privileged information (MDA-HS+) to effectively incorporate the valuable textual information into our MDA-HS method, based on the recent learning using privileged information paradigm. MDA-HS+ can be further extended by using a new elastic-net-like regularization. We solve our MDA-HS and MDA-HS+ methods by using the cutting-plane algorithm, in which a multiple kernel learning problem is derived and solved. Extensive experiments on three benchmark data sets demonstrate that our proposed approaches are effective for action and event recognition without requiring any labeled samples from the target domain.
Li Niu 0002, Xinxing Xu, Lin Chen 0021, Lixin Duan, Dong Xu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2015 Video Object Segmentation Via Dense Trajectories
abstract
In this paper, we propose a novel approach to segment moving object in video by utilizing improved point trajectories . First, point trajectories are densely sampled from video and tracked through optical flow, which provides information of long-term temporal interactions among objects in the video sequence . Second, a novel affinity measurement method considering both global and local information of point trajectories is proposed to cluster trajectories into groups. Finally, we propose a new graph-based segmentation method which adopts both local and global motion information encoded by the tracked dense point trajectories. The proposed approach achieves good performance on trajectory clustering, and it also obtains accurate video object segmentation results on both the Moseg dataset and our new dataset containing more challenging videos.
Lin Chen 0021, Jianbing Shen, Wenguan Wang, Bingbing Ni
IEEE Trans. Multim.1
2014 Recognizing RGB Images by Learning from RGB-D Data
abstract
In this work, we propose a new framework for recognizing RGB images captured by the conventional cameras by leveraging a set of labeled RGB-D data, in which the depth features can be additionally extracted from the depth images. We formulate this task as a new unsupervised domain adaptation (UDA) problem, in which we aim to take advantage of the additional depth features in the source domain and also cope with the data distribution mismatch between the source and target domains. To effectively utilize the additional depth features, we seek two optimal projection matrices to map the samples from both domains into a common space by preserving as much as possible the correlations between the visual features and depth features. To effectively employ the training samples from the source domain for learning the target classifier, we reduce the data distribution mismatch by minimizing the Maximum Mean Discrepancy (MMD) criterion, which compares the data distributions for each type of feature in the common space. Based on the above two motivations, we propose a new SVM based objective function to simultaneously learn the two projection matrices and the optimal target classifier in order to well separate the source samples from different classes when using each type of feature in the common space. An efficient alternating optimization algorithm is developed to solve our new objective function. Comprehensive experiments for object recognition and gender recognition demonstrate the effectiveness of our proposed approach for recognizing RGB images by learning from RGB-D data.
Lin Chen 0021, Wen Li 0001, Dong Xu 0001
CVPR1
2014 Incorporating Privileged Genetic Information for Fundus Image Based Glaucoma Detection
Lixin Duan, Yanwu Xu 0001, Wen Li 0001, Lin Chen 0021, Damon Wing Kee Wong, Tien Yin Wong, Jiang Liu 0001
MICCAI (2)4
2014 Re-texturing by intrinsic video
Jianbing Shen, Lin Chen 0021, Hanqiu Sun, Xuelong Li 0001
Inf. Sci.3
2014 Spectral Embedded Hashing for Scalable Image Retrieval
abstract
We propose a new graph based hashing method called spectral embedded hashing (SEH) for large-scale image retrieval. We first introduce a new regularizer into the objective function of the recent work spectral hashing to control the mismatch between the resultant hamming embedding and the low-dimensional data representation, which is obtained by using a linear regression function. This linear regression function can be employed to effectively handle the out-of-sample data, and the introduction of the new regularizer makes SEH better cope with the data sampled from a nonlinear manifold. Considering that SEH cannot efficiently cope with the high dimensional data, we further extend SEH to kernel SEH (KSEH) to improve the efficiency and effectiveness, in which a nonlinear regression function can also be employed to obtain the low dimensional data representation. We also develop a new method to efficiently solve the approximate solution for the eigenvalue decomposition problem in SEH and KSEH. Moreover, we show that some existing hashing methods are special cases of our KSEH. Our comprehensive experiments on CIFAR, Tiny-580K, NUS-WIDE, and Caltech-256 datasets clearly demonstrate the effectiveness of our methods.
Lin Chen 0021, Dong Xu 0001, Ivor W. Tsang, Xuelong Li 0001
IEEE Trans. Cybern.1
2013 Event Recognition in Videos by Learning from Heterogeneous Web Sources
abstract
In this work, we propose to leverage a large number of loosely labeled web videos (e.g., from YouTube) and web images (e.g., from Google/Bing image search) for visual event recognition in consumer videos without requiring any labeled consumer videos. We formulate this task as a new multi-domain adaptation problem with heterogeneous sources, in which the samples from different source domains can be represented by different types of features with different dimensions (e.g., the SIFT features from web images and space-time (ST) features from web videos) while the target domain samples have all types of features. To effectively cope with the heterogeneous sources where some source domains are more relevant to the target domain, we propose a new method called Multi-domain Adaptation with Heterogeneous Sources (MDA-HS) to learn an optimal target classifier, in which we simultaneously seek the optimal weights for different source domains with different types of features as well as infer the labels of unlabeled target domain data based on multiple types of features. We solve our optimization problem by using the cutting-plane algorithm based on group based multiple kernel learning. Comprehensive experiments on two datasets demonstrate the effectiveness of MDA-HS for event recognition in consumer videos.
Lin Chen 0021, Lixin Duan, Dong Xu 0001
CVPR1
2012 Efficient Discriminative Learning of Class Hierarchy for Many Class Prediction
Lin Chen 0021, Lixin Duan, Ivor W. Tsang, Dong Xu 0001
ACCV (1)1
2012 Tag-Based Image Retrieval Improved by Augmented Features and Group-Based Refinement
abstract
In this paper, we propose a new tag-based image retrieval framework to improve the retrieval performance of a group of related personal images captured by the same user within a short period of an event by leveraging millions of training web images and their associated rich textual descriptions. For any given query tag (e.g., “car”), the inverted file method is employed to automatically determine the relevant training web images that are associated with the query tag and the irrelevant training web images that are not associated with the query tag. Using these relevant and irrelevant web images as positive and negative training data respectively, we propose a new classification method called support vector machine (SVM) with augmented features (AFSVM) to learn an adapted classifier by leveraging the prelearned SVM classifiers of popular tags that are associated with a large number of relevant training web images. Treating the decision values of one group of test photos from AFSVM classifiers as the initial relevance scores, in the subsequent group-based refinement process, we propose to use the Laplacian regularized least squares method to further refine the relevance scores of test photos by utilizing the visual similarity of the images within the group. Based on the refined relevance scores, our proposed framework can be readily applied to tag-based image retrieval for a group of raw consumer photos without any textual descriptions or a group of Flickr photos with noisy tags. Moreover, we propose a new method to better calculate the relevance scores for Flickr photos. Extensive experiments on two datasets demonstrate the effectiveness of our framework.
Lin Chen 0021, Dong Xu 0001, Ivor W. Tsang, Jiebo Luo 0001
IEEE Trans. Multim.1
2012 Laplacian Embedded Regression for Scalable Manifold Regularization
abstract
Semi-supervised learning (SSL), as a powerful tool to learn from a limited number of labeled data and a large number of unlabeled data, has been attracting increasing attention in the machine learning community. In particular, the manifold regularization framework has laid solid theoretical foundations for a large family of SSL algorithms, such as Laplacian support vector machine (LapSVM) and Laplacian regularized least squares (LapRLS). However, most of these algorithms are limited to small scale problems due to the high computational cost of the matrix inversion operation involved in the optimization problem. In this paper, we propose a novel framework called Laplacian embedded regression by introducing an intermediate decision variable into the manifold regularization framework. By using ∈-insensitive loss, we obtain the Laplacian embedded support vector regression (LapESVR) algorithm, which inherits the sparse solution from SVR. Also, we derive Laplacian embedded RLS (LapERLS) corresponding to RLS under the proposed framework. Both LapESVR and LapERLS possess a simpler form of a transformed kernel, which is the summation of the original kernel and a graph kernel that captures the manifold structure. The benefits of the transformed kernel are two-fold: (1) we can deal with the original kernel matrix and the graph Laplacian matrix in the graph kernel separately and (2) if the graph Laplacian matrix is sparse, we only need to perform the inverse operation for a sparse matrix, which is much more efficient when compared with that for a dense one. Inspired by kernel principal component analysis, we further propose to project the introduced decision variable into a subspace spanned by a few eigenvectors of the graph Laplacian matrix in order to better reflect the data manifold, as well as accelerate the calculation of the graph kernel, allowing our methods to efficiently and effectively cope with large scale SSL problems. Extensive experiments on both toy and real world data sets show the effectiveness and scalability of the proposed framework.
Lin Chen 0021, Ivor W. Tsang, Dong Xu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2010 Tag-based web photo retrieval improved by batch mode re-tagging
abstract
Web photos in social media sharing websites such as Flickr are generally accompanied by rich but noisy textual descriptions (tags, captions, categories, etc.). In this paper, we proposed a tag-based photo retrieval framework to improve the retrieval performance for Flickr photos by employing a novel batch mode re-tagging method. The proposed batch mode re-tagging method can automatically refine noisy tags of a group of Flickr photos uploaded by the same user within a short period by leveraging millions of training web images and their associated rich textual descriptions. Specifically, for one group of Flickr photos, we construct a group-specific lexicon which contains only the tags of all photos within the group. For each query tag, we employ the inverted file method to automatically find loosely labeled training web images. We propose a SVM with Augmented Features, referred to as AFSVM, to learn adapted classifiers to refine the annotation tags of photos by leveraging the existing SVM classifiers of popular tags, which are associated with a large amount of positive training web images. Moreover, to further refine the annotation tags of photos in the same group, we additionally introduce an objective function that utilizes the visual similarities of photos within the group as well as the semantic proximities of their tags. Based on the refined tags, photos can be retrieved according to more reliable relevance scores. Extensive experiments demonstrate the effectiveness of our framework.
Lin Chen 0021, Dong Xu 0001, Ivor W. Tsang, Jiebo Luo 0001
CVPR1