EDBT 2026 Demo / reviewers in the wild / expert
Linjun Yang
dblp:65/5947
· DBLP profile ↗
70ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 6 first-authorArtificial intelligence and machine learning · 19 · 7 since 2021Databases, data management, data science and information retrieval · 14 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Web-scale Recommendations with LLMs: From Quality-aware Ranking to Candidate GenerationabstractExplore Further @ Bing is a webpage-to-webpage recommendation product, enhancing the search experience on Bing by surfacing engaging webpage recommendations tied to the search result URLs. In this paper, we present our approach for leveraging Large Language Models (LLMs) for enhancing our web-scale recommendation system. We describe the development and validation of our LLM-powered recommendation quality metric RecoDCG. We discuss our core techniques for utilizing LLMs to make our ranking stage quality-aware. Furthermore, we detail Q' recall, a recall path that enhances our system's candidate generation stage by leveraging LLMs to produce complementary and engaging recommendation candidates. We also address how we optimize our system for multiple objectives, balancing recommendation quality with click metrics. We deploy our work to production, achieving a significant improvement in recommendation quality. We share results from offline and online experiments as well as insights and steps we took to ensure our approaches scale effectively for our web-scale needs. Jaidev Shah, Iman Barjasteh, Amey Barapatre, Rana Forsati, Xue Deng, Blake Shepard, Ronak Shah, Linjun Yang |
KDD (1) | 11 |
| 2024 | Improving Text Embeddings with Large Language ModelsabstractIn this paper, we introduce a novel and simple method for obtaining high-quality text embeddings using only synthetic data and less than 1k training steps.Unlike existing methods that often depend on multi-stage intermediate pretraining with billions of weakly-supervised text pairs, followed by fine-tuning with a few labeled datasets, our method does not require building complex training pipelines or relying on manually collected datasets that are often constrained by task diversity and language coverage.We leverage proprietary LLMs to generate diverse synthetic data for hundreds of thousands of text embedding tasks across 93 languages.We then fine-tune open-source decoder-only LLMs on the synthetic data using standard contrastive loss.Experiments demonstrate that our method achieves strong performance on highly competitive text embedding benchmarks without using any labeled data.Furthermore, when fine-tuned with a mixture of synthetic and labeled data, our model sets new state-of-the-art results on the BEIR and MTEB benchmarks. Liang Wang 0046, Nan Yang 0002, Xiaolong Huang 0002, Linjun Yang, Rangan Majumder, Furu Wei |
ACL (1) | 4 |
| 2024 | AMPO: Automatic Multi-Branched Prompt OptimizationabstractSheng Yang, Yurong Wu, Yan Gao, Zineng Zhou, Bin Benjamin Zhu, Xiaodi Sun, Jian-Guang Lou, Zhiming Ding, Anbang Hu, Yuan Fang, Yunsong Li, Junyan Chen, Linjun Yang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yurong Wu, Yan Gao 0002, Zineng Zhou, Bin B. Zhu, Xiaodi Sun, Jian-Guang Lou, Zhiming Ding, Anbang Hu, Linjun Yang |
EMNLP | 13 |
| 2024 | LEAD: Liberal Feature-based Distillation for Dense RetrievalabstractKnowledge distillation is often used to transfer knowledge from a strong teacher model to a relatively weak student model. Traditional methods include response-based methods and feature-based methods. Response-based methods are widely used but suffer from lower upper limits of performance due to their ignorance of intermediate signals, while feature-based methods have constraints on vocabularies, tokenizers and model architectures. In this paper, we propose a liberal feature-based distillation method (LEAD). LEAD aligns the distribution between the intermediate layers of teacher model and student model, which is effective, extendable, portable and has no requirements on vocabularies, tokenizers, or model architectures. Extensive experiments show the effectiveness of LEAD on widely-used benchmarks, including MS MARCO Passage Ranking, TREC 2019 DL Track, MS MARCO Document Ranking and TREC 2020 DL Track. Our code is available in https://github.com/microsoft/SimXNS/tree/main/LEAD. Hao Sun 0015, Xiao Liu 0029, Yeyun Gong, Anlei Dong, Jingwen Lu, Yan Zhang 0117, Linjun Yang, Rangan Majumder, Nan Duan 0001 |
WSDM | 7 |
| 2023 | SimLM: Pre-training with Representation Bottleneck for Dense Passage RetrievalabstractLiang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liang Wang 0046, Nan Yang 0002, Xiaolong Huang 0002, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei |
ACL (1) | 5 |
| 2023 | LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval
Tao Shen 0001, Xiubo Geng, Chongyang Tao, Can Xu 0002, Xiaolong Huang 0002, Binxing Jiao, Linjun Yang, Daxin Jiang |
ICLR | 7 |
| 2022 | Less is Less: When are Snippets Insufficient for Human vs Machine Relevance Estimation?
Gabriella Kazai, Bhaskar Mitra 0001, Anlei Dong, Nick Craswell, Linjun Yang |
ECIR (2) | 5 |
| 2021 | xMoCo: Cross Momentum Contrastive Learning for Open-Domain Question AnsweringabstractNan Yang, Furu Wei, Binxing Jiao, Daxing Jiang, Linjun Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Nan Yang 0002, Furu Wei, Binxing Jiao, Daxing Jiang, Linjun Yang |
ACL/IJCNLP (1) | 5 |
| 2020 | Embedding-based Retrieval in Facebook SearchabstractSearch in social networks such as Facebook poses different challenges than in classical web search: besides the query text, it is important to take into account the searcher's context to provide relevant results. Their social graph is an integral part of this context and is a unique aspect of Facebook search. While embedding-based retrieval (EBR) has been applied in web search engines for years, Facebook search was still mainly based on a Boolean matching model. In this paper, we discuss the techniques for applying EBR to a Facebook Search system. We introduce the unified embedding framework developed to model semantic embeddings for personalized search, and the system to serve embedding-based retrieval in a typical search system based on an inverted index. We discuss various tricks and experiences on end-to-end optimization of the whole system, including ANN parameter tuning and full-stack optimization. Finally, we present our progress on two selected advanced topics about modeling. We evaluated EBR on verticals for Facebook Search with significant metrics gains observed in online A/B experiments. We believe this paper will provide useful insights and experiences to help people on developing embedding-based retrieval systems in search engines. Jui-Ting Huang, Shuying Sun, Philip Pronin, Janani Padmanabhan, Giuseppe Ottaviano, Linjun Yang |
KDD | 9 |
| 2018 | CleanNet: Transfer Learning for Scalable Image Classifier Training With Label NoiseabstractIn this paper, we study the problem of learning image classification models with label noise. Existing approaches depending on human supervision are generally not scalable as manually identifying correct or incorrect labels is time-consuming, whereas approaches not relying on human supervision are scalable but less effective. To reduce the amount of human supervision for label noise cleaning, we introduce CleanNet, a joint neural embedding network, which only requires a fraction of the classes being manually verified to provide the knowledge of label noise that can be transferred to other classes. We further integrate CleanNet and conventional convolutional neural network classifier into one framework for image classification learning. We demonstrate the effectiveness of the proposed algorithm on both of the label noise detection task and the image classification on noisy data task on several large-scale datasets. Experimental results show that CleanNet can reduce label noise detection error rate on held-out classes where no human supervision available by 41.5% compared to current weakly supervised methods. It also achieves 47% of the performance gain of verifying all images with only 3.2% images verified on an image classification task. Source code and dataset will be available at kuanghuei.github.io/CleanNetProject. Kuang-Huei Lee, Xiaodong He 0001, Lei Zhang 0001, Linjun Yang |
CVPR | 4 |
| 2018 | Web-Scale Responsive Visual Search at BingabstractIn this paper, we introduce a web-scale general visual search system deployed in Microsoft Bing. The system accommodates tens of billions of images in the index, with thousands of features for each image, and can respond in less than 200 ms. In order to overcome the challenges in relevance, latency, and scalability in such large scale of data, we employ a cascaded learning-to-rank framework based on various latest deep learning visual features, and deploy in a distributed heterogeneous computing platform. Quantitative and qualitative experiments show that our system is able to support various applications on Bing website and apps. Houdong Hu, Linjun Yang, Pavel Komlev, Xi Stephen Chen, Jiapei Huang, Meenaz Merchant, Arun Sacheti |
KDD | 3 |
| 2015 | Exploration of Image Search Results Quality AssessmentabstractImage retrieval plays an increasingly important role in our daily lives. There are many factors which affect the quality of image search results, including chosen search algorithms, ranking functions, and indexing features. Applying different settings for these factors generates search result lists with varying levels of quality. However, no setting can always perform optimally for all queries. Therefore, given a set of search result lists generated by different settings, it is crucial to automatically determine which result list is the best in order to present it to users. This paper aims to solve this problem and makes four main innovations. First, a preference learning model is proposed to quantitatively study and formulate the best image search result list identification problem. Second, a set of valuable preference learning related features is proposed by exploring the visual characters of returned images. Third, a query-dependent preference learning model is further designed for building a more precise and query-specific model. Fourth, the proposed approach has been tested on a variety of applications including re-ranking ability assessment, optimal search engine selection, and synonymous query suggestion. Extensive experimental results on three image search datasets demonstrate the effectiveness and promising potential of the proposed method. Xinmei Tian 0001, Yijuan Lu, Nate Stender, Linjun Yang, Dacheng Tao |
IEEE Trans. Big Data | 4 |
| 2015 | Image Search Reranking With Hierarchical Topic AwarenessabstractWith much attention from both academia and industrial communities, visual search reranking has recently been proposed to refine image search results obtained from text-based image search engines. Most of the traditional reranking methods cannot capture both relevance and diversity of the search results at the same time. Or they ignore the hierarchical topic structure of search result. Each topic is treated equally and independently. However, in real applications, images returned for certain queries are naturally in hierarchical organization, rather than simple parallel relation. In this paper, a new reranking method "topic-aware reranking (TARerank)" is proposed. TARerank describes the hierarchical topic structure of search results in one model, and seamlessly captures both relevance and diversity of the image search results simultaneously. Through a structured learning framework, relevance and diversity are modeled in TARerank by a set of carefully designed features, and then the model is learned from human-labeled training samples. The learned model is expected to predict reranking results with high relevance and diversity for testing queries. To verify the effectiveness of the proposed method, we collect an image search dataset and conduct comparison experiments on it. The experimental results demonstrate that the proposed TARerank outperforms the existing relevance-based and diversified reranking methods. Xinmei Tian 0001, Linjun Yang, Yijuan Lu, Qi Tian 0001, Dacheng Tao |
IEEE Trans. Cybern. | 2 |
| 2015 | Mining Latent Attributes From Click-Through Logs for Image RecognitionabstractAttribute-based image representation, which represents an image by projecting it into a space spanned by attributes, has attracted increasing attention from both computer vision and multimedia communities for its compactness and potential to bridge the semantic gap. While many works focus on learning attribute models and utilizing them in image recognition and retrieval, few touch on the problem of how to effectively construct a vocabulary of attributes, which is an essential part of effective attribute-based representation. Most existing approaches define the attribute vocabulary by human experts or through existing ontology, which is often limited in coverage of general concept space. In this paper, we propose automatically constructing the attribute vocabulary by mining latent topics from the click-through log of a commercial image search engine. These attributes are referred to as latent topic attributes (LTA), which take advantage of tens of millions of interactions between user submitted queries and images, thereby providing better coverage for the concept space than existing approaches. The mining of latent topics from the click log is formulated as a matrix factorization problem, and further improved by weighted terms-based matrix factorization to address the extreme sparsity of the click-through matrix. Both qualitative results of the mined LTA and quantitative results on the standard image recognition benchmark demonstrate the mined LTA's effectiveness. Yi-Jie Lu, Linjun Yang, Kuiyuan Yang, Yong Rui |
IEEE Trans. Multim. | 2 |
| 2014 | Predicting Failing Queries in Video SearchabstractThe ability to predict when a video search query is not likely to deliver satisfying search results is expected to enable more effective search results optimizations and improved search experience for users. In this paper, we propose a novel context-aware query failure prediction approach that predicts whether a particular query submitted in a user's search session is likely to fail. The approach builds on the well-known concept of query performance prediction introduced in conventional text-based Web search to estimate the query's retrieval performance, but extends this concept with two novel characteristics, user indicators and engine indicators. User indicators are derived from transaction logs, capture the patterns of user interactions with the video search engine, and exploit the context in which a particular query was submitted. Engine indicators are derived from the search results list and measure the consistency of visual search results at the level of visual concepts and textual metadata associated with videos. Extensive evaluation of the approach on a test set containing over one million video search queries shows its effectiveness and demonstrates a significant improvement over traditional and state-of-the-art baseline approaches. Christoph Kofler, Linjun Yang, Martha A. Larson, Tao Mei 0001, Alan Hanjalic, Shipeng Li 0001 |
IEEE Trans. Multim. | 2 |
| 2014 | Image Relevance Prediction Using Query-Context Bag-of-Object Retrieval ModelabstractImage search reranking and image research result summarization are two effective approaches which enhance text-based image search results using visual information. Since the existing approaches optimize search relevance in terms of average performance, they usually cannot achieve satisfactory results for some particular classes of queries, like “object queries,” which is defined as the queries with the intent of searching for some kinds of objects. One possible reason is that the generic approaches such as , , are mostly built based on the global statistics of images as features while ignoring the fact that the relevance between the image and the query sometimes depends on an image patch instead of the whole image. In this paper, we therefore design a novel bag-of-object retrieval model to predict image relevance, which is particularly effective for object queries. First, we construct an object vocabulary containing query-relative objects by mining frequent object patches from the result image collection of the expanded query set. After representing each image as a bag of objects, our retrieval model can be derived from a risk-minimization framework for language modeling. To demonstrate the effectiveness of the proposed model, this paper also present two related applications: for image search reranking, we adopt a supervised framework to combine multiple ranking features from different assumptions; for image search result summarization, we propose a two-step ranking process which optimizes not only representativeness but also image attractiveness. The experimental results show that the proposed methods can significantly outperform the existing approaches. Yang Yang 0222, Linjun Yang, Gangshan Wu, Shipeng Li 0001 |
IEEE Trans. Multim. | 2 |
| 2013 | Semantic-Spatial Matching for image classificationabstractSpatial Pyramid Matching (SPM) has been proven a simple but effective extension to bag-of-visual-words image representation for spatial layout information compensation. SPM describes image in coarse-to-fine scale by partitioning the image into blocks over multiple levels and the features extracted from each block are concatenated into a long vector representation. Based on the assumption that images from the same class have similar spatial configurations, SPM matches the blocks from different images according to their spatial layout, by aligning all blocks from an image in a fixed spatial order. However, target objects may appear at any location in the image with various backgrounds. Therefore, the fixed spatial matching in SPM fails to match similar objects located different locations. To solve this problem, we propose an effective and efficient block matching method, Semantic-Spatial Matching (SSM). In this method, not only the spatial layout but also the semantic content is considered for block matching. The experiments on two benchmark image classification datasets demonstrate the effectiveness of SSM. Yupeng Yan, Xinmei Tian 0001, Linjun Yang, Yijuan Lu, Houqiang Li |
ICME | 3 |
| 2013 | Clickage: towards bridging semantic and intent gaps via mining click logs of search enginesabstractThe semantic gap between low-level visual features and high-level semantics has been investigated for decades but still remains a big challenge in multimedia. When "search" became one of the most frequently used applications, "intent gap", the gap between query expressions and users' search intents, emerged. Researchers have been focusing on three approaches to bridge the semantic and intent gaps: 1) developing more representative features, 2) exploiting better learning approaches or statistical models to represent the semantics, and 3) collecting more training data with better quality. However, it remains a challenge to close the gaps. In this paper, we argue that the massive amount of click data from commercial search engines provides a data set that is unique in the bridging of the semantic and intent gap. Search engines generate millions of click data (a.k.a. image-query pairs), which provide almost "unlimited" yet strong connections between semantics and images, as well as connections between users' intents and queries. To study the intrinsic properties of click data and to investigate how to effectively leverage this huge amount of data to bridge semantic and intent gap is a promising direction to advance multimedia research. In the past, the primary obstacle is that there is no such dataset available to the public research community. This changes as Microsoft has released a new large-scale real-world image click data to public. This paper presents preliminary studies on the power of large-scale click data with a variety of experiments, such as building large-scale concept detectors, tag processing, search, definitive tag detection, intent analysis, etc., with the goal to inspire deeper researches based on this dataset. Xian-Sheng Hua 0001, Linjun Yang, Jingdong Wang 0001, Jing Wang 0068, Kuansan Wang, Yong Rui, Jin Li 0001 |
ACM Multimedia | 2 |
| 2013 | GPS Estimation from Users' Photos
Jing Li 0049, Xueming Qian, Yuan Yan Tang, Linjun Yang, Chaoteng Liu |
MMM (1) | 4 |
| 2013 | GPS Estimation for Places of Interest From Social Users' Uploaded PhotosabstractSocial media has become a very popular way for people to share their photos with friends. Because most of the social images are attached with GPS (geo-tags), a photo's GPS information can be estimated with the help of the large geo-tagged image set while using a visual searching based approach. This paper proposes an unsupervised image GPS location estimation approach with hierarchical global feature clustering and local feature refinement. It consists of two parts: an offline system and an online system. In the offline system, a hierarchical structure is constructed for a large-scale offline social image set with GPS information. Representative images are selected for each GPS location refined cluster, and an inverted file structure is proposed. In the online system, when given an input image, its GPS information can be estimated by hierarchical global clusters selection and local feature refinement in the online system. Both the computational cost and GPS estimation performance demonstrates the effectiveness of the proposed hierarchical structure and inverted file structure in our approach. Jing Li 0049, Xueming Qian, Yuan Yan Tang, Linjun Yang, Tao Mei 0001 |
IEEE Trans. Multim. | 4 |
| 2012 | Constrained keypoint quantization: towards better bag-of-words model for large-scale multimedia retrievalabstractBag-of-words models are among the most widely used and successful representations in multimedia retrieval. However, the quantization error which is introduced when mapping keypoints to visual words is one of the main drawbacks of the bag-of-words model. Although some techniques, such as soft-assignment to bags [23] and query expansion [27], have been introduced to deal with the problem, the performance gain is always at the cost of longer query response time, which makes them difficult to apply to large-scale multimedia retrieval applications. In this paper, we propose a simple "constrained keypoint quantization" method which can effectively reduce the overall quantization error of the bag-of-words representation and greatly improve the retrieval efficiency at the same time. The central idea of the proposed quantization method is that if a keypoint is far away from all visual words, we simply remove it. At first glance, this simple strategy seems naive and dangerous. However, we show that the proposed method has a solid theoretical background. Our experimental results on three widely used datasets for near duplicate image and video retrieval confirm that by removing a large amount of keypoints which have high quantization error, we obtain comparable or even better retrieval performance while dramatically boosting retrieval efficiency. Yang Cai 0002, Linjun Yang, Alex Hauptmann 0001 |
ICMR | 3 |
| 2012 | When video search goes wrong: predicting query failure using search engine logs and visual search resultsabstractThe recent increase in the volume and variety of video content available online presents growing challenges for video search. Users face increased difficulty in formulating effective queries and search engines must deploy highly effective algorithms to provide relevant results. Although lately much effort has been invested in optimizing video search engine results, relatively little attention has been given to predicting for which queries results optimization is most useful, i.e., predicting which queries will fail. Being able to predict when a video search query would fail is likely to make the video search result optimization more efficient and effective, improve the search experience for the user by providing support in the query formulation process and in this way boost the development of video search engines in general. While insight about a query's performance in general could be obtained using the well-known concept of query performance prediction (QPP), we propose a novel approach for predicting a failure of a video search query in the specific context of a search session. Our 'context-aware query failure' prediction approach uses a combination of 'user indicators' and 'engine indicators' to predict whether a particular query is likely to fail in the context of a particular search session. User indicators are derived from the search log and capture the patterns of query (re)formulation behavior and the click-through data of a user during a typical video search session. Engine indicators are derived from the video search results list and capture the visual variance of search results that would be offered to the user for the given query. We validate our approach experimentally on a test set containing 1+ million video search queries and show its effectiveness compared to a set of conventional QPP baselines. Our approach achieves a 13% relative improvement over the baseline. Christoph Kofler, Linjun Yang, Martha A. Larson, Tao Mei 0001, Alan Hanjalic, Shipeng Li 0001 |
ACM Multimedia | 2 |
| 2012 | A bag-of-objects retrieval model for web image searchabstractImage search reranking has been an active research topic in recent years to boost the performance of the existing web image search engine which is mostly based on textual metadata of images. Various approaches have been proposed to rerank images for general queries and argue that, they may not necessarily be optimal for queries in specific domain, e.g., object queries, since the reranking algorithms are operated on whole images, instead of the relevant parts of images. In this paper, we propose a novel bag-of-objects retrieval model for image search reranking of object queries. Firstly, we employ a common object discovery algorithm to discover query-relevant objects from the search results returned by text-based image search engine. Then, the query and its result images are represented as a language model on the query relevant object vocabulary, based on which the ranking function can be derived. As the common object discovery is unreliable and may introduce noises, we propose to incorporate the attributes of the discovered objects, e.g., size, position, etc., into the ranking function through a linear model, and the weights on the object attributes can be learned. The experiments on two subsets of Web Queries dataset comprising object queries demonstrate that our approach can significantly outperform the existing reranking methods on object queries. Yang Yang 0222, Linjun Yang, Gangshan Wu, Shipeng Li 0001 |
ACM Multimedia | 2 |
| 2012 | Query difficulty estimation for image retrieval
Yangxi Li, Bo Geng, Linjun Yang, Chao Xu 0006 |
Neurocomputing | 3 |
| 2012 | Ensemble Manifold RegularizationabstractWe propose an automatic approximation of the intrinsic manifold for general semi-supervised learning (SSL) problems. Unfortunately, it is not trivial to define an optimization function to obtain optimal hyperparameters. Usually, cross validation is applied, but it does not necessarily scale up. Other problems derive from the suboptimality incurred by discrete grid search and the overfitting. Therefore, we develop an ensemble manifold regularization (EMR) framework to approximate the intrinsic manifold by combining several initial guesses. Algorithmically, we designed EMR carefully so it 1) learns both the composite manifold and the semi-supervised learner jointly, 2) is fully automatic for learning the intrinsic manifold hyperparameters implicitly, 3) is conditionally optimal for intrinsic manifold approximation under a mild and reasonable assumption, and 4) is scalable for a large number of candidate manifold hyperparameters, from both time and space perspectives. Furthermore, we prove the convergence property of EMR to the deterministic matrix at rate root-n. Extensive experiments over both synthetic and real data sets demonstrate the effectiveness of the proposed framework. Bo Geng, Dacheng Tao, Chao Xu 0006, Linjun Yang, Xian-Sheng Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Ranking Model Adaptation for Domain-Specific SearchabstractWith the explosive emergence of vertical search domains, applying the broad-based ranking model directly to different domains is no longer desirable due to domain differences, while building a unique ranking model for each domain is both laborious for labeling data and time consuming for training models. In this paper, we address these difficulties by proposing a regularization-based algorithm called ranking adaptation SVM (RA-SVM), through which we can adapt an existing ranking model to a new domain, so that the amount of labeled data and the training cost is reduced while the performance is still guaranteed. Our algorithm only requires the prediction from the existing ranking models, rather than their internal representations or the data from auxiliary domains. In addition, we assume that documents similar in the domain-specific feature space should have consistent rankings, and add some constraints to control the margin and slack variables of RA-SVM adaptively. Finally, ranking adaptability measurement is proposed to quantitatively estimate if an existing ranking model can be adapted to a new domain. Experiments performed over Letor and two large scale data sets crawled from a commercial search engine demonstrate the applicabilities of the proposed ranking adaptation algorithms and the ranking adaptability measurement. Bo Geng, Linjun Yang, Chao Xu 0006, Xian-Sheng Hua 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Visually Summarizing Web Pages Through Internal and External ImagesabstractVisually summarizing web pages is an attractive approach that provides users an effective and friendly interface to identify desired contents at a first glance for search and re-finding tasks. Using dominant images in web pages is generally reliable for this purpose. However, dominant images are often unavailable in many web pages. To solve this problem, we first propose a new approach to summarize those web pages without any dominant images by retrieving relevant external images from the Internet. However, relevant external images are sometimes unreliable. To take the advantages of these two kinds of images, we further propose a clustering based algorithm to select the best summarization among all of internal and external images. This algorithm leverages relevance and dominance of images as the prior information. Experimental results show that our approach achieves 0.098 and 0.082 NDCG1gain on a human labeled data set, compared with relevant external image and dominant image, respectively. Our user study also indicates that the images selected by our algorithm are useful as the summarization of web pages. Binxing Jiao, Linjun Yang, Jizheng Xu, Qi Tian 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2012 | Difficulty Guided Image Retrieval Using Linear Multiple Feature EmbeddingabstractExisting image retrieval systems suffer from a performance variance for different queries. Severe performance variance may greatly degrade the effectiveness of the subsequent query-dependent ranking optimization algorithms, especially those that utilize the information mined from the initial search results. In this paper, we tackle this problem by proposing a query difficulty guided image retrieval system, which can predict the queries' ranking performance in terms of their difficulties and adaptively apply ranking optimization approaches. We estimate the query difficulty by comprehensively exploring the information residing in the query image, the retrieval results, and the target database. To handle the high-dimensional and multi-model image features in the large-scale image retrieval setting, we propose a linear multiple feature embedding algorithm which learns a linear transformation from a small set of data by integrating a joint subspace in which the neighborhood information is preserved. The transformation can be effectively and efficiently used to infer the subspace features of the newly observed data in the online setting. We prove the significance of query difficulty to image retrieval by applying it to guide the conduction of three retrieval refinement applications, i.e., reranking, federated search, and query suggestion. Thorough empirical studies on three datasets suggest the effectiveness and scalability of the proposed image query difficulty estimation algorithm, as well as the promising of the image difficulty guided retrieval system. Yangxi Li, Bo Geng, Dacheng Tao, Zhengjun Zha, Linjun Yang, Chao Xu 0006 |
IEEE Trans. Multim. | 5 |
| 2012 | Query Difficulty Prediction for Web Image SearchabstractImage search plays an important role in our daily life. Given a query, the image search engine is to retrieve images related to it. However, different queries have different search difficulty levels. For some queries, they are easy to be retrieved (the search engine can return very good search results). While for others, they are difficult (the search results are very unsatisfactory). Thus, it is desirable to identify those “difficult” queries in order to handle them properly. Query difficulty prediction (QDP) is an attempt to predict the quality of the search result for a query over a given collection. QDP problem has been investigated for many years in text document retrieval, and its importance has been recognized in the information retrieval (IR) community. However, little effort has been conducted on the image query difficulty prediction problem for image search. Compared with QDP in document retrieval, QDP in image search is more challenging due to the noise of textual features and the well-known semantic gap of visual features. This paper aims to investigate the QDP problem in Web image search. A novel method is proposed to automatically predict the quality of image search results for an arbitrary query. This model is built based on a set of valuable features that are designed by exploring the visual characteristic of images in the search results. The experiments on two real image search datasets demonstrate the effectiveness of the proposed query difficulty prediction method. Two applications, including optimal image search engine selection and search results merging, are presented to show the promising applicability of QDP. Xinmei Tian 0001, Yijuan Lu, Linjun Yang |
IEEE Trans. Multim. | 3 |
| 2012 | Correction to "Bayesian Visual Reranking"abstractIn the above titled paper (ibid., vol. 13, no. 4, pp. 639-652, Aug. 2011), the first author's name appears incorrectly in the byline as "Xinmie Tian" instead of "Xinmei Tian." The name appears correctly in the biography section. Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 2 |
| 2012 | Prototype-Based Image Search RerankingabstractThe existing methods for image search reranking suffer from the unreliability of the assumptions under which the initial text-based image search result is employed in the reranking process. In this paper, we propose a prototype-based reranking method to address this problem in a supervised, but scalable fashion. The typical assumption that the top-N images in the text-based search result are equally relevant is relaxed by linking the relevance of the images to their initial rank positions. Then, we employ a number of images from the initial search result as the prototypes that serve to visually represent the query and that are subsequently used to construct meta rerankers. By applying different meta rerankers to an image from the initial result, reranking scores are generated, which are then aggregated using a linear model to produce the final relevance score and the new rank position for an image in the reranked search result. Human supervision is introduced to learn the model weights offline, prior to the online reranking process. While model learning requires manual labeling of the results for a few queries, the resulting model is query independent and therefore applicable to any other query. The experimental results on a representative web image search dataset comprising 353 queries demonstrate that the proposed method outperforms the existing supervised and unsupervised reranking approaches. Moreover, it improves the performance over the text-based image search engine by more than 25.48%. Linjun Yang, Alan Hanjalic |
IEEE Trans. Multim. | 1 |
| 2012 | A unified context model for web image retrievalabstractContent-based web image retrieval based on the query-by-example (QBE) principle remains a challenging problem due to the semantic gap as well as the gap between a user's intent and the representativeness of a typical image query. In this article, we propose to address this problem by integrating query-related contextual information into an advanced query model to improve the performance of QBE-based web image retrieval. We consider both the local and global context of the query image. The local context can be inferred from the web pages and the click-through log associated with the query image, while the global context is derived from the entire corpus comprising all web images and the associated web pages. To effectively incorporate the local query context we propose a language modeling based approach to deal with the combined structured query representation from the contextual and visual information. The global query context is integrated by the multi-modal relevance model to “reconstruct” the query from the document models indexed in the corpus. In this way, the global query context is employed to address the noise or missing information in the query and its local context, so that a comprehensive and robust query model can be obtained. We evaluated the proposed approach on a representative product image dataset collected from the web and demonstrated that the inclusion of the local and global query contexts significantly improves the performance of QBE-based web image retrieval. Linjun Yang, Bo Geng, Alan Hanjalic, Xian-Sheng Hua 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2011 | Million-scale near-duplicate video retrieval systemabstractIn this paper, we present a novel near-duplicate video retrieval system serving one million web videos. To achieve both the effectiveness and efficiency, a visual word based approach is proposed, which quantizes each video frame into a word and represents the whole video as a bag of words. The system can respond to a query in 41ms with 78.4% MAP on average. Linjun Yang, Wei Ping, Tao Mei 0001, Xian-Sheng Hua 0001, Shipeng Li 0001 |
ACM Multimedia | 2 |
| 2011 | The role of attractiveness in web image searchabstractExisting web image search engines are mainly designed to optimize topical relevance. However, according to our user study, attractiveness is becoming a more and more important factor for web image search engines to satisfy users' search intentions. Important as it can be, web image attractiveness from the search users' perspective has not been sufficiently recognized in both the industry and the academia. In this paper, we present a definition of web image attractiveness with three levels according to the end users' feedback, including perceptual quality, aesthetic sensitivity and affective tune. Corresponding to each level of the definition, various visual features are investigated on their applicability to attractiveness estimation of web images. To further deal with the unreliability of visual features induced by the large variations of web images, we propose a contextual approach to integrate the visual features with contextual cues mined from image EXIF information and the associated web pages. We explore the role of attractiveness by applying it to various stages of a web image search engine, including the online ranking and the interactive reranking, as well as the offline index selection. Experimental results on three large-scale web image search datasets demonstrate that the incorporation of attractiveness can bring more satisfaction to 80% of the users for ranking/reranking search results and 30.5% index coverage improvement for index selection, compared to the conventional relevance based approaches. Bo Geng, Linjun Yang, Chao Xu 0006, Xian-Sheng Hua 0001, Shipeng Li 0001 |
ACM Multimedia | 2 |
| 2011 | Difficulty guided image retrieval using linear multiview embeddingabstractExisting image retrieval systems suffer from a radical performance variance for different queries. The bad initial search results for "difficult" queries may greatly degrade the performance of their subsequent refinements, especially the refinement that utilizes the information mined from the search results, e.g., pseudo relevance feedback based reranking. In this paper, we tackle this problem by proposing a query difficulty guided image retrieval system, which selectively performs reranking according to the estimated query difficulty. To improve the performance of both reranking and difficulty estimation, we apply multiview embedding (ME) to images represented by multiple different features for integrating a joint subspace by preserving the neighborhood information in each feature space. However, existing ME approaches suffer from both "out of sample" and huge computational cost problems, and cannot be applied to online reranking or offline large-scale data processing for practical image retrieval systems. Therefore, we propose a linear multiview embedding algorithm which learns a linear transformation from a small set of data and can effectively infer the subspace features of new data. Empirical evaluations on both Oxford and 500K ImageNet datasets suggest the effectiveness of the proposed difficulty guided retrieval system with LME. Yangxi Li, Bo Geng, Zhengjun Zha, Dacheng Tao, Linjun Yang, Chao Xu 0006 |
ACM Multimedia | 5 |
| 2011 | Learning to judge image search resultsabstractGiven the explosive growth of the Web and the popularity of image sharing Web sites, image retrieval plays an increasingly important role in our daily lives. Search engines aim to provide beneficial image search results to users in response to queries. The quality of image search results depends on many factors: chosen search algorithms, ranking functions, indexing features, the base image database, etc. Applying different settings for these factors generates search result lists with varying levels of quality. Previous research has shown that no setting can always perform optimally for all queries. Therefore, given a set of search result lists generated by different settings, it is crucial to automatically determine which result list is the best in order to present it to users. This paper proposes a novel method to automatically identify the best search result list from a number of candidates. There are three main innovations in this paper. First, we propose a preference learning model to quantitatively study the best image search result identification problem. Second, we propose a set of valuable preference learning related features by exploring the visual characters of returned images. Third, our method shows promising potential in applications such as reranking ability assessment and optimal search engine selection. Experiments on two image search datasets show that our method achieves about 80% prediction accuracy for reranking ability assessment, and selects optimal search engine for about 70% queries correctly. Xinmei Tian 0001, Yijuan Lu, Linjun Yang, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2011 | Video-based image retrievalabstractLikely variations in the capture conditions (e.g. light, blur, scale, occlusion) and in the viewpoint between the query image and the images in the collection are the factors due to which image retrieval based on the Query-by-Example (QBE) principle is still not reliable enough. In this paper, we propose a novel QBE-based image retrieval system where users are allowed to submit a short video clip as a query to improve the retrieval reliability. Improvement is achieved by integrating the information about different viewpoints and conditions under which object and scene appearances can be captured across different video frames. Rich information extracted from a video can be exploited to generate a more complete query representation than in the case of a single-image query and to improve the relevance of the retrieved results. Our experimental results show that video-based image retrieval (VBIR) is significantly more reliable than the retrieval using a single image as a query. Linjun Yang, Alan Hanjalic, Xian-Sheng Hua 0001, Shipeng Li 0001 |
ACM Multimedia | 1 |
| 2011 | Learning from search engine and human supervision for web image searchabstractVisual reranking aims at improving the precision of text-based Web image search. In this paper we propose to combine two learning strategies for deriving the reranking model: learning from search engine and learning from human supervision. The first strategy learns the reranking model in a pseudo-supervised fashion by interpreting parts of the initial text-based search result as pseudo-relevant. The second strategy involves manual relevance labeling of the text-based search results obtained for a limited number of representative queries. While learning from search engine is query dependent and can therefore adapt better to individual queries, it is essentially unsupervised and noisy. While human supervision can better relate the search results to true relevance criteria, it needs to be deployed in a way to keep the reranking scalable. A combination of the two is expected to benefit from their respective advantages and reduce the impact of their individual deficiencies. We propose a two-stage learning approach to visual reranking, where in the online stage multiple query-relative meta rerankers are learned in a pseudo-supervised fashion from the search results and in the offline stage human supervision is used to derive the final reranking function based on these meta rerankers. The experimental results demonstrate that the proposed method significantly outperforms the existing reranking approaches. Linjun Yang, Alan Hanjalic |
ACM Multimedia | 1 |
| 2011 | Latent visual context learning for web image applications
Wengang Zhou 0001, Qi Tian 0001, Yijuan Lu, Linjun Yang, Houqiang Li |
Pattern Recognit. | 4 |
| 2011 | Bayesian Visual RerankingabstractVisual reranking has been proven effective to refine text-based video and image search results. It utilizes visual information to recover “true” ranking list from the noisy one generated by text-based search, by incorporating both textual and visual information. In this paper, we model the textual and visual information from the probabilistic perspective and formulate visual reranking as an optimization problem in the Bayesian framework, termed Bayesian visual reranking. In this method, the textual information is modeled as a likelihood, to reflect the disagreement between reranked results and text-based search results which is called ranking distance. The visual information is modeled as a conditional prior, to indicate the ranking score consistency among visually similar samples which is called visual consistency. Bayesian visual reranking derives the best reranking results by maximizing visual consistency while minimizing ranking distance. To model the ranking distance more precisely, we propose a novel pair-wise method which measure the ranking distance based on the disagreement in terms of pair-wise orders. For visual consistency, we study three different regularizers to mine the best way for its modeling. We conduct extensive experiments on both video and image search datasets. Experimental results demonstrate the effectiveness of our proposed Bayesian visual reranking. Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 2 |
| 2011 | Object Retrieval Using Visual Query ContextabstractObject retrieval aims at retrieving images containing objects similar to the query object captured in the region of interest (ROI) of the query image. Boosted by the invention and wide popularity of SIFT image features and bag-of-visual-words image representation, object retrieval has progressed significantly in the past years and has already found deployment in real-life applications and products. While existing object retrieval methods perform well in many cases, they may fail to return satisfactory results if the ROI specified by the user is inaccurate or if the object captured there is too small to be represented using discriminative features and consequently to be matched with similar objects in the image collection. In order to improve the object retrieval performance also in these difficult cases, we propose in this paper an object retrieval method that exploits the information about the visual context of the query object and employ it to compensate for possible uncertainty in feature-based query object representation. Contextual information is drawn from the visual elements surrounding the query object in the query image. We consider the ROI as an uncertain observation of the latent search intent and the saliency map detected for the query image as a prior. Then a language modeling approach is employed to devise a contextual object retrieval (COR) model. There, the relevance score is determined based on the search intent scores that are inferred from the uncertain ROI and the saliency prior. The usefulness of the contextual information for object retrieval and the effectiveness of the proposed COR model are demonstrated and evaluated on three representative image datasets. Linjun Yang, Bo Geng, Alan Hanjalic, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 1 |
| 2010 | Content-aware Ranking for visual searchabstractThe ranking models of existing image/video search engines are generally based on associated text while the visual content is actually neglected. Imperfect search results frequently appear due to the mismatch between the textual features and the actual visual content. Visual reranking, in which visual information is applied to refine text based search results, has been proven to be effective. However, the improvement brought by visual reranking is limited, and the main reason is that the errors in the text-based results will propagate to the refinement stage. In this paper, we propose a Content-Aware Ranking model based on “learning to rank” framework, in which textual and visual information are simultaneously leveraged in the ranking learning process. We formulate the Content-Aware Ranking based on large margin structured output learning, by modeling the visual information into a regularization term. The direct optimization of the learning problem is nearly infeasible since the number of constraints is huge. The efficient cutting plane algorithm is adopted to learn the model by iteratively adding the most violated constraints. Extensive experimental results on a large-scale dataset collected from a commercial Web image search engine, as well as the TRECVID 2007 video search dataset, demonstrate that the proposed ranking model significantly outperforms the state-of-the-art ranking and reranking methods. Bo Geng, Linjun Yang, Chao Xu 0006, Xian-Sheng Hua 0001 |
CVPR | 2 |
| 2010 | Compact projection: Simple and efficient near neighbor search with practical memory requirementsabstractImage similarity search is a fundamental problem in computer vision. Efficient similarity search across large image databases depends critically on the availability of compact image representations and good data structures for indexing them. Numerous approaches to the problem of generating and indexing image codes have been presented in the literature, but existing schemes generally lack explicit estimates of the number of bits needed to effectively index a given large image database. We present a very simple algorithm for generating compact binary representations of imagery data, based on random projections. Our analysis gives the first explicit bound on the number of bits needed to effectively solve the indexing problem. When applied to real image search tasks, these theoretical improvements translate into practical performance gains: experimental results show that the new method, while using significantly less memory, is several times faster than existing alternatives. Kerui Min, Linjun Yang, John Wright 0001, Lei Wu 0017, Xian-Sheng Hua 0001, Yi Ma 0001 |
CVPR | 2 |
| 2010 | Large-scale robust visual codebook constructionabstractThe web-scale image retrieval system demands a large-scale visual codebook, which is difficult to be generated by the commonly adopted K-means vector quantization due to the applicability issue. While approximate K-means is proposed to scale up the visual codebook construction it needs to employ a high-precision approximate nearest neighbor search in the assignment step and is difficult to converge, which limits its scalability. In this paper, we propose an improved approximate K-means, by leveraging the assignment information in the history, namely the previous iterations, to improve the assignment precision. By further randomizing the employed approximate nearest neighbor search in each iteration, the proposed algorithm can improve the assignment precision conceptually similarly as the randomized k-d trees, while nearly no additional cost is introduced. The algorithm can be proved to be convergent and we demonstrate that the proposed algorithm improves the quality of the generated visual codebook as well as the scalability experimentally and analytically. Darui Li, Linjun Yang, Xian-Sheng Hua 0001, HongJiang Zhang |
ACM Multimedia | 2 |
| 2010 | Real-time large scale near-duplicate web video retrievalabstractNear-duplicate video retrieval is becoming more and more important with the exponential growth of the Web. Though various approaches have been proposed to address this problem, they are mainly focusing on the retrieval accuracy while infeasible to query on Web scale video database in real time. This paper proposes a novel method to address the efficiency and scalability issues for near-duplicate We video retrieval. We introduce a compact spatiotemporal feature to represent videos and construct an efficient data structure to index the feature to achieve real-time retrieving performance. This novel feature leverages relative gray-level intensity distribution within a frame and temporal structure of videos along frame sequence. The new index structure is proposed based on inverted file to allow for fast histogram intersection computation between videos. To demonstrate the effectiveness and efficiency of the proposed methods we evaluate its performance on an open Web video data set containing about 10K videos and compare it with four existing methods in terms of precision and time complexity. We also test our method on a data set containing about 50K videos and 11M key-frames. It takes on average 17ms to execute a query against the whole 50K Web video data set. Lifeng Shang, Linjun Yang, Kwok-Ping Chan, Xian-Sheng Hua 0001 |
ACM Multimedia | 2 |
| 2010 | Supervised reranking for web image searchabstractVisual search reranking that aims to improve the text-based image search with the help from visual content analysis has rapidly grown into a hot research topic. The interestingness of the topic stems mainly from the fact that the search reranking is an unsupervised process and therefore has the potential to scale better than its main alternative, namely the search based on offline-learned semantic concepts. However, the unsupervised nature of the reranking paradigm also makes it suffer from problems, the main of which can be identified as the difficulty to optimally determine the role of visual modality over different application scenarios. Inspired by the success of the "learning-to-rank" idea proposed in the field of information retrieval, we propose in this paper the "learning-to-rerank" paradigm, which derives the reranking function in a supervised fashion from the human-labeled training data. Although supervised learning is introduced, our approach does not suffer from scalability issues since a unified reranking model is learned that can be applied to all queries. In other words, a query-independent reranking model will be learned for all queries using query-dependent reranking features. The query-dependent reranking feature extraction is challenging since the textual query and the visual documents have different representation. In this paper, 11 lightweight reranking features are proposed by representing the textual query using visual context and pseudo relevant images from the initial search result. The experiments performed on two representative Web image datasets demonstrate that the proposed learning-to-rerank algorithm outperforms the state-of-the-art unsupervised reranking methods, which makes the learning-to-rerank paradigm a promising alternative for robust and reliable Web-scale image search. Linjun Yang, Alan Hanjalic |
ACM Multimedia | 1 |
| 2010 | Visual Reranking with Local Learning Consistency
Xinmei Tian 0001, Linjun Yang, Xiuqing Wu, Xian-Sheng Hua 0001 |
MMM | 2 |
| 2010 | Visual summarization of web pagesabstractVisual summarization is an attractive new scheme to summarize web pages, which can help achieve a more friendly user experience in search and re-finding tasks by allowing users quickly get the idea of what the web page is about and helping users recall the visited web page. In this paper, we perform a careful study on the recently proposed visual summarization approaches, including the thumbnail of the web page snapshot, the internal image in the web page which is representative of the content in the page, and the visual snippet which is a synthesized image based on the internal image, the title, and the logo found in the web page. Moreover, since the internal image based summarization approach hardly works when the representative internal images are unavailable, we propose a new strategy, which retrieves the representative image from the external to summarize the web page. The experimental results suggest that the various summarization approaches have respective advantages on different types of web pages. While internal images and thumbnails can provide a reliable summarization on web pages with dominant images and web pages with simple structure respectively, the external images are regarded as a useful information to complement the internal images and are demonstrated very useful in helping users understanding new web pages. The visual snippet performs well on the re-finding tasks since it incorporates the title and logo which are advantageous on identifying the visited web pages. Binxing Jiao, Linjun Yang, Jizheng Xu, Feng Wu 0001 |
SIGIR | 2 |
| 2010 | Visual query suggestion: Towards capturing user intent in internet image searchabstractQuery suggestion is an effective approach to bridge the Intention Gap between the users' search intents and queries. Most existing search engines are able to automatically suggest a list of textual query terms based on users' current query input, which can be called Textual Query Suggestion. This article proposes a new query suggestion scheme named Visual Query Suggestion (VQS) which is dedicated to image search. VQS provides a more effective query interface to help users to precisely express their search intents by joint text and image suggestions. When a user submits a textual query, VQS first provides a list of suggestions, each containing a keyword and a collection of representative images in a dropdown menu. Once the user selects one of the suggestions, the corresponding keyword will be added to complement the initial query as the new textual query, while the image collection will be used as the visual query to further represent the search intent. VQS then performs image search based on the new textual query using text search techniques, as well as content-based visual retrieval to refine the search results by using the corresponding images as query examples. We compare VQS against three popular image search engines, and show that VQS outperforms these engines in terms of both the quality of query suggestion and the search performance. Zhengjun Zha, Linjun Yang, Tao Mei 0001, Meng Wang 0001, Zengfu Wang, Tat-Seng Chua, Xian-Sheng Hua 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2009 | Image Search Result Summarization with Informative Priors
Linjun Yang, Xian-Sheng Hua 0001 |
ACCV (3) | 2 |
| 2009 | Ranking model adaptation for domain-specific searchabstractRecently, various domain-specific search engines emerge, which are restricted to specific topicalities or document formats, and vertical to the broad-based search. Simply applying the ranking model trained for the broad-based search to the verticals cannot achieve a sound performance due to the domain differences, while building different ranking models for each domain is both laborious for labeling sufficient training samples and time-consuming or the training process. In this paper, to address the above difficulties, we investigate two problems: (1) whether we can adapt the ranking model learned for existing Web page search or verticals, to the new domain, so that the amount of labeled data and the training cost is reduced, while the performance requirement is still satisfied; and (2) how to adapt the ranking model from auxiliary domains to a new target domain. We address the second problem from the regularization framework and an algorithm called ranking adaptation SVM is proposed. Our algorithm is flexible enough, which needs only the prediction from the existing ranking model, rather than the internal representation of the model or the data from auxiliary domains. The first problem is addressed by the proposed ranking adaptability measurement, which quantitatively estimates if an existing ranking model can be adapted to the new domain. Extensive experiments are performed over Letor benchmark dataset and two large scale datasets crawled from different domains through a commercial internet search engine, where the ranking model learned for one domain will be adapted to the other. The results demonstrate the applicabilities of the proposed ranking model adaptation algorithm and the ranking adaptability measurement. Bo Geng, Linjun Yang, Chao Xu 0006, Xian-Sheng Hua 0001 |
CIKM | 2 |
| 2009 | Ensemble manifold regularizationabstractWe propose an automatic approximation of the intrinsic manifold for general semi-supervised learning problems. Unfortunately, it is not trivial to define an optimization function to obtain optimal hyperparameters. Usually, pure cross-validation is considered but it does not necessarily scale up. A second problem derives from the suboptimality incurred by discrete grid search and overfitting problems. As a consequence, we developed an ensemble manifold regularization (EMR) framework to approximate the intrinsic manifold by combining several initial guesses. Algorithmically, we designed EMR very carefully so that it (a) learns both the composite manifold and the semi-supervised classifier jointly; (b) is fully automatic for learning the intrinsic manifold hyperparameters implicitly; (c) is conditionally optimal for intrinsic manifold approximation under a mild and reasonable assumption; and (d) is scalable for a large number of candidate manifold hyperparameters, from both time and space perspectives. Extensive experiments over both synthetic and real datasets show the effectiveness of the proposed framework. Bo Geng, Chao Xu 0006, Dacheng Tao, Linjun Yang, Xian-Sheng Hua 0001 |
CVPR | 4 |
| 2009 | Tag quality improvement for social imagesabstractOnline social media sharing Web sites like Flickr allow users to manually annotate images with tags, which can facilitate image search and organization. However, the tags provided by users are often imprecise and incomplete, which severely limits the application of tags to image search and browse. In this paper, we propose a scheme to improve poorly annotated tags associated with social images. Two properties are exploited and integrated in an unified optimization framework: (1) consistency between visual and semantic similarities, where the semantic similarity is estimated using tags; (2) compatibility of tags before and after improvement, since the initial user provided tags carry valuable information. An iterative bound optimization method is derived to solve the optimization problem. Experimental results on Flickr dataset show that the proposed method can significantly improve the quality of tags. Dong Liu 0001, Meng Wang 0001, Linjun Yang, Xian-Sheng Hua 0001, HongJiang Zhang |
ICME | 3 |
| 2009 | Efficient image and video re-coloring for colorblindnessabstractThere are about 8% of men and 0.8% of women suffering from colorblindness. These viewers have difficulty in discriminating several colors, and thus many colorful images and videos that have high qualities for normal viewers may not be readily perceptible for them. In this paper, we propose an efficient re-coloring approach which can modify the colors of images and videos to enhance their perceptibility for colorblind users. Given an image, the re-coloring is accomplished by two color rotation steps in CIELAB color space. We first perform a local color rotation such that information of b*axis can be enhanced, and then adopt a global color rotation to refine the results. We will show that this method is simple yet effective, and it is able to outperform the traditional re-coloring algorithms in both performance and computational efficiency. We apply the method to process video frames and adopt several strategies to further reduce computational cost to realize real-time video re-coloring. We also explore the structure information of video data to avoid the color inconsistency problem. Specifically, we enforce the color mapping function to be identical in each shot and vary smoothly in adjacent shots. We conduct experiments on real-world images and videos with diverse content, and empirical results demonstrate the effectiveness of the proposed methods. Bo Liu 0005, Meng Wang 0001, Linjun Yang, Xiuqing Wu, Xian-Sheng Hua 0001 |
ICME | 3 |
| 2009 | Query aware visual similarity propagation for image search rerankingabstractImage search reranking is an effective approach to refining the text-based image search result. In the reranking process, the estimation of visual similarity is critical to the performance. However, the existing measures, based on global or local features, cannot be adapted to different queries. In this paper, we propose to estimate a query aware image similarity by incorporating the global visual similarity, local visual similarity and visual word co-occurrence into an iterative propagation framework. After the propagation, a query aware image similarity combining the advantages of both global and local similarities is achieved and applied to image search reranking. The experiments on a real-world Web image dataset demonstrate that the proposed query aware similarity outperforms the global, local similarity and their linear combination, for image search reranking. Linjun Yang, Xinmei Tian 0001 |
ACM Multimedia | 2 |
| 2009 | Visual query suggestionabstractQuery suggestion is an effective approach to improve the usability of image search. Most existing search engines are able to automatically suggest a list of textual query terms based on users' current query input, which can be called Textual Query Suggestion. This paper proposes a new query suggestion scheme named Visual Query Suggestion (VQS) which is dedicated to image search. It provides a more effective query interface to formulate an intent-specific query by joint text and image suggestions. We show that VQS is able to more precisely and more quickly help users specify and deliver their search intents. When a user submits a text query, VQS first provides a list of suggestions, each containing a keyword and a collection of representative images in a dropdown menu. If the user selects one of the suggestions, the corresponding keyword will be added to complement the initial text query as the new text query, while the image collection will be formulated as the visual query. VQS then performs image search based on the new text query using text search techniques, as well as content-based visual retrieval to refine the search results by using the corresponding images as query examples. We compare VQS with three popular image search engines, and show that VQS outperforms these engines in terms of both the quality of query suggestion and search performance. Zhengjun Zha, Linjun Yang, Tao Mei 0001, Meng Wang 0001, Zengfu Wang |
ACM Multimedia | 2 |
| 2009 | Multiple-Instance Active Learning for Image Categorization
Dong Liu 0001, Xian-Sheng Hua 0001, Linjun Yang, HongJiang Zhang |
MMM | 3 |
| 2009 | Accommodating colorblind users in image searchabstractThere are about 8% of men and 0.8% of women suffering from colorblindness. Due to certain loss of color information, the existing image search techniques may not provide satisfactory results for these users. In this demonstration, we show an image search system that can accommodate colorblind users. It can help these special users find and enjoy what they want by providing multiple services for them, including search results reranking, image recoloring and color indication. Meng Wang 0001, Bo Liu 0005, Linjun Yang, Xian-Sheng Hua 0001 |
SIGIR | 3 |
| 2009 | Query sampling for ranking learning in web searchabstractLearning to rank has become a popular approach to build a ranking model for Web search recently. Based on our observation, the constitution of the training set will greatly influence the performance of the learned ranking model. Meanwhile, the number of queries in Web search is nearly infinite and the human labeling cost is expensive, hence a subset of queries need to be carefully selected for training. In this paper, we develop a greedy algorithm to sample the queries, by simultaneously taking the query density, difficulty and diversity into consideration. The experimental results on a collected Web search dataset comprising 2024 queries show that the proposed method can lead to a more informative training set for building an effective model. Linjun Yang, Bo Geng, Xian-Sheng Hua 0001 |
SIGIR | 1 |
| 2009 | Tag rankingabstractSocial media sharing web sites like Flickr allow users to annotate images with free tags, which significantly facilitate Web image search and organization. However, the tags associated with an image generally are in a random order without any importance or relevance information, which limits the effectiveness of these tags in search and other applications. In this paper, we propose a tag ranking scheme, aiming to automatically rank the tags associated with a given image according to their relevance to the image content. We first estimate initial relevance scores for the tags based on probability density estimation, and then perform a random walk over a tag similarity graph to refine the relevance scores. Experimental results on a 50, 000 Flickr photo collection Dong Liu 0001, Xian-Sheng Hua 0001, Linjun Yang, Meng Wang 0001, HongJiang Zhang |
WWW | 3 |
| 2009 | Learning to tagabstractSocial tagging provides valuable and crucial information for large-scale web image retrieval. It is ontology-free and easy to obtain; however, irrelevant tags frequently appear, and users typically will not tag all semantic objects in the image, which is also called semantic loss. To avoid noises and compensate for the semantic loss, tag recommendation is proposed in literature. However, current recommendation simply ranks the related tags based on the single modality of tag co-occurrence on the whole dataset, which ignores other modalities, such as visual correlation. This paper proposes a multi-modality recommendation based on both tag and visual correlation, and formulates the tag recommendation as a learning problem. Each modality is used to generate a ranking feature, and Rankboost algorithm is applied to learn an optimal combination of these ranking features from different modalities. Experiments on Flickr data demonstrate the effectiveness of this learning-based multi-modality recommendation strategy. Lei Wu 0017, Linjun Yang, Nenghai Yu, Xian-Sheng Hua 0001 |
WWW | 2 |
| 2008 | Unbiased active learning for image retrievalabstractIn transductive active learning, after selecting the samples for labeling using existing sample selection strategy such as close-to-boundary, the constructed labeled set will be under a different distribution from the unlabeled set, which violates the i.i.d assumption of existing classifier. In this paper, by explicitly considering the distribution difference, we propose an algorithm called unbiased active learning. In such algorithm, the distribution difference, so-called sample selection bias, is not only considered into the classifier, but also incorporated into the sample selection process for introducing a better sample selection strategy. We apply the proposed method to image retrieval and the experimental results show that our unbiased active learning algorithm outperforms existing approaches. Bo Geng, Linjun Yang, Zhengjun Zha, Chao Xu 0006, Xian-Sheng Hua 0001 |
ICME | 2 |
| 2008 | Transductive video annotation via local learnable kernel classifierabstractOne crucial problem in transductive video annotation is how to estimate the label from the neighboring samples. Existing methods such as graph-based Gaussian random filed only considered the pair-wise similarity and then propagated the labels based on it. In this paper, we propose a new method from the perspective of local learning, which formulate the prediction of labels from the neighbors into a learning problem. Our contributions lie in two-fold: (1) we propose a new transductive video annotation method based on local kernel classifier; (2) local learnable is proposed to measure whether a sample can be learned from the neighbors well and we employ this measure into the optimization objective. Experiments on TRECVID 2005 dataset prove that the proposed method is effective and the local learning perspective is promising for video annotation. Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001 |
ICME | 2 |
| 2008 | Optimized video scene segmentationabstractIn this paper, we propose an optimized video scene segmentation approach with considering both content coherence and temporally contextual dissimilarity. First, a chain structure is constructed by connecting temporally adjacent shots to represent a video. Then the chain is partitioned such that the content within a chain segment is coherent enough and the contextual similarity of temporally adjacent chain segments is small enough. This task is formulated as a ratio function of content coherence and contextual similarity. Finally, we present an effective and efficient hierarchical chain partitioning approach to find the optimal scene segmentation. Experimental results on a set of home videos and feature movies demonstrate the superiority of the proposed approach over several existing key approaches. Jingdong Wang 0001, Xinmei Tian 0001, Linjun Yang, Zhengjun Zha, Xian-Sheng Hua 0001 |
ICME | 3 |
| 2008 | Bayesian video search rerankingabstractContent-based video search reranking can be regarded as a process that uses visual content to recover the "true" ranking list from the noisy one generated based on textual information. This paper explicitly formulates this problem in the Bayesian framework, i.e., maximizing the ranking score consistency among visually similar video shots while minimizing the ranking distance, which represents the disagreement between the objective ranking list and the initial text-based. Different from existing point-wise ranking distance measures, which compute the distance in terms of the individual scores, two new methods are proposed in this paper to measure the ranking distance based on the disagreement in terms of pair-wise orders. Specifically, hinge distance penalizes the pairs with reversed order according to the degree of the reverse, while preference strength distance further considers the preference degree. By incorporating the proposed distances into the optimization objective, two reranking methods are developed which are solved using quadratic programming and matrix computation respectively. Evaluation on TRECVID video search benchmark shows that the performance improvement up to 21% on TRECVID 2006 and 61.11% on TRECVID 2007 are achieved relative to text search baseline. Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Yichen Yang 0001, Xiuqing Wu, Xian-Sheng Hua 0001 |
ACM Multimedia | 2 |
| 2008 | A comprehensive human computation framework: with application to image labelingabstractImage and video labeling is important for computers to understand images and videos and for image and video search. Manual labeling is tedious and costly. Automatically image and video labeling is yet a dream. In this paper, we adopt a Web 2.0 approach to labeling images and videos efficiently: Internet users around the world are mobilized to apply their "common sense" to solve problems that are hard for today's computers, such as labeling images and videos. We first propose a general human computation framework that binds problem providers, Web sites, and Internet users together to solve large-scale common sense problems efficiently and economically. The framework addresses the technical challenges such as preventing a malicious party from attacking others, removing answers from bots, and distilling human answers to produce high-quality solutions to the problems. The framework is then applied to labeling images. Three incremental refinement stages are applied. The first stage collects candidate labels of objects in an image. The second stage refines the candidate labels using multiple choices. Synonymic labels are also correlated in this stage. To prevent bots and lazy humans from selecting all the choices, trap labels are generated automatically and intermixed with the candidate labels. Semantic distance is used to ensure that the selected trap labels would be different enough from the candidate labels so that no human users would mistakenly select the trap labels. The last stage is to ask users to locate an object given a label from a segmented image. The experimental results are also reported in this paper. They indicate that our proposed schemes can successfully remove spurious answers from bots and distill human answers to produce high-quality image labels. Yang Yang 0059, Bin B. Zhu, Linjun Yang, Shipeng Li 0001, Nenghai Yu |
ACM Multimedia | 4 |
| 2007 | VideoSense: towards effective online video advertisingabstractWith Internet delivery of video content surging to an unprecedented level, online video advertising is becoming increasingly pervasive. In this paper, we present a novel advertising system for online video service called VideoSense, which automatically associates the most relevant video ads with online videos and seamlessly inserts the ads at the most appropriate positions within each individual video. Unlike most current video-oriented sites that only display a video ad at the beginning or the end of a video, VideoSense aims to embed more contextually relevant ads at less intrusive positions within the video stream. Given an online video, VideoSense is able to detect a set of candidate ad insertion points based on content discontinuity and attractiveness, select a list of relevant candidate ads ranked according to global textual relevance, and compute local visual-aural relevance between each pair of insertion points and ads. To support contextually relevant and less intrusive advertising, the ads are expected to be inserted at the positions with highest discontinuity and lowest attractiveness, while the overall global and local relevance is maximized. We formulate this task as a nonlinear 0-1 integer programming problem and embed these rules as constraints. The experiments have proved the effectiveness of VideoSense for online video advertising. Tao Mei 0001, Xian-Sheng Hua 0001, Linjun Yang, Shipeng Li 0001 |
ACM Multimedia | 3 |
| 2007 | VideoSense: a contextual video advertising systemabstractThis demonstration presents a novel contextual advertising platform for online video service, called VideoSense. Unlike most current video-oriented sites that only display a video ad at the beginning or the end of a video, VideoSense aims to embed more contextually relevant ads at less intrusive positions within the video stream. Given an online video, VideoSense is able to detect a set of candidate ad insertion points based on content analysis, select a list of relevant candidate ads ranked according to textual relevance, and find the best match between insertion points and ads which maximizes the overall multimodal relevance. The effectiveness of VideoSense supporting contextually relevant and less intrusive advertising is validated by the user studies conducted on a variety of online video documents. Tao Mei 0001, Linjun Yang, Xian-Sheng Hua 0001, Shipeng Li 0001 |
ACM Multimedia | 2 |
| 2007 | VideoReach: an online video recommendation systemabstractThis paper presents a novel online video recommendation system called VideoReach, which alleviates users' efforts on finding the most relevant videos according to current viewings without a sufficient collection of user profiles as required in traditional recommenders. In this system, video recommendation is formulated as finding a list of relevant videos in terms of multimodal relevance (i.e. textual, visual, and aural relevance) and user click-through. Since different videos have different intra-weights of relevance within an individual modality and inter-weights among different modalities, we adopt relevance feedback to automatically find optimal weights by user click-though, as well as an attention fusion function to fuse multimodal relevance. We use 20 clips as the representative test videos, which are searched by top 10 queries from more than 13k online videos, and report superior performance compared with an existing video site. Tao Mei 0001, Bo Yang 0008, Xian-Sheng Hua 0001, Linjun Yang, Shiqiang Yang, Shipeng Li 0001 |
SIGIR | 4 |
| 2005 | Efficient Video Clip Retrieval Using Index StructureabstractRetrieving similar video clips from large video database requires high query efficiency, precision and recall, which remains a challenging problem since the traditional query algorithms are inefficient and time-consuming. In this paper, we adopt the high-dimensional index structure vector-approximation file (VA-file) to organize the video database, and propose a new similarity measure which takes the temporal order among the video representations into account to improve the accuracy of query. Based on the VA-file and similarity measure, a new video clip retrieval algorithm is proposed in our method to achieve high query efficiency by using restricted sliding window to construct candidate video clips. Experimental results show that the proposed video retrieval method is efficient and effective Linjun Yang, Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
MMSP | 2 |