EDBT 2026 Demo / reviewers in the wild / expert
Bin Bi
dblp:28/3252
· DBLP profile ↗
28ranked-venue papers
11as first author
12since 2021 · last 2023
0000-0003-2207-9146ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 8 first-author · 9 since 2021Databases, data management, data science and information retrieval · 13 · 8 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | BUS : Efficient and Effective Vision-language Pre-training with Bottom-Up Patch SummarizationabstractVision Transformer (ViT) based Vision-Language Pre-training (VLP) models have demonstrated impressive performance in various tasks. However, the lengthy visual token sequences fed into ViT can lead to training inefficiency and ineffectiveness. Existing efforts address the challenge by either bottom-level patch extraction in the ViT backbone or top-level patch abstraction outside, not balancing training efficiency and effectiveness well. Inspired by text summarization in natural language processing, we propose a Bottom-Up Patch Summarization approach named BUS, coordinating bottom-level extraction and top-level abstraction to learn a concise summary of lengthy visual token sequences efficiently. Specifically, We incorporate a Text-Semantics-Aware Patch Selector (TSPS) into the ViT backbone to perform a coarse-grained visual token extraction and then attach a flexible Transformer-based Patch Abstraction Decoder (PAD) upon the backbone for top-level visual abstraction. This bottom-up collaboration enables our BUS to yield high training efficiency while maintaining or even improving effectiveness. We evaluate our approach on various visual-language understanding and generation tasks and show competitive downstream task performance while boosting the training efficiency by 50%. Additionally, our model achieves state-of-the-art performance on many downstream tasks by increasing input image resolution without increasing computational costs over baselines. Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Bin Bi, Shikun Zhang, Fei Huang 0002, Songfang Huang |
ICCV | 7 |
| 2023 | mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and VideoabstractRecent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for multi-modal pretraining, which can benefit from modality collaboration while addressing the problem of modality entanglement. In contrast to predominant paradigms of solely relying on sequence-to-sequence generation or encoder-based instance discrimination, mPLUG-2 introduces a multi-module composition network by sharing common universal modules for modality collaboration and disentangling different modality modules to deal with modality entanglement. It is flexible to select different modules for different understanding and generation tasks across all modalities including text, image, and video. Empirical study shows that mPLUG-2 achieves state-of-the-art or competitive results on a broad range of over 30 downstream tasks, spanning multi-modal tasks of image-text and video-text understanding and generation, and uni-modal tasks of text-only, image-only, and video-only understanding. Notably, mPLUG-2 shows new state-of-the-art results of 48.0 top-1 accuracy and 80.3 CIDEr on the challenging MSRVTT video QA and video caption tasks with a far smaller model size and data scale. It also demonstrates strong zero-shot transferability on vision-language and video-language tasks. Code and models will be released in https://github.com/X-PLUG/mPLUG-2. Haiyang Xu 0001, Qinghao Ye, Ming Yan 0008, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li 0003, Bin Bi, Qi Qian 0001, Wei Wang 0225, Guohai Xu, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Jingren Zhou 0001 |
ICML | 8 |
| 2023 | LightToken: A Task and Model-agnostic Lightweight Token Embedding Framework for Pre-trained Language Models
Haoyu Wang 0004, Ruirui Li 0002, Haoming Jiang, Xianfeng Tang, Bin Bi, Monica Xiao Cheng, Yaqing Wang 0001, Tuo Zhao, Jing Gao 0004 |
KDD | 6 |
| 2023 | COPA : Efficient Vision-Language Pre-training through Collaborative Object- and Patch-Text AlignmentabstractVision-Language Pre-training (VLP) methods based on object detection enjoy the rich knowledge of fine-grained object-text alignment but at the cost of computationally expensive inference. Recent Visual-Transformer (ViT)-based approaches circumvent this issue while struggling with long visual sequences without detailed cross-modal alignment information. This paper introduces a ViT-based VLP technique that efficiently incorporates object information through a novel patch-text alignment mechanism. Specifically, we convert object-level signals into patch-level ones and devise a Patch-Text Alignment pre-training task (PTA) to learn a text-aware patch detector. By using off-the-shelf delicate object annotations in 5% training images, we jointly train PTA with other conventional VLP objectives in an end-to-end manner, bypassing the high computational cost of object detection and yielding an effective patch detector that accurately detects text-relevant patches, thus considerably reducing patch sequences and accelerating computation within the ViT backbone. Our experiments on a variety of widely-used benchmarks reveal that our method achieves a speedup of nearly 88% compared to prior VLP models while maintaining competitive or superior performance on downstream tasks with similar model size and data scale. Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Bin Bi, Shikun Zhang, Fei Huang 0002, Ji Zhang 0011 |
ACM Multimedia | 7 |
| 2023 | Achieving Human Parity on Visual Question AnsweringabstractThe Visual Question Answering (VQA) task utilizes both visual image and language analysis to answer a textual question with respect to an image. It has been a popular research topic with an increasing number of real-world applications in the last decade. This paper introduces a novel hierarchical integration of vision and language AliceMind-MMU (ALIbaba’s Collection of Encoder-decoders from Machine IntelligeNce lab of Damo academy - MultiMedia Understanding) , which leads to similar or even slightly better results than a human being does on VQA. A hierarchical framework is designed to tackle the practical problems of VQA in a cascade manner including: (1) diverse visual semantics learning for comprehensive image content understanding; (2) enhanced multi-modal pre-training with modality adaptive attention; and (3) a knowledge-guided model integration with three specialized expert modules for the complex VQA task. Treating different types of visual questions with corresponding expertise needed plays an important role in boosting the performance of our VQA architecture up to the human level. An extensive set of experiments and analysis are conducted to demonstrate the effectiveness of the new research work. Ming Yan 0008, Haiyang Xu 0001, Chenliang Li 0003, Bin Bi, Wei Wang 0225, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Luo Si, Rong Jin 0001 |
ACM Trans. Inf. Syst. | 5 |
| 2022 | TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch SelectionabstractVision Transformers (ViTs) have been widely used in large-scale Vision and Language Pretraining (VLP) models.Though previous VLP works have proved the effectiveness of ViTs, they still suffer from computational efficiency brought by the long visual sequence.To tackle this problem, in this paper, we propose an efficient vision-and-language pre-training model with Text-Relevant Image Patch Selection, namely TRIPS, which reduces the visual sequence progressively with a text-guided patchselection layer in the visual backbone for efficient training and inference.The patchselection layer can dynamically compute textdependent visual attention to identify the attentive image tokens with text guidance and fuse inattentive ones in an end-to-end manner.Meanwhile, TRIPS does not introduce extra parameters to ViTs.Experimental results on a variety of popular benchmark datasets demonstrate that TRIPS gain a speedup of 40% over previous similar VLP models, yet with competitive or better downstream task performance. Chaoya Jiang, Haiyang Xu 0001, Chenliang Li 0003, Ming Yan 0008, Wei Ye 0004, Shikun Zhang, Bin Bi, Songfang Huang |
EMNLP | 7 |
| 2022 | mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connectionsabstractChenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, He Chen, Guohai Xu, Zheng Cao, Ji Zhang, Songfang Huang, Fei Huang, Jingren Zhou, Luo Si. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Chenliang Li 0003, Haiyang Xu 0001, Wei Wang 0225, Ming Yan 0008, Bin Bi, Jiabo Ye, Guohai Xu, Zheng Cao 0003, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Jingren Zhou 0001, Luo Si |
EMNLP | 6 |
| 2021 | A Unified Pretraining Framework for Passage Ranking and ExpansionabstractPretrained language models have recently advanced a wide range of natural language processing tasks. Nowadays, the application of pretrained language models to IR tasks has also achieved impressive results. Typical methods either directly apply a pretrained model to improve the re-ranking stage, or use it to conduct passage expansion and term weighting for first-stage retrieval. We observe that the passage ranking and passage expansion tasks share certain inherent relations, and can benefit from each other. Therefore, in this paper, we propose a general pretraining framework to enhance both tasks with Unified Encoder-Decoder networks (UED). The overall ranking framework consists of two parts in a cascade manner: (1) passage expansion with a pretraining-based query generation method; (2) re-ranking of passage candidates from a traditional retrieval method with a pretrained transformer encoder. Both the two parts are based on the same pretrained UED model, where we jointly train the passage ranking and query generation tasks for further improving the full ranking pipeline. An extensive set of experiments have been conducted on two large-scale passage retrieval datasets to demonstrate the state-of-the-art results of the proposed framework in both the first-stage retrieval and the final re-ranking. In addition, we successfully deploy the framework to our online production system, which can stably serve industrial applications with a request volume of up to 100 QPS in less than 300ms. Ming Yan 0008, Chenliang Li 0003, Bin Bi, Wei Wang 0225, Songfang Huang |
AAAI | 3 |
| 2021 | StructuralLM: Structural Pre-training for Form UnderstandingabstractChenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chenliang Li 0003, Bin Bi, Ming Yan 0008, Wei Wang 0225, Songfang Huang, Fei Huang 0002, Luo Si |
ACL/IJCNLP (1) | 2 |
| 2021 | VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and GenerationabstractFuli Luo, Wei Wang, Jiahao Liu, Yijia Liu, Bin Bi, Songfang Huang, Fei Huang, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fuli Luo, Wei Wang 0225, Bin Bi, Songfang Huang, Fei Huang 0002, Luo Si |
ACL/IJCNLP (1) | 5 |
| 2021 | E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual LearningabstractHaiyang Xu, Ming Yan, Chenliang Li, Bin Bi, Songfang Huang, Wenming Xiao, Fei Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Haiyang Xu 0001, Ming Yan 0008, Chenliang Li 0003, Bin Bi, Songfang Huang, Wenming Xiao, Fei Huang 0002 |
ACL/IJCNLP (1) | 4 |
| 2021 | AliMe DA: A Data Augmentation Framework for Question Answering in Cold-start ScenariosabstractCold-start is the most difficult and time-consuming phase when building a question answering based chatbot for a new business scenario because of the collection of sufficient training data. In this paper, we propose AliMe DA, a practical data augmentation (DA) framework that consists of data production, denoising and consumption, to alleviate this problem. We show how our DA approach can be used to substantially enhance annotation productivity and also improve downstream model performance. More importantly, we provide best practices for data augmentation, including how to choose and employ appropriate methods at each stage of our framework, and share our observation on the applicable scene of data augmentation in the era of pre-trained language models. Guohai Xu, Chenliang Li 0003, Feng-Lin Li, Bin Bi, Ji Zhang 0011, Haiqing Chen |
SIGIR | 5 |
| 2020 | Generating Well-Formed Answers by Machine Reading with Stochastic Selector Networks
Bin Bi, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Jiangnan Xia, Chenliang Li 0003 |
AAAI | 1 |
| 2020 | PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned GenerationabstractSelf-supervised pre-training, such as BERT (Devlin et al., 2018), MASS (Song et al., 2019) and BART (Lewis et al., 2019), has emerged as a powerful technique for natural language understanding and generation.Existing pre-training techniques employ autoencoding and/or autoregressive objectives to train Transformer-based models by recovering original word tokens from corrupted text with some masked tokens.The training goals of existing techniques are often inconsistent with the goals of many language generation tasks, such as generative question answering and conversational response generation, for producing new text given context.This work presents PALM with a novel scheme that jointly pre-trains an autoencoding and autoregressive language model on a large unlabeled corpus, specifically designed for generating new text conditioned on context.The new scheme alleviates the mismatch introduced by the existing denoising scheme between pre-training and fine-tuning where generation is more than reconstructing original text.An extensive set of experiments show that PALM achieves new state-of-theart results on a variety of language generation benchmarks covering generative question answering (Rank 1 on the official MARCO leaderboard), abstractive summarization on CNN/DailyMail as well as Gigaword, question generation on SQuAD, and conversational response generation on Cornell Movie Dialogues. Bin Bi, Chenliang Li 0003, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Songfang Huang, Fei Huang 0002, Luo Si |
EMNLP (1) | 1 |
| 2020 | StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
Wei Wang 0225, Bin Bi, Ming Yan 0008, Chen Wu 0006, Jiangnan Xia, Zuyi Bao, Liwei Peng, Luo Si |
ICLR | 2 |
| 2020 | Latent Template Induction with Gumbel-CRFsabstractLearning to control the structure of sentences is a challenging problem in text generation. Existing work either relies on simple deterministic approaches or RL-based hard structures. We explore the use of structured variational autoencoders to infer latent templates for sentence generation using a soft, continuous relaxation in order to utilize reparameterization for training. Specifically, we propose a Gumbel-CRF, a continuous relaxation of the CRF sampling algorithm using a relaxed Forward-Filtering Backward-Sampling (FFBS) approach. As a reparameterized gradient estimator, the Gumbel-CRF gives more stable gradients than score-function based estimators. As a structured inference network, we show that it learns interpretable templates during training, which allows us to control the decoder during testing. We demonstrate the effectiveness of our methods with experiments on data-to-text generation and unsupervised paraphrase generation. Chuanqi Tan, Bin Bi, Mosha Chen, Yansong Feng 0002, Alexander M. Rush |
NeurIPS | 3 |
| 2019 | A Deep Cascade Model for Multi-Document Reading ComprehensionabstractA fundamental trade-off between effectiveness and efficiency needs to be balanced when designing an online question answering system. Effectiveness comes from sophisticated functions such as extractive machine reading comprehension (MRC), while efficiency is obtained from improvements in preliminary retrieval components such as candidate document selection and paragraph ranking. Given the complexity of the real-world multi-document MRC scenario, it is difficult to jointly optimize both in an end-to-end system. To address this problem, we develop a novel deep cascade learning model, which progressively evolves from the documentlevel and paragraph-level ranking of candidate texts to more precise answer extraction with machine reading comprehension. Specifically, irrelevant documents and paragraphs are first filtered out with simple functions for efficiency consideration. Then we jointly train three modules on the remaining texts for better tracking the answer: the document extraction, the paragraph extraction and the answer extraction. Experiment results show that the proposed method outperforms the previous state-of-the-art methods on two large-scale multidocument benchmark datasets, i.e., TriviaQA and DuReader. In addition, our online system can stably serve typical scenarios with millions of daily requests in less than 50ms. Ming Yan 0008, Jiangnan Xia, Chen Wu 0006, Bin Bi, Zhongzhou Zhao, Ji Zhang 0011, Luo Si, Rui Wang 0005, Wei Wang 0225, Haiqing Chen |
AAAI | 4 |
| 2019 | Incorporating External Knowledge into Machine Reading for Generative Question AnsweringabstractBin Bi, Chen Wu, Ming Yan, Wei Wang, Jiangnan Xia, Chenliang Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bin Bi, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Jiangnan Xia, Chenliang Li 0003 |
EMNLP/IJCNLP (1) | 1 |
| 2016 | Modeling a Retweet Network via an Adaptive Bayesian ApproachabstractTwitter (and similar microblogging services) has become a central nexus for discussion of the topics of the day. Twitter data contains rich content and structured information on users' topics of interest and behavior patterns. Correctly analyzing and modeling Twitter data enables the prediction of the user behavior and preference in a variety of practical applications, such as tweet recommendation and followee recommendation. Although a number of models have been developed on Twitter data in prior work, most of these only model the tweets from users, while neglecting their valuable retweet information in the data. Models would enhance their predictive power by incorporating users' retweet content as well as their retweet behavior. In this paper, we propose two novel Bayesian nonparametric models, URM and UCM, on retweet data. Both of them are able to integrate the analysis of tweet text and users' retweet behavior in the same probabilistic framework. Moreover, they both jointly model users' interest in tweet and retweet. As nonparametric models, URM and UCM can automatically determine the parameters of the models based on input data, avoiding arbitrary parameter settings. Extensive experiments on real-world Twitter data show that both URM and UCM are superior to all the baselines, while UCM further outperforms URM, confirming the appropriateness of our models in retweet modeling. Bin Bi, Junghoo Cho |
WWW | 1 |
| 2015 | Learning to Recommend Related Entities to Search UsersabstractOver the past few years, major web search engines have introduced knowledge bases to offer popular facts about people, places, and things on the entity pane next to regular search results. In addition to information about the entity searched by the user, the entity pane often provides a ranked list of related entities. To keep users engaged, it is important to develop a recommendation model that tailors the related entities to individual user interests. We propose a probabilistic Three-way Entity Model (TEM) that provides personalized recommendation of related entities using three data sources: knowledge base, search click log, and entity pane log. Specifically, TEM is capable of extracting hidden structures and capturing underlying correlations among users, main entities, and related entities. Moreover, the TEM model can also exploit the click signals derived from the entity pane log. We further provide an inference technique to learn the parameters in TEM, and propose a principled preference learning method specifically designed for ranking related entities. Extensive experiments with two real-world datasets show that TEM with our probabilistic framework significantly outperforms a state of the art baseline, confirming the effectiveness of TEM and our probabilistic framework in related entity recommendation. Bin Bi, Hao Ma 0001, Bo-June Paul Hsu, Kuansan Wang, Junghoo Cho |
WSDM | 1 |
| 2014 | Who are experts specializing in landscape photography?: analyzing topic-specific authority on content sharing servicesabstractWith the rapid growth of Web 2.0, a variety of content sharing services, such as Flickr, YouTube, Blogger, and TripAdvisor etc, have become extremely popular over the last decade. On these websites, users have created and shared with each other various kinds of resources, such as photos, video, and travel blogs. The sheer amount of user-generated content varies greatly in quality, which calls for a principled method to identify a set of authorities, who created high-quality resources, from a massive number of contributors of content. Since most previous studies only infer global authoritativeness of a user, there is no way to differentiate the authoritativeness in different aspects of life (topics). Bin Bi, Ben Kao, Chang Wan, Junghoo Cho |
KDD | 1 |
| 2014 | Scalable topic-specific influence analysis on microblogsabstractSocial influence analysis on microblog networks, such as Twitter, has been playing a crucial role in online advertising and brand management. While most previous influence analysis schemes rely only on the links between users to find key influencers, they omit the important text content created by the users. As a result, there is no way to differentiate the social influence in different aspects of life (topics). Although a few prior works do support topic-specific influence analysis, they either separate the analysis of content from the analysis of network structure, or assume that content is the only cause of links, which is clearly an inappropriate assumption for microblog networks. Bin Bi, Yuanyuan Tian 0001, Yannis Sismanis, Andrey Balmin, Junghoo Cho |
WSDM | 1 |
| 2013 | Automatically generating descriptions for resources by tag modelingabstractWe have been witnessing an increasing number of social tagging systems on the web. Tags help users understand a resource readily and accurately. In a social tagging system, however, there are typically a fairly large number of resources each associated with a long list of tags. When browsing resources, users are reluctant to read these tags one by one. Instead, users prefer a shorter list of tags as a compact description of a resource. Such a tag description facilitates users to understand the resource accurately and effortlessly. This calls for a generator for a tag description, which selects a set of high-quality tags for a given resource. The tag description condenses the original tag list by retaining the most important tags of the long list. Bin Bi, Junghoo Cho |
CIKM | 1 |
| 2013 | Incorporating popularity in topic models for social network analysisabstractTopic models are used to group words in a text dataset into a set of relevant topics. Unfortunately, when a few words frequently appear in a dataset, the topic groups identified by topic models become noisy because these frequent words repeatedly appear in "irrelevant" topic groups. This noise has not been a serious problem in a text dataset because the frequent words (e.g., the and is) do not have much meaning and have been simply removed before a topic model analysis. However, in a social network dataset we are interested in, they correspond to popular persons (e.g., Barack Obama and Justin Bieber) and cannot be simply removed because most people are interested in them. Youngchul Cha, Bin Bi, Chu-Cheng Hsieh, Junghoo Cho |
SIGIR | 2 |
| 2013 | Inferring the demographics of search users: social data meets search queriesabstractKnowing users' views and demographic traits offers a great potential for personalizing web search results or related services such as query suggestion and query completion. Such signals however are often only available for a small fraction of search users, namely those who log in with their social network account and allow its use for personalization of search results. In this paper, we offer a solution to this problem by showing how user demographic traits such as age and gender, and even political and religious views can be efficiently and accurately inferred based on their search query histories. This is accomplished in two steps; we first train predictive models based on the publically available myPersonality dataset containing users' Facebook Likes and their demographic information. We then match Facebook Likes with search queries using Open Directory Project categories. Finally, we apply the model trained on Facebook Likes to large-scale query logs of a commercial search engine while explicitly taking into account the difference between the traits distribution in both datasets. We find that the accuracy of classifying age and gender, expressed by the area under the ROC curve (AUC), are 77% and 84% respectively for predictions based on Facebook Likes, and only degrade to 74% and 80% when based on search queries. On a US state-by-state basis we find a Pearson correlation of 0.72 for political views between the predicted scores and Gallup data, and 0.54 for affiliation with Judaism between predicted scores and data from the US Religious Landscape Survey. We conclude that it is indeed feasible to infer important demographic data of users from their query history based on labelled Likes data and believe that this approach could provide valuable information for personalization and monetization even in the absence of demographic data. Bin Bi, Milad Shokouhi, Michal Kosinski, Thore Graepel |
WWW | 1 |
| 2012 | DQR: a probabilistic approach to diversified query recommendationabstractWeb search queries issued by casual users are often short and with limited expressiveness. Query recommendation is a popular technique employed by search engines to help users refine their queries. Traditional similarity-based methods, however, often result in redundant and monotonic recommendations. We identify five basic requirements of a query recommendation system. In particular, we focus on the requirements of redundancy-free and diversified recommendations. We propose the DQR framework, which mines a search log to achieve two goals: (1) It clusters search log queries to extract query concepts, based on which recommended queries are selected. (2) It employs a probabilistic model and a greedy heuristic algorithm to achieve recommendation diversification. Through a comprehensive user study we compare DQR against five other recommendation methods. Our experiment shows that DQR outperforms the other methods in terms of relevancy, diversity, and ranking performance of the recommendations. Ben Kao, Bin Bi, Reynold Cheng, Eric Lo 0001 |
CIKM | 3 |
| 2011 | CubeLSI: An effective and efficient method for searching resources in social tagging systemsabstractIn a social tagging system, resources (such as photos, video and web pages) are associated with tags. These tags allow the resources to be effectively searched through tag-based keyword matching using traditional IR techniques. We note that in many such systems, tags of a resource are often assigned by a diverse audience of causal users (taggers). This leads to two issues that gravely affect the effectiveness of resource retrieval: (1) Noise: tags are picked from an uncontrolled vocabulary and are assigned by untrained taggers. The tags are thus noisy features in resource retrieval. (2) A multitude of aspects: different taggers focus on different aspects of a resource. Representing a resource using a flattened bag of tags ignores this important diversity of taggers. To improve the effectiveness of resource retrieval in social tagging systems, we propose CubeLSI - a technique that extends traditional LSI to include taggers as another dimension of feature space of resources. We compare CubeLSI against a number of other tag-based retrieval models and show that CubeLSI significantly outperforms the other models in terms of retrieval accuracy. We also prove two interesting theorems that allow CubeLSI to be very efficiently computed despite the much enlarged feature space it employs. Bin Bi, Sau Dan Lee, Ben Kao, Reynold Cheng |
ICDE | 1 |
| 2009 | Collaborative resource discovery in social tagging systemsabstractSocial tagging systems which allow users to create, edit and share collections of internet resources associated with tags in a collaborative fashion are growing in popularity in recent years. The rapidly growing amount of shared data in these folksonomies, i.e., taxonomies created by the folk, presents new technical challenges involved with discovering resources which are likely of interest to the user. Social tags which reflect the meaning of resources from the user's points of view provide an opportunity to enhance the quality of retrieval. In this paper, we introduce a novel framework to search relevant resources to the user query by incorporating information obtained from folksonomies' underlying data structures consisting of a set of user/tag/resource triplets. In contrast to traditional retrieval and recommendation techniques which represent a collection by a matrix, we represent our data as a third-order tensor on which a novel Cube Latent Semantic Indexing (CubeLSI) technique is proposed to capture latent semantic associations between tags. With the latent semantic representation we show how to rank relevant resources according to their relevance to user queries. The excellent performance of the method is demonstrated by an experimental evaluation on the deli.cio.us dataset. Copyright 2009 ACM. Bin Bi, Lifeng Shang, Ben Kao |
CIKM | 1 |