Xiaochi Wei

dblp:131/2938 · DBLP profile ↗
← Back
36ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-4359-4024ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 17 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Thought Space Exploration Through Strategic Intervention
abstract
While large language models (LLMs) demonstrate emerging reasoning capabilities, current inference-time expansion methods incur prohibitive computational costs through exhaustive sampling. Through analyzing decoding trajectories, we observe that most next-token predictions align well with the golden output, except for a few critical tokens that lead to deviations. Inspired by this phenomenon, we propose a novel Hint-Practice Reasoning (HPR) framework that operationalizes this insight through two synergistic components: 1) a hinter (powerful LLM) that provides probabilistic guidance at critical decision points, and 2) a practitioner (efficient smaller model) that executes major reasoning steps. The framework's core innovation lies in Distributional Inconsistency Reduction (DIR), a theoretically-grounded metric that dynamically identifies intervention points by quantifying the divergence between practitioner's reasoning trajectory and hinter's expected distribution in a tree-structured probabilistic space. Through iterative tree updates guided by DIR, HPR reweights promising reasoning paths while deprioritizing low-probability branches. Experiments across arithmetic and commonsense reasoning benchmarks demonstrate HPR's state-of-the-art efficiency-accuracy tradeoffs: it achieves comparable performance to self-consistency and MCTS baselines while decoding only 1/5 tokens, and outperforms existing methods by at most 5.1% absolute accuracy while maintaining similar or lower FLOPs.
Ziheng Li 0003, Hengyi Cai, Xiaochi Wei, Yuchen Li 0006, Shuaiqiang Wang, Zhi-Hong Deng 0001, Dawei Yin 0001
AAAI3
2025 LLMs + Persona-Plug = Personalized LLMs
abstract
Jiongnan Liu, Yutao Zhu, Shuting Wang, Xiaochi Wei, Erxue Min, Yu Lu, Shuaiqiang Wang, Dawei Yin, Zhicheng Dou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jiongnan Liu 0001, Yutao Zhu 0001, Shuting Wang 0002, Xiaochi Wei, Erxue Min, Yu Lu 0009, Shuaiqiang Wang, Dawei Yin 0001, Zhicheng Dou
ACL (1)4
2025 Enhancing Retrieval-Augmented Generation via Evidence Tree Search
abstract
Retrieval-Augmented Generation (RAG) is widely used to enhance Large Language Models (LLMs) by grounding responses in external knowledge. However, in real-world applications, retrievers often return lengthy documents with redundant or irrelevant content, confusing downstream readers. While evidence retrieval aims to address this by extracting key information, it faces critical challenges: (1) inability to model synergistic inter-dependencies among evidence sentences, (2) lack of supervision for evaluating multi-sentence evidence quality, and (3) computational inefficiency in navigating exponentially growing search spaces of candidate evidence sets. To tackle these challenges, we propose ETS (Evidence Tree Search), a novel framework that reformulates evidence retrieval as a dynamic tree expansion process. Our approach first constructs an evidence tree where each path represents a candidate evidence set, explicitly modeling inter-sentence dependencies through context-aware node selection. We then leverage Monte Carlo Tree Search (MCTS) to efficiently assess evidence quality and introduce an Early-Terminating Beam Search strategy to efficiently accelerate the model inference. Extensive experiments on five datasets demonstrate that ETS significantly outperforms existing methods across different readers. Our code and datasets will be released to facilitate future research.
Hao Sun 0015, Hengyi Cai, Yuchen Li 0006, Xuanbo Fan, Xiaochi Wei, Shuaiqiang Wang, Yan Zhang 0117, Dawei Yin 0001
ACL (1)5
2025 From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions
abstract
Tool learning enables Large Language Models (LLMs) to interact with external environments by invoking tools, serving as an effective strategy to mitigate the limitations inherent in their pre-training data. In this process, tool documentation plays a crucial role by providing usage instructions for LLMs, thereby facilitating effective tool utilization. This paper concentrates on the critical challenge of bridging the comprehension gap between LLMs and external tools due to the inadequacies and inaccuracies inherent in existing human-centric tool documentation. We propose a novel framework, DRAFT, aimed at Dynamically Refining tool documentation through the Analysis of Feedback and Trials emanating from LLMs' interactions with external tools. This methodology pivots on an innovative trial-and-error approach, consisting of three distinct learning phases: experience gathering, learning from experience, and documentation rewriting, to iteratively enhance the tool documentation. This process is further optimized by implementing a diversity-promoting exploration strategy to ensure explorative diversity and a tool-adaptive termination mechanism to prevent overfitting while enhancing efficiency. Extensive experiments on multiple datasets demonstrate that DRAFT's iterative, feedback-based refinement significantly ameliorates documentation quality, fostering a deeper comprehension and more effective utilization of tools by LLMs. Notably, our analysis reveals that the tool documentation refined via our approach demonstrates robust cross-model generalization capabilities.
Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin 0001, Jun Xu 0001, Ji-Rong Wen
ICLR3
2025 Tool learning with large language models: a survey
Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin 0001, Jun Xu 0001, Ji-Rong Wen
Frontiers Comput. Sci.3
2024 Towards Completeness-Oriented Tool Retrieval for Large Language Models
abstract
Recently, integrating external tools with Large Language Models (LLMs) has gained significant attention as an effective strategy to mitigate the limitations inherent in their pre-training data. However, real-world systems often incorporate a wide array of tools, making it impractical to input all tools into LLMs due to length limitations and latency constraints. Therefore, to fully exploit the potential of tool-augmented LLMs, it is crucial to develop an effective tool retrieval system. Existing tool retrieval methods primarily focus on semantic matching between user queries and tool descriptions, frequently leading to the retrieval of redundant, similar tools. Consequently, these methods fail to provide a complete set of diverse tools necessary for addressing the multifaceted problems encountered by LLMs. In this paper, we propose a novel modelagnostic CO llaborative L earning-based T ool Retrieval approach, COLT, which captures not only the semantic similarities between user queries and tool descriptions but also takes into account the collaborative information of tools. Specifically, we first fine-tune the PLM-based retrieval models to capture the semantic relationships between queries and tools in the semantic learning stage. Subsequently, we construct three bipartite graphs among queries, scenes, and tools and introduce a dual-view graph collaborative learning framework to capture the intricate collaborative relationships among tools during the collaborative learning stage. Extensive experiments on both the open benchmark and the newly introduced ToolLens dataset show that COLT achieves superior performance. Notably, the performance of BERT-mini (11M) with our proposed model framework outperforms BERT-large (340M), which has 30 times more parameters. Furthermore, we will release ToolLens publicly to facilitate future research on tool retrieval.
Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin 0001, Jun Xu 0001, Ji-Rong Wen
CIKM3
2024 AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning
abstract
Recent advancements in large language models (LLMs) have been remarkable.Users face a choice between using cloud-based LLMs for generation quality and deploying local-based LLMs for lower computational cost.The former option is typically costly and inefficient, while the latter usually fails to deliver satisfactory performance for reasoning steps requiring deliberate thought processes.In this work, we propose a novel LLM utilization paradigm that facilitates the collaborative operation of large cloud-based LLMs and smaller local-deployed LLMs.Our framework comprises two primary modules: the local agent instantiated with a relatively smaller LLM, handling less complex reasoning steps, and the cloud agent equipped with a larger LLM, managing more intricate reasoning steps.This collaborative processing is enabled through an adaptive mechanism where the local agent introspectively identifies errors and proactively seeks assistance from the cloud agent, thereby effectively integrating the strengths of both locally-deployed and cloudbased LLMs, resulting in significant enhancements in task completion performance and efficiency.We evaluate ADASWITCH across 7 benchmarks, ranging from mathematical reasoning and complex question answering, using various types of LLMs to instantiate the local and cloud agents.The empirical results show that ADASWITCH effectively improves the performance of the local agent, and sometimes achieves competitive results compared to the cloud agent while utilizing much less computational overhead.Question: Joy is 2 years older than twice the age of Tom.If Tom is 10 years old, how old is Joy?Thought: Calculate twice the age of Tom.Action: R1 = Calculator(2 * 2) Observation: 4 Reflection: Is previous step wrong: Yes Thought: Calculate twice the age of Tom who is 10.Action: R1 = Calculator(10 * 2) Observation: 20 Reflection: Is previous step wrong: No Thought: Calculate the age of Joy.
Hao Sun 0015, Jiayi Wu 0001, Hengyi Cai, Xiaochi Wei, Yue Feng 0002, Bo Wang 0134, Shuaiqiang Wang, Yan Zhang 0117, Dawei Yin 0001
EMNLP4
2024 Towards Verifiable Text Generation with Evolving Memory and Self-Reflection
abstract
Despite the remarkable ability of large language models (LLMs) in language comprehension and generation, they often suffer from producing factually incorrect information, also known as hallucination.A promising solution to this issue is verifiable text generation, which prompts LLMs to generate content with citations for accuracy verification.However, verifiable text generation is non-trivial due to the focus-shifting phenomenon, the intricate reasoning needed to align the claim with correct citations, and the dilemma between the precision and breadth of retrieved documents.In this paper, we present VTG, an innovative framework for Verifiable Text Generation with evolving memory and self-reflection.VTG introduces evolving long short-term memory to retain both valuable documents and recent documents.A two-tier verifier equipped with an evidence finder is proposed to rethink and reflect on the relationship between the claim and citations.Furthermore, active retrieval and diverse query generation are utilized to enhance both the precision and breadth of the retrieved documents.We conduct extensive experiments on five datasets across three knowledge-intensive tasks and the results reveal that VTG significantly outperforms baselines.
Hao Sun 0015, Hengyi Cai, Bo Wang 0134, Yingyan Hou, Xiaochi Wei, Shuaiqiang Wang, Yan Zhang 0117, Dawei Yin 0001
EMNLP5
2023 Knowing Before Seeing: Incorporating Post-retrieval Information into Pre-retrieval Query Intention Classification
Xueqing Ma, Xiaochi Wei, Yixing Gao 0001, Runyang Feng, Dawei Yin 0001, Yi Chang 0001
KSEM (2)2
2022 BIT-WOW at NLPCC-2022 Task5 Track1: Hierarchical Multi-label Classification via Label-Aware Graph Convolutional Network
Bo Wang 0134, Yi-Fan Lu, Xiaochi Wei, Xiao Liu 0029, Ge Shi 0002, Changsen Yuan, Heyan Huang, Chong Feng 0001, Xianling Mao
NLPCC (2)3
2021 Multi-Modal Relational Graph for Cross-Modal Video Moment Retrieval
abstract
Given an untrimmed video and a query sentence, cross-modal video moment retrieval aims to rank a video moment from pre-segmented video moment candidates that best matches the query sentence. Pioneering work typically learns the representations of the textual and visual content separately and then obtains the interactions or alignments between different modalities. However, the task of cross-modal video moment retrieval is not yet thoroughly addressed as it needs to further identify the fine-grained differences of video moment candidates with high repeatability and similarity. Moveover, the relation among objects in both video and sentence is intuitive and efficient for understanding semantics but is rarely considered.Toward this end, we contribute a multi-modal relational graph to capture the interactions among objects from the visual and textual content to identify the differences among similar video moment candidates. Specifically, we first introduce a visual relational graph and a textual relational graph to form relation-aware representations via message propagation. Thereafter, a multi-task pre-training is designed to capture domain-specific knowledge about objects and relations, enhancing the structured visual representation after explicitly defined relation. Finally, the graph matching and boundary regression are employed to perform the cross-modal retrieval. We conduct extensive experiments on two datasets about daily activities and cooking activities, demonstrating significant improvements over state-of-the-art solutions.
Yawen Zeng, Da Cao, Xiaochi Wei, Meng Liu 0006, Zhou Zhao 0001, Zheng Qin 0001
CVPR3
2021 Document-level relation extraction with Entity-Selection Attention
Changsen Yuan, Heyan Huang, Chong Feng 0001, Ge Shi 0002, Xiaochi Wei
Inf. Sci.5
2020 Adversarial Video Moment Retrieval by Jointly Modeling Ranking and Localization
abstract
Retrieving video moments from an untrimmed video given a natural language as the query is a challenging task in both academia and industry. Although much effort has been made to address this issue, traditional video moment ranking methods are unable to generate reasonable video moment candidates and video moment localization approaches are not applicable to large-scale retrieval scenario. How to combine ranking and localization into a unified framework to overcome their drawbacks and reinforce each other is rarely considered. Toward this end, we contribute a novel solution to thoroughly investigate the video moment retrieval issue under the adversarial learning paradigm. The key of our solution is to formulate the video moment retrieval task as an adversarial learning problem with two tightly connected components. Specifically, a reinforcement learning is employed as a generator to produce a set of possible video moments. Meanwhile, a pairwise ranking model is utilized as a discriminator to rank the generated video moments and the ground truth. Finally, the generator and the discriminator are mutually reinforced in the adversarial learning framework, which is able to jointly optimize the performance of both video moment ranking and video moment localization. Extensive experiments on two well-known datasets have well verified the effectiveness and rationality of our proposed solution.
Da Cao, Yawen Zeng, Xiaochi Wei, Liqiang Nie, Richang Hong, Zheng Qin 0001
ACM Multimedia3
2020 Video-based recipe retrieval
Da Cao, Ning Han 0005, Hao Chen 0051, Xiaochi Wei, Xiangnan He 0001
Inf. Sci.4
2020 Similarity-aware neural machine translation: reducing human translator efforts by leveraging high-potential sentences with translation memory
Tianfu Zhang, Heyan Huang, Chong Feng 0001, Xiaochi Wei
Neural Comput. Appl.4
2020 A Discriminative Convolutional Neural Network with Context-aware Attention
abstract
Feature representation and feature extraction are two crucial procedures in text mining. Convolutional Neural Networks (CNN) have shown overwhelming success for text-mining tasks, since they are capable of efficiently extracting n -gram features from source data. However, vanilla CNN has its own weaknesses on feature representation and feature extraction. A certain amount of filters in CNN are inevitably duplicate and thus hinder to discriminatively represent a given text. In addition, most existing CNN models extract features in a fixed way (i.e., max pooling) that either limit the CNN to local optimum nor without considering the relation between all features, thereby unable to learn a contextual n -gram features adaptively. In this article, we propose a discriminative CNN with context-aware attention to solve the challenges of vanilla CNN. Specifically, our model mainly encourages discrimination across different filters via maximizing their earth mover distances and estimates the salience of feature candidates by considering the relation between context features. We validate carefully our findings against baselines on five benchmark datasets of classification and two datasets of summarization. The results of the experiments verify the competitive performance of our proposed model.
Lejian Liao, Yang Gao 0016, Heyan Huang, Xiaochi Wei
ACM Trans. Intell. Syst. Technol.5
2019 Distant Supervision for Relation Extraction with Linear Attenuation Simulation and Non-IID Relevance Embedding
abstract
Distant supervision for relation extraction is an efficient method to reduce labor costs and has been widely used to seek novel relational facts in large corpora, which can be identified as a multi-instance multi-label problem. However, existing distant supervision methods suffer from selecting important words in the sentence and extracting valid sentences in the bag. Towards this end, we propose a novel approach to address these problems in this paper. Firstly, we propose a linear attenuation simulation to reflect the importance of words in the sentence with respect to the distances between entities and words. Secondly, we propose a non-independent and identically distributed (non-IID) relevance embedding to capture the relevance of sentences in the bag. Our method can not only capture complex information of words about hidden relations, but also express the mutual information of instances in the bag. Extensive experiments on a benchmark dataset have well-validated the effectiveness of the proposed method.
Changsen Yuan, Heyan Huang, Chong Feng 0001, Xiao Liu 0029, Xiaochi Wei
AAAI5
2019 HSDS: An Abstractive Model for Automatic Survey Generation
Xiao-Jian Jiang, Xianling Mao, Bo-Si Feng, Xiaochi Wei, Bin-Bin Bian, Heyan Huang
DASFAA (1)4
2019 Neural Variational Correlated Topic Modeling
abstract
With the rapid development of the Internet, millions of documents, such as news and web pages, are generated everyday. Mining the topics and knowledge on them has attracted a lot of interest on both academic and industrial areas. As one of the prevalent unsupervised data mining tools, topic models are usually explored as probabilistic generative models for large collections of texts. Traditional probabilistic topic models tend to find a closed form solution of model parameters and approach the intractable posteriors via approximation methods, which usually lead to the inaccurate inference of parameters and low efficiency when it comes to a quite large volume of data. Recently, an emerging trend of neural variational inference can overcome the above issues, which offers a scalable and powerful deep generative framework for modeling latent topics via neural networks. Interestingly, a common assumption for the most neural variational topic models is that topics are independent and irrelevant to each other. However, this assumption is unreasonable in many practical scenarios. In this paper, we propose a novel Centralized Transformation Flow to capture the correlations among topics by reshaping topic distributions. Furthermore, we present the Transformation Flow Lower Bound to improve the performance of the proposed model. Extensive experiments on two standard benchmark datasets have well-validated the effectiveness of the proposed approach.
Heyan Huang, Yang Gao 0016, Xiaochi Wei
WWW5
2019 Mapping sentences to concept transferred space for semantic textual similarity
Heyan Huang, Hao Wu 0066, Xiaochi Wei, Yang Gao 0016, Shumin Shi
Knowl. Inf. Syst.3
2019 From Question to Text: Question-Oriented Feature Attention for Answer Selection
abstract
Understanding unstructured texts is an essential skill for human beings as it enables knowledge acquisition. Although understanding unstructured texts is easy for we human beings with good education, it is a great challenge for machines. Recently, with the rapid development of artificial intelligence techniques, researchers put efforts to teach machines to understand texts and justify the educated machines by letting them solve the questions upon the given unstructured texts, inspired by the reading comprehension test as we humans do. However, feature effectiveness with respect to different questions significantly hinders the performance of answer selection, because different questions may focus on various aspects of the given text and answer candidates. To solve this problem, we propose a question-oriented feature attention (QFA) mechanism, which learns to weight different engineering features according to the given question, so that important features with respect to the specific question is emphasized accordingly. Experiments on MCTest dataset have well-validated the effectiveness of the proposed method. Additionally, the proposed QFA is applicable to various IR tasks, such as question answering and answer selection. We have verified the applicability on a crawled community-based question-answering dataset.
Heyan Huang, Xiaochi Wei, Liqiang Nie, Xianling Mao, Xin-Shun Xu
ACM Trans. Inf. Syst.2
2018 Task-oriented Word Embedding for Text Classification
abstract
Distributed word representation plays a pivotal role in various natural language processing tasks. In spite of its success, most existing methods only consider contextual information, which is suboptimal when used in various tasks due to a lack of task-specific features. The rational word embeddings should have the ability to capture both the semantic features and task-specific features of words. In this paper, we propose a task-oriented word embedding method and apply it to the text classification task. With the function-aware component, our method regularizes the distribution of words to enable the embedding space to have a clear classification boundary. We evaluate our method using five text classification datasets. The experiment results show that our method significantly outperforms the state-of-the-art methods.
Qian Liu 0012, Heyan Huang, Yang Gao 0016, Xiaochi Wei
COLING4
2018 Quality Matters: Assessing cQA Pair Quality via Transductive Multi-View Learning
abstract
Community-based question answering (cQA) sites have become important knowledge sharing platforms, as massive cQA pairs are archived, but the uneven quality of cQA pairs leaves information seekers unsatisfied. Various efforts have been dedicated to predicting the quality of cQA contents. Most of them concatenate different features into single vectors and then feed them into regression models. In fact, the quality of cQA pairs is influenced by different views, and the agreement among them is essential for quality assessment. Besides, the lacking of labeled data significantly hinders the quality prediction performance. Toward this end, we present a transductive multi-view learning model. It is designed to find a latent common space by unifying and preserving information from various views, including question, answer, QA relevance, asker, and answerer. Additionally, rich information in the unlabeled test cQA pairs are utilized via transductive learning to enhance the representation ability of the common space. Extensive experiments on real-world datasets have well-validated the proposed model.
Xiaochi Wei, Heyan Huang, Liqiang Nie, Fuli Feng, Richang Hong, Tat-Seng Chua
IJCAI1
2017 Leveraging Pattern Associations for Word Embedding Models
Qian Liu 0012, Heyan Huang, Yang Gao 0016, Xiaochi Wei, Ruiying Geng
DASFAA (1)4
2017 Embedding Factorization Models for Jointly Recommending Items and User Generated Lists
abstract
Existing recommender algorithms mainly focused on recommending individual items by utilizing user-item interactions. However, little attention has been paid to recommend user generated lists (e.g., playlists and booklists). On one hand, user generated lists contain rich signal about item co-occurrence, as items within a list are usually gathered based on a specific theme. On the other hand, a user's preference over a list also indicate her preference over items within the list. We believe that 1) if the rich relevance signal within user generated lists can be properly leveraged, an enhanced recommendation for individual items can be provided, and 2) if user-item and user-list interactions are properly utilized, and the relationship between a list and its contained items is discovered, the performance of user-item and user-list recommendations can be mutually reinforced.
Da Cao, Liqiang Nie, Xiangnan He 0001, Xiaochi Wei, Shunzhi Zhu, Tat-Seng Chua
SIGIR4
2017 Version-sensitive mobile App recommendation
Da Cao, Liqiang Nie, Xiangnan He 0001, Xiaochi Wei, Jialie Shen 0001, Shunxiang Wu, Tat-Seng Chua
Inf. Sci.4
2017 Data-Driven Answer Selection in Community QA Systems
abstract
Finding similar questions from historical archives has been applied to question answering, with well theoretical underpinnings and great practical success. Nevertheless, each question in the returned candidate pool often associates with multiple answers, and hence users have to painstakingly browse a lot before finding the correct one. To alleviate such problem, we present a novel scheme to rank answer candidates via pairwise comparisons. In particular, it consists of one offline learning component and one online search component. In the offline learning component, we first automatically establish the positive, negative, and neutral training samples in terms of preference pairs guided by our data-driven observations. We then present a novel model to jointly incorporate these three types of training samples. The closed-form solution of this model is derived. In the online search component, we first collect a pool of answer candidates for the given question via finding its similar questions. We then sort the answer candidates by leveraging the offline trained model to judge the preference orders. Extensive experiments on the real-world vertical and general community-based question answering datasets have comparatively demonstrated its robustness and promising performance. Also, we have released the codes and data to facilitate other researchers.
Liqiang Nie, Xiaochi Wei, Dongxiang Zhang, Xiang Wang 0010, Zhipeng Gao 0002, Yi Yang 0001
IEEE Trans. Knowl. Data Eng.2
2017 I Know What You Want to Express: Sentence Element Inference by Incorporating External Knowledge Base
abstract
Sentence auto-completion is an important feature that saves users many keystrokes in typing the entire sentence by providing suggestions as they type. Despite its value, the existing sentence auto-completion methods, such as query completion models, can hardly be applied to solving the object completion problem in sentences with the form of (subject, verb, object), due to the complex natural language description and the data deficiency problem. Towards this goal, we treat an SVO sentence as a three-element triple (subject, sentence pattern, object), and cast the sentence object completion problem as an element inference problem. These elements in all triples are encoded into a unified low-dimensional embedding space by our proposed TRANSFER model, which leverages the external knowledge base to strengthen the representation learning performance. With such representations, we can provide reliable candidates for the desired missing element by a linear model. Extensive experiments on a real-world dataset have well-validated our model. Meanwhile, we have successfully applied our proposed model to factoid question answering systems for answer candidate selection, which further demonstrates the applicability of the TRANSFER model.
Xiaochi Wei, Heyan Huang, Liqiang Nie, Hanwang Zhang, Xianling Mao, Tat-Seng Chua
IEEE Trans. Knowl. Data Eng.1
2017 Cross-Platform App Recommendation by Jointly Modeling Ratings and Texts
abstract
Over the last decade, the renaissance of Web technologies has transformed the online world into an application (App) driven society. While the abundant Apps have provided great convenience, their sheer number also leads to severe information overload, making it difficult for users to identify desired Apps. To alleviate the information overloading issue, recommender systems have been proposed and deployed for the App domain. However, existing work on App recommendation has largely focused on one single platform (e.g., smartphones), while it ignores the rich data of other relevant platforms (e.g., tablets and computers). In this article, we tackle the problem of cross-platform App recommendation, aiming at leveraging users’ and Apps’ data on multiple platforms to enhance the recommendation accuracy. The key advantage of our proposal is that by leveraging multiplatform data, the perpetual issues in personalized recommender systems—data sparsity and cold-start—can be largely alleviated. To this end, we propose a hybrid solution, STAR (short for “croSs-plaTform App Recommendation”) that integrates both numerical ratings and textual content from multiple platforms. In STAR, we innovatively represent an App as an aggregation of common features across platforms (e.g., App’s functionalities) and specific features that are dependent on the resided platform. In light of this, STAR can discriminate a user’s preference on an App by separating the user’s interest into two parts (either in the App’s inherent factors or platform-aware features). To evaluate our proposal, we construct two real-world datasets that are crawled from the App stores of iPhone, iPad, and iMac. Through extensive experiments, we show that our STAR method consistently outperforms highly competitive recommendation methods, justifying the rationality of our cross-platform App recommendation proposal and the effectiveness of our solution.
Da Cao, Xiangnan He 0001, Liqiang Nie, Xiaochi Wei, Xia Ben Hu, Shunxiang Wu, Tat-Seng Chua
ACM Trans. Inf. Syst.4
2016 A novel unsupervised method for new word extraction
Lili Mei, Heyan Huang, Xiaochi Wei, Xianling Mao
Sci. China Inf. Sci.3
2015 Re-Ranking Voting-Based Answers by Discarding User Behavior Biases
Xiaochi Wei, Heyan Huang, Chin-Yew Lin, Xin Xin 0001, Xianling Mao, Shangguang Wang
IJCAI1
2015 Cross-Domain Collaborative Filtering with Review Text
Xin Xin 0001, Zhirun Liu, Chin-Yew Lin, Heyan Huang, Xiaochi Wei, Ping Guo 0002
IJCAI5
2015 When Factorization Meets Heterogeneous Latent Topics: An Interpretable Cross-Site Recommendation Framework
Xin Xin 0001, Chin-Yew Lin, Xiaochi Wei, Heyan Huang
J. Comput. Sci. Technol.3
2014 Tri-Rank: An Authority Ranking Framework in Heterogeneous Academic Networks by Mutual Reinforce
abstract
Recently, authority ranking has received increasing interests in both academia and industry, and it is applicable to many problems such as discovering influential nodes and building recommendation systems. Various graph-based ranking approaches like PageRank have been used to rank authors and papers separately in homogeneous networks. In this paper, we take venue information into consideration and propose a novel graph-based ranking framework, Tri-Rank, to co-rank authors, papers and venues simultaneously in heterogeneous networks. This approach is a flexible framework and it ranks authors, papers and venues iteratively in a mutually reinforcing way to achieve a more synthetic, fair ranking result. We conduct extensive experiments using the data collected from ACM Digital Library. The experimental results show that Tri-Rank is more effective and efficient than the state-of-the-art baselines including PageRank, HITS and Co-Rank in ranking authors. The papers and venues ranked by Tri-Rank also demonstrate that Tri-Rank is rational.
Zhirun Liu, Heyan Huang, Xiaochi Wei, Xianling Mao
ICTAI3
2013 A Unified Generative Model for Characterizing Microblogs' Topics
Kun Zhuang, Heyan Huang, Xin Xin 0001, Xiaochi Wei, Xianxiang Yang, Chong Feng 0001
WAIM4
2013 Distinguishing Social Ties in Recommender Systems by Graph-Based Algorithms
Xiaochi Wei, Heyan Huang, Xin Xin 0001, Xianxiang Yang
WISE (1)1