Yuan Lin 0001

dblp:07/3483-1 · DBLP profile ↗
← Back
31ranked-venue papers
9as first author
9since 2021 · last 2024
0000-0001-7452-5270ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 13 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2024 Bi-preference Learning Heterogeneous Hypergraph Networks for Session-based Recommendation
abstract
Session-based recommendation intends to predict next purchased items based on anonymous behavior sequences. Numerous economic studies have revealed that item price is a key factor influencing user purchase decisions. Unfortunately, existing methods for session-based recommendation only aim at capturing user interest preference, while ignoring user price preference. Actually, there are primarily two challenges preventing us from accessing price preference. First, the price preference is highly associated to various item features (i.e., category and brand), which asks us to mine price preference from heterogeneous information. Second, price preference and interest preference are interdependent and collectively determine user choice, necessitating that we jointly consider both price and interest preference for intent modeling. To handle above challenges, we propose a novel approach Bi-Preference Learning Heterogeneous Hypergraph Networks (BiPNet) for session-based recommendation. Specifically, the customized heterogeneous hypergraph networks with a triple-level convolution are devised to capture user price and interest preference from heterogeneous features of items. Besides, we develop a Bi-Preference Learning schema to explore mutual relations between price and interest preference and collectively learn these two preferences under the multi-task learning architecture. Extensive experiments on multiple public datasets confirm the superiority of BiPNet over competitive baselines. Additional research also supports the notion that the price is crucial for the task.
Xiaokun Zhang 0001, Bo Xu 0009, Fenglong Ma, Chenliang Li 0005, Yuan Lin 0001, Hongfei Lin
ACM Trans. Inf. Syst.5
2022 Patient Condition Change Network for Safe Medication Recommendation
abstract
As an important task of natural language processing, medication recommendation aims to recommend medication combinations according to the electronic health record, which can also be regarded as a multi-label classification task. But patients often have multiple diseases simultaneously, and the model must consider drug-drug interactions (DDI) of medication combinations when recommending medications, making medication recommendation more difficult. There is little existing work to explore the changes in patient conditions. However, these changes may point to future trends in patient conditions that are critical for reducing DDI rates in recommended drug combinations. In this paper, we proposed the Patient Condition Change Network (PCCNet), which models the current core medications of patient by mining the temporal and spatial changes of patient medication order and patient condition vector, and allocates some auxiliary medications as the currently recommended medication combination. The experimental results show that the proposed model greatly reduces the recommended DDI of medications while achieving results no lower than the state-of-the-art results.11The code is available at https://github.com/master032/PCCNet
Ruobing Li, Jian Wang 0021, Hongfei Lin, Yuan Lin 0001, Huiyi Lu
BIBM4
2022 Dynamic intent-aware iterative denoising network for session-based recommendation
Xiaokun Zhang 0001, Hongfei Lin, Bo Xu 0009, Chenliang Li 0005, Yuan Lin 0001, Haifeng Liu 0002, Fenglong Ma
Inf. Process. Manag.5
2022 Perceived individual fairness with a molecular representation for medicine recommendations
Haifeng Liu 0002, Hongfei Lin, Bo Xu 0009, Nan Zhao 0001, Dongzhen Wen, Xiaokun Zhang 0001, Yuan Lin 0001
Knowl. Based Syst.7
2022 Dual constraints and adversarial learning for fair recommenders
Haifeng Liu 0002, Nan Zhao 0001, Xiaokun Zhang 0001, Hongfei Lin, Liang Yang 0003, Bo Xu 0009, Yuan Lin 0001, Wenqi Fan
Knowl. Based Syst.7
2022 Context-aware ranking refinement with attentive semi-supervised autoencoders
Bo Xu 0009, Hongfei Lin, Yuan Lin 0001, Kan Xu
Soft Comput.3
2021 Info-flow Enhanced GANs for Recommender
abstract
Recommendation systems can help users process large amounts of information, and generative adversarial networks (GANs) show great potential in recommendation systems. In this paper, we propose a new GAN model to enhance the information flow within the generator based on the information flow between the original generator and discriminator. Our experimental results indicate that our model reduces the discrepancy between the generator and the discriminator. Both the generator and discriminator yield considerable performance improvements compared to other strong baselines. The improvements by [email protected] and MRR are significant, which can reach 30.98% and 30.17%, respectively.
Yuan Lin 0001, Zhang Xie, Bo Xu 0009, Kan Xu, Hongfei Lin
SIGIR1
2021 Knowledge-enhanced recommendation using item embedding and path attention
Yuan Lin 0001, Bo Xu 0009, Jiaojiao Feng, Hongfei Lin, Kan Xu
Knowl. Based Syst.1
2021 Two-stage supervised ranking for emotion cause extraction
Bo Xu 0009, Hongfei Lin, Yuan Lin 0001, Kan Xu
Knowl. Based Syst.3
2020 Drug Repositioning for SARS-CoV-2 Based on Graph Neural Network
abstract
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is the strain of coronavirus that causes coronavirus disease 2019 (COVID-19), which leads to over 800,000 deaths and is still no specific medicines. Drug repositioning aiming to infer potential drugs for diseases and achieve much attention during the SARS-CoV-2 epidemic. However, find a specific drug of SARS-CoV-2 is still a large challenge that cannot be addressed well with current methods. To overcome this problem, we present a novel drug repositioning framework of heterogeneous graph convolutional networks for SARS-CoV2. The deep2CoV model can effectively search the potential drugs for SARS-CoV-2, which reduce the number of clinical trials and drug development cycles. The experimental results demonstrate the effectiveness and feasibility of our proposed deep2CoV framework.
Haifeng Liu 0002, Hongfei Lin, Chen Shen 0001, Liang Yang 0003, Yuan Lin 0001, Bo Xu 0009, Jian Wang 0021, Yuanyuan Sun 0002
BIBM5
2020 Improving Social Recommendations with Item Relationships
Haifeng Liu 0002, Hongfei Lin, Bo Xu 0009, Liang Yang 0003, Yuan Lin 0001, Yonghe Chu, Wenqi Fan, Nan Zhao 0001
ICONIP (4)5
2020 Integrating social annotations into topic models for personalized document retrieval
Bo Xu 0009, Hongfei Lin, Yuan Lin 0001, Yizhou Guan
Soft Comput.3
2019 A supervised term ranking model for diversity enhanced biomedical information retrieval
abstract
BACKGROUND: The number of biomedical research articles have increased exponentially with the advancement of biomedicine in recent years. These articles have thus brought a great difficulty in obtaining the needed information of researchers. Information retrieval technologies seek to tackle the problem. However, information needs cannot be completely satisfied by directly introducing the existing information retrieval techniques. Therefore, biomedical information retrieval not only focuses on the relevance of search results, but also aims to promote the completeness of the results, which is referred as the diversity-oriented retrieval. RESULTS: We address the diversity-oriented biomedical retrieval task using a supervised term ranking model. The model is learned through a supervised query expansion process for term refinement. Based on the model, the most relevant and diversified terms are selected to enrich the original query. The expanded query is then fed into a second retrieval to improve the relevance and diversity of search results. To this end, we propose three diversity-oriented optimization strategies in our model, including the diversified term labeling strategy, the biomedical resource-based term features and a diversity-oriented group sampling learning method. Experimental results on TREC Genomics collections demonstrate the effectiveness of the proposed model in improving the relevance and the diversity of search results. CONCLUSIONS: The proposed three strategies jointly contribute to the improvement of biomedical retrieval performance. Our model yields more relevant and diversified results than the state-of-the-art baseline models. Moreover, our method provides a general framework for improving biomedical retrieval performance, and can be used as the basis for future work.
Bo Xu 0009, Hongfei Lin, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001, Dongyu Zhang 0001, Jian Wang 0021, Yuan Lin 0001, Fuliang Yin
BMC Bioinform.9
2019 FGFIREM: A feature generation framework based on information retrieval evaluation measures
Yuan Lin 0001, Bo Xu 0009, Hongfei Lin, Kan Xu
Expert Syst. Appl.1
2019 Incorporating query constraints for autoencoder enhanced ranking
Bo Xu 0009, Hongfei Lin, Yuan Lin 0001, Kan Xu
Neurocomputing3
2019 Learning to Refine Expansion Terms for Biomedical Information Retrieval Using Semantic Resources
abstract
With the rapid development of biomedicine, the number of biomedical articles has increased accordingly, which presents a great challenge for biologists trying to keep up with the latest research. Information retrieval seeks to meet this challenge by searching among a large number of articles based on given queries and providing the most relevant ones to fulfill information needs. As an effective information retrieval technique, query expansion has some room for improvement to achieve the desired performance when directly applied for biomedical information retrieval because there exist many domain-related terms both in users' queries and in related articles. To solve this problem, we propose a biomedical query expansion framework based on learning-to-rank methods, in which we refine candidate expansion terms by training term-ranking models to select the most relevant terms. To train the term-ranking models, we first propose a pseudo-relevance feedback method based on MeSH to select candidate expansion terms and then represent the candidate terms as feature vectors by defining both the corpus-based term features and the resource-based term features. Experimental results obtained for TREC genomics datasets show that our method can capture more relevant terms to expand the original query and effectively improve biomedical information retrieval performance.
Bo Xu 0009, Hongfei Lin, Yuan Lin 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 Improve Diversity-oriented Biomedical Information Retrieval using Supervised Query Expansion
Bo Xu 0009, Hongfei Lin, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001, Dongyu Zhang 0001, Jian Wang 0021, Yuan Lin 0001, Fuliang Yin
BIBM9
2018 Improve Biomedical Information Retrieval Using Modified Learning to Rank Methods
abstract
In these years, the number of biomedical articles has increased exponentially, which becomes a problem for biologists to capture all the needed information manually. Information retrieval technologies, as the core of search engines, can deal with the problem automatically, providing users with the needed information. However, it is a great challenge to apply these technologies directly for biomedical retrieval, because of the abundance of domain specific terminologies. To enhance biomedical retrieval, we propose a novel framework based on learning to rank. Learning to rank is a series of state-of-the-art information retrieval techniques, and has been proved effective in many information retrieval tasks. In the proposed framework, we attempt to tackle the problem of the abundance of terminologies by constructing ranking models, which focus on not only retrieving the most relevant documents, but also diversifying the searching results to increase the completeness of the resulting list for a given query. In the model training, we propose two novel document labeling strategies, and combine several traditional retrieval models as learning features. Besides, we also investigate the usefulness of different learning to rank approaches in our framework. Experimental results on TREC Genomics datasets demonstrate the effectiveness of our framework for biomedical information retrieval.
Bo Xu 0009, Hongfei Lin, Yuan Lin 0001, Liang Yang 0003, Jian Wang 0021
IEEE ACM Trans. Comput. Biol. Bioinform.3
2017 Learning to Rank with Query-level Semi-supervised Autoencoders
abstract
Learning to rank utilizes machine learning methods to solve ranking problems by constructing ranking models in a supervised way, which needs fixed-length feature vectors of documents as inputs, and outputs the ranking models learned by iteratively reducing the pre-defined ranking loss. The document features are always extracted based on classic textual statistics, and different features contribute differently to ranking performance. Given that well-defined features would contribute more to the retrieval performance, we investigate the usage of autoencoders to enrich the feature representations of documents. Autoencoders, as basic building blocks of deep neural networks, have been successfully used in many text mining tasks for generating effective features. To enrich the feature space for learning to rank, we introduce supervision into the loss functions of autoencoders. Specifically, we first train a linear ranking model on the training data, and then incorporate the learned weights into the reconstruction costs of an autoencoder. Meanwhile, we accumulate the costs of documents for a given query with query-level constraints for producing more useful features. We evaluate the effectiveness of our model on three LETOR datasets, and show that our model can generate effective document features to improve the retrieval performance.
Bo Xu 0009, Hongfei Lin, Yuan Lin 0001, Kan Xu
CIKM3
2016 Combining local and global information for product feature extraction in opinion documents
Liang Yang 0003, Bing Liu 0001, Hongfei Lin, Yuan Lin 0001
Inf. Process. Lett.4
2016 Assessment of learning to rank methods for query expansion
abstract
Pseudo relevance feedback, as an effective query expansion method, can significantly improve information retrieval performance. However, the method may negatively impact the retrieval performance when some irrelevant terms are used in the expanded query. Therefore, it is necessary to refine the expansion terms. Learning to rank methods have proven effective in information retrieval to solve ranking problems by ranking the most relevant documents at the top of the returned list, but few attempts have been made to employ learning to rank methods for term refinement in pseudo relevance feedback. This article proposes a novel framework to explore the feasibility of using learning to rank to optimize pseudo relevance feedback by means of reranking the candidate expansion terms. We investigate some learning approaches to choose the candidate terms and introduce some state‐of‐the‐art learning to rank methods to refine the expansion terms. In addition, we propose two term labeling strategies and examine the usefulness of various term features to optimize the framework. Experimental results with three TREC collections show that our framework can effectively improve retrieval performance.
Bo Xu 0009, Hongfei Lin, Yuan Lin 0001
J. Assoc. Inf. Sci. Technol.3
2015 Learning to rank for biomedical information retrieval
abstract
Research articles in biomedicine domain have increased exponentially, which makes it more and more difficult for biologists to manually capture all the information they need. Information retrieval technologies can help to obtain the users' needed information automatically. However, it is a great challenge to apply these technologies to biomedicine domain directly because of some domain specific characteristics, such as the abundance of terminologies. To enhance the effectiveness of the biomedical information retrieval, we propose a novel framework based on the state-of-the-art information retrieval methods, called learning to rank, which has been proved effective to rank documents based on their relevance degree. In the framework, we attempt to tackle the problem of the abundance of terminologies by constructing ranking models, which focus on not only retrieving the most relevant documents but also diversifying the searching results to increase the completeness of the resulting list for a given query. In the model training, we propose two novel document labeling strategies, and combine several traditional retrieval models as learning features. Besides, we also investigate the usefulness of different learning to rank approaches in our framework. Experimental results on TREC Genomics datasets demonstrate our proposed framework is effective in improving the performance of biomedical information retrieval.
Bo Xu 0009, Hongfei Lin, Yuan Lin 0001, Liang Yang 0003, Jian Wang 0021
BIBM3
2015 Group-enhanced ranking
Yuan Lin 0001, Hongfei Lin, Kan Xu, Ajith Abraham, Hongbo Liu 0001
Neurocomputing1
2014 GPQ: Directly Optimizing Q-measure based on Genetic Programming
abstract
Ranking plays an important role in information retrieval system. In recent years, a kind of research named 'learning to rank' becomes more and more popular, which applies machine learning technology to solve ranking problems. Lots of ranking models belonged to learning to rank have been proposed, such as Regression, RankNet, and ListNet. Inspired by this, we proposed a novel learning to rank algorithm named GPQ in this paper, in which genetic programming was employed to directly optimize Q-measure evaluation metric. Experimental results on OHSUMED benchmark dataset indicated that our method GPQ could be competitive with Ranking SVM, SVMMAP and ListNet, and improve the ranking accuracies.
Yuan Lin 0001, Hongfei Lin, Bo Xu 0009
CIKM1
2013 Learning to rank using smoothing methods for language modeling
abstract
The central issue in language model estimation is smoothing, which is a technique for avoiding zero probability estimation problem and overcoming data sparsity. There are three representative smoothing methods: Jelinek‐Mercer (JM) method; Bayesian smoothing using Dirichlet priors (Dir) method; and absolute discounting (Dis) method, whose parameters are usually estimated empirically. Previous research in information retrieval (IR) on smoothing parameter estimation tends to select a single value from optional values for the collection, but it may not be appropriate for all the queries. The effectiveness of all the optional values should be considered to improve the ranking performance. Recently, learning to rank has become an effective approach to optimize the ranking accuracy by merging the existing retrieval methods. In this article, the smoothing methods for language modeling in information retrieval (LMIR) with different parameters are treated as different retrieval methods, then a learning to rank approach to learn a ranking model based on the features extracted by smoothing methods is presented. In the process of learning, the effectiveness of all the optional smoothing parameters is taken into account for all queries. The experimental results on the Learning to Rank for Information Retrieval (LETOR) LETOR3.0 and LETOR4.0 data sets show that our approach is effective in improving the performance of LMIR.
Yuan Lin 0001, Hongfei Lin, Kan Xu, Xiaoling Sun 0002
J. Assoc. Inf. Sci. Technol.1
2011 Learning to rank with cross entropy
abstract
Learning to rank algorithms are usually grouped into three types: the point wise approach, the pairwise approach, and the listwise approach, according to the input spaces. Much of the prior work is based on the three approaches to learn the ranking model to predict the relevance of a document to a query. In this paper, we focus on the problem of constructing new input space based on groups of documents with the same relevance judgment. A novel approach is proposed based on cross entropy to improve the existing ranking method. The experimental results show that our approach leads to significant improvements in retrieval effectiveness.
Yuan Lin 0001, Hongfei Lin, Jiajin Wu, Kan Xu
CIKM1
2011 Selecting related terms in query-logs using two-stage SimRank
abstract
It is commonly believed that query logs from Web search are a gold mine for search business, because they reflect users' preference over Web pages presented by search engines, so a lot of studies based on query logs have been carried out in the last few years. In this study, we assume that two queries are relevant to each other when they have same clicked page in their result lists, and we also consider the queries' topics of user's need. Thus, we propose a Two-Stage SimRank (called TSS in this paper) algorithm based on SimRank and some clustering algorithms to compute the similarity among queries, and then use it to discover relevant terms for query expansion, considering the information of topics and the global relationships of queries concurrently, with a query log collected by a practical search engine. Experimental results on two TREC test collections show that our approach can discover qualified terms effectively and improve retrieval performance.
Hongfei Lin, Yuan Lin 0001
CIKM3
2011 Social annotation in query expansion: a machine learning approach
abstract
Automatic query expansion technologies have been proven to be effective in many information retrieval tasks. Most existing approaches are based on the assumption that the most informative terms in top-retrieved documents can be viewed as context of the query and thus can be used for query expansion. One problem with these approaches is that some of the expansion terms extracted from feedback documents are irrelevant to the query, and thus may hurt the retrieval performance. In social annotations, users provide different keywords describing the respective Web pages from various aspects. These features may be used to boost IR performance. However, to date, the potential of social annotation for this task has been largely unexplored. In this paper, we explore the possibility and potential of social annotation as a new resource for extracting useful expansion terms. In particular, we propose a term ranking approach based on social annotation resource. The proposed approach consists of two phases: (1) in the first phase, we propose a term-dependency method to choose the most likely expansion terms; (2) in the second phase, we develop a machine learning method for term ranking, which is learnt from the statistics of the candidate expansion terms, using ListNet. Experimental results on three TREC test collections show that the retrieval performance can be improved when the term ranking method is used. In addition, we also demonstrate that terms selected by the term-dependency method from social annotation resources are beneficial to improve the retrieval performance.
Yuan Lin 0001, Hongfei Lin
SIGIR1
2011 Learning to rank using query-level regression
abstract
In this paper, we use query-level regression as the loss function. The regression loss function has been used in pointwise methods, however pointwise methods ignore the query boundaries and treat the data equally across queries, and thus the effectiveness is limited. We show that regression is an effective loss function for learning to rank when used in query-level. We use neural network to model the ranking function and gradient descent for optimization and refer our method as ListReg. Experimental results show that ListReg significantly outperforms pointwise Regression and the state-of-the-art listwise method in most cases.
Jiajin Wu, Yuan Lin 0001, Hongfei Lin, Kan Xu
SIGIR3
2010 Ranking SVM for multiple kernels output combination in protein-protein interaction extraction from biomedical literature
abstract
Knowledge about protein-protein interactions unveils the molecular mechanisms of biological processes. This paper presents a multiple kernels learning-based approach to automatically extracting protein-protein interactions from biomedical literature. Experimental evaluations show that our approach can achieve state-of-the-art performance with respect to comparable evaluations, with 64.88% F-score and 88.02% area under the receiver operating characteristics curve (AUC) on the AImed corpus.
Yuan Lin 0001, Jiajin Wu, Nan Tang 0002, Hongfei Lin
BIBM2
2010 Learning to rank with groups
abstract
An essential issue in document retrieval is ranking, and the documents are ranked by their expected relevance to a given query. Multiple labels are used to represent different level of relevance for documents to a given query, and the corresponding label values are used to quantify the relevance of the documents. According to the training set for a given query, the documents can be divided into several groups. Specifically, the documents with the same label are assigned to the same group. If the documents in the group with higher relevance label can always be ranked higher over the ones in groups with lower relevance label by a ranking model, it is reasonable to expect perfect ranking performance. Inspired by this idea, we propose a novel framework for learning to rank, which depends on two new samples. The first one is one-group constituted by one document with higher level label and a group of documents with lower level label; the second one is group-group constituted by a group of documents with higher level label and a group of documents with lower level label. A novel loss function is proposed based on the likelihood loss similar to ListMLE. We demonstrate the advantages of our approaches on the Letor 3.0 data set. Experimental results show that our approaches are effective in improving the ranking performance.
Yuan Lin 0001, Hongfei Lin, Xiaoling Sun 0002
CIKM1