Rui Li 0044

dblp:96/4282-44 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0002-4595-0881ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Enhancing CTR Prediction through Sequential Recommendation Pre-training: Introducing the SRP4CTR framework
abstract
In sequential recommendation, pre-training from user historical behaviors through self-supervised learning can better comprehend user dynamic preferences, presenting the potential for direct integration with Click-Through Rate (CTR) prediction tasks. Previous methods have integrated pre-trained models into downstream tasks with the sole purpose of extracting semantic information or well-represented user features, which are then incorporated as new features. However, these approaches tend to ignore the additional inference costs and do not consider how to transfer the effective information from the pre-trained models for specific estimated items in CTR prediction. In this paper, we propose a Sequential Recommendation Pre-training framework for CTR prediction (SRP4CTR) to tackle the above problems. Initially, we discuss the impact of introducing pre-trained models on inference costs. Subsequently, we introduced a pre-trained method to encode sequence side information concurrently. During the fine-tuning process, we incorporate a cross-attention block to establish a bridge between estimated items and the pre-trained model at a low cost. Moreover, we develop a querying transformer technique to facilitate the knowledge transfer from the pre-trained model. Offline and online experiments show that our method outperforms previous baseline models.
Ruidong Han, Qianzhong Li, Rui Li 0044, Yurou Zhao, Xiang Li 0067, Wei Lin 0022
CIKM4
2023 Interactive Recommendation System for Meituan Waimai
abstract
As the largest local retail & instant delivery platform in China, Meituan Waimai has deployed a personalized recommender system on server and recommend nearby stores to users through APP homepage. To capture real-time intention of users and flexibly adjust the recommendation results on the homepage, we further add an interactive recommender system. The existing interactive recommender systems in the industry mainly capture intention of users based on their feedback on a specific UI of questions. However, we find that it will undermine use fluency and increase use complexity by rashly inserting a new question UI when users browse the homepage. Therefore, we develop an Embedded Interactive Recommender System (EIRS) that directly infers users' intention according to their click behaviors on the homepage and dynamically inserts a new recommendation result into the homepage1. To demonstrate the effectiveness of EIRS, we conduct systematic online A/B Tests, where click-through & conversion rate of the inserted EIRS result is 132% higher than that of the initial result on the homepage, and the overall gross merchandise volume is effectively enhanced by 0.43%.
Rui Li 0044, Fei Jiang 0009, Xiang Li 0067, Wei Lin 0022, Wei Wang 0468
SIGIR3
2022 A Dual-Expert Framework for Event Argument Extraction
abstract
Event argument extraction (EAE) is an important information extraction task, which aims to identify the arguments of an event described in a given text and classify the roles played by them. A key characteristic in realistic EAE data is that the instance numbers of different roles follow an obvious long-tail distribution. However, the training and evaluation paradigms of existing EAE models either prone to neglect the performance on "tail roles'', or change the role instance distribution for model training to an unrealistic uniform distribution. Though some generic methods can alleviate the class imbalance in long-tail datasets, they usually sacrifice the performance of "head classes'' as a trade-off. To address the above issues, we propose to train our model on realistic long-tail EAE datasets, and evaluate the average performance over all roles. Inspired by the Mixture of Experts (MOE), we propose a Routing-Balanced Dual Expert Framework (RBDEF), which divides all roles into "head" and "tail" two scopes and assigns the classifications of head and tail roles to two separate experts. In inference, each encoded instance will be allocated to one of the two experts by a routing mechanism. To reduce routing errors caused by the imbalance of role instances, we design a Balanced Routing Mechanism (BRM), which transfers several head roles to the tail expert to balance the load of routing, and employs a tri-filter routing strategy to reduce the misallocation of the tail expert's instances. To enable an effective learning of tail roles with scarce instances, we devise Target-Specialized Meta Learning (TSML) to train the tail expert. Different from other meta learning algorithms that only search a generic parameter initialization equally applying to infinite tasks, TSML can adaptively adjust its search path to obtain a specialized initialization for the tail expert, thereby expanding the benefits to the learning of tail roles. In experiments, RBDEF significantly outperforms the state-of-the-art EAE models and advanced methods for long-tail data.
Rui Li 0044, Wenlin Zhao, Cheng Yang 0002, Sen Su
SIGIR1
2022 MiDTD: A Simple and Effective Distillation Framework for Distantly Supervised Relation Extraction
abstract
Relation extraction (RE), an important information extraction task, faced the great challenge brought by limited annotation data. To this end, distant supervision was proposed to automatically label RE data, and thus largely increased the number of annotated instances. Unfortunately, lots of noise relation annotations brought by automatic labeling become a new obstacle. Some recent studies have shown that the teacher-student framework of knowledge distillation can alleviate the interference of noise relation annotations via label softening. Nevertheless, we find that they still suffer from two problems: propagation of inaccurate dark knowledge and constraint of a unified distillation temperature . In this article, we propose a simple and effective Multi-instance Dynamic Temperature Distillation (MiDTD) framework, which is model-agnostic and mainly involves two modules: multi-instance target fusion (MiTF) and dynamic temperature regulation (DTR). MiTF combines the teacher’s predictions for multiple sentences with the same entity pair to amend the inaccurate dark knowledge in each student’s target. DTR allocates alterable distillation temperatures to different training instances to enable the softness of most student’s targets to be regulated to a moderate range. In experiments, we construct three concrete MiDTD instantiations with BERT, PCNN, and BiLSTM-based RE models, and the distilled students significantly outperform their teachers and the state-of-the-art (SOTA) methods.
Rui Li 0044, Cheng Yang 0002, Tingwei Li, Sen Su
ACM Trans. Inf. Syst.1
2021 Treasures Outside Contexts: Improving Event Detection via Global Statistics
abstract
Event detection (ED) aims at identifying event instances of specified types in given texts, which has been formalized as a sequence labeling task.As far as we know, existing neural-based ED models make decisions relying on the contextual semantic features of each word in the input text, which we find is easy to get confused by varied contexts in the test stage.To this end, we come up with the idea of introducing a set of statistical features from word-event co-occurrence frequencies in the entire training set to cooperate with the contextual features.Specifically, we propose a Semantic and Statistic-Joint Discriminative Network (S 2 -JDN) consisting of a semantic feature extractor, a statistical feature extractor, and a joint event discriminator.In experiments, S 2 -JDN effectively exceeds ten recent state-ofthe-art (SOTA) baseline methods on ACE2005 and KBP2015 benchmark datasets.Further, we perform extensive experiments to investigate the effectiveness of S 2 -JDN.
Rui Li 0044, Wenlin Zhao, Cheng Yang 0002, Sen Su
EMNLP (1)1
2021 Interactive POS-aware network for aspect-level sentiment classification
Kai Shuang, Mengyu Gu, Rui Li 0044, Jonathan Loo, Sen Su
Neurocomputing3
2021 FGCAN: Filter-based Gated Contextual Attention Network for event detection
Shunyu Yao 0001, Kai Shuang, Rui Li 0044, Sen Su
Knowl. Based Syst.3
2020 Major-Minor Long Short-Term Memory for Word-Level Language Model
abstract
Language model (LM) plays an important role in natural language processing (NLP) systems, such as machine translation, speech recognition, learning token embeddings, natural language generation, and text classification. Recently, the multilayer long short-term memory (LSTM) models have been demonstrated to achieve promising performance on word-level language modeling. For each LSTM layer, larger hidden size usually means more diverse semantic features, which enables the LM to perform better. However, we have observed that when a certain LSTM layer reaches a sufficiently large scale, the promotion of overall effect will slow down, as its hidden size increases. In this article, we analyze that an important factor leading to this phenomenon is the high correlation between the newly extended hidden states and the original hidden states, which hinders diverse feature expression of the LSTM. As a result, when the scale is large enough, simply lengthening the LSTM hidden states will cost tremendous extra parameters but has little effect. We propose a simple yet effective improvement on each LSTM layer consisting of a large-scale Major LSTM and a small-scale Minor LSTM to break the high correlation between the two parts of hidden states, which we call Major-Minor LSTMs (MMLSTMs). In experiments, we demonstrate the LM with MMLSTMs surpasses the existing state-of-the-art model on Penn Treebank (PTB) and WikiText-2 (WT2) data sets and outperforms the baseline by 3.3 points in perplexity on WikiText-103 data set without increasing model parameter counts.
Kai Shuang, Rui Li 0044, Mengyu Gu, Jonathan Loo, Sen Su
IEEE Trans. Neural Networks Learn. Syst.2
2019 AELA-DLSTMs: Attention-Enabled and Location-Aware Double LSTMs for aspect-level sentiment classification
Kai Shuang, Xintao Ren, Rui Li 0044, Jonathan Loo
Neurocomputing4