Feifei Dai

dblp:206/9926 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0003-1268-1715ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1Security and privacy · 1
YearPublicationVenuePosition
2026 Incentivizing Retrieval-Augmented Generation via Inner Adaptive Context Selection
abstract
Retrieval-Augmented Generation (RAG) techniques have emerged as a promising direction to merge the non-parametric knowledge into Large Language Models (LLMs), thereby alleviating factual errors, hallucinations and outdated knowledge. Existing RAG methods, which append multiple retrieved documents or passages to the input of LLMs, will inevitably increase the context length, resulting in not only significant computational overhead and inference latency, but also performance degradation. Although reranking or compression modules have been introduced to address these challenges, they overlook the contextual preferences of the generative LLMs itself and may inadvertently discard information that is crucial for generation accuracy. To this end, we introduce InnerRAG, which incentivizes RAG via Inner Adaptive Context Selection. InnerRAG is a novel paradigm that empowers LLMs to autonomously select relevant context during generation. Our proposed InnerRAG endows the model to accurately identify the documents that are most helpful for generation from long contexts. By endowing the model with this capability, InnerRAG facilitates more effective exploitation of external knowledge without being misled by disturbed information, leading to substantial improvements in generation quality while maintaining computational efficiency. Extensive experiments across multiple benchmarks and human evaluations demonstrate that our method consistently outperforms state-of-the-art RAG baselines. Moreover, our framework is orthogonal and complementary to in-context RAG approaches, offering further performance improvements when combined.
Chenxu Cui, Haihui Fan, Sa Zhu, Feifei Dai, Bo Li 0063
SIGIR5
2025 Towards Confidential and Efficient LLM Inference with Dual Privacy Protection
Honglan Yu, Feifei Dai, Haihui Fan, Xiaoyan Gu 0001
DASFAA (5)3
2025 Teach Structure Features to Cooperate with Node Embeddings in Link Prediction
abstract
Structure-enhanced models get a leading performance on the link prediction task as they utilize selected structure features and Graph Neural Network (GNN) based node embeddings simultaneously. However, we observe that when graphs get sparser, these methods perform worse than classical GNN-based methods, which has a severe impact on their practical use. We prove that when the graph gets sparser, the distance between structure features gets smaller. We induce this is the underlying reason for hindering the model from giving a reasonable prediction and leading to performance degeneration. To overcome this problem, we first claim that models need to learn the importance of node embeddings based on the distance between structure features. However, there is a lack of research to efficiently estimate the distance information, and existing models fail to assign the importance properly. Then we design a method called DIP, which satisfies the relation requirement with a weighted term of node embeddings, and use node degree to estimate the distance information to get the weight. Experimental results show that DIP can significantly improve the accuracy of structure-enhanced link prediction models and solve the performance degeneration problem effectively. The code of DIP is publicly available at https://github.com/lzwqbh/DIP.
Feifei Dai, Yucan Zhou, Haihui Fan, Xiaoyan Gu 0001, Dan Meng 0002
IJCNN2
2024 Domain-aware and Co-adaptive Feature Transformation for Domain Adaption Few-shot Relation Extraction
abstract
Few-shot relation extraction (FSRE) can alleviate the data scarcity problem in relation extraction. However, FSRE models often suffer a significant decline in performance when adapting to new domains. To overcome this issue, many researchers have focused on domain adaption FSRE (DAFSRE). Nevertheless, existing approaches primarily concentrate on the source domain, which makes it difficult to accurately transfer useful knowledge to the target domain. Additionally, the lack of distinction between relations further restricts the model performance. In this paper, we propose the domain-aware and co-adaptive feature transformation approach to address these issues. Specifically, we introduce a domain-aware transformation module that leverages the target domain distribution features to guide the domain-aware feature transformations. This can enhance the model’s adaptability across domains, leading to improved target domain performance. Furthermore, we design co-adaptive prototypical networks to perform co-adaptive feature transformation through a transformer mechanism. This results in more robust and distinguishable relation prototypes. Experiments on DAFSRE benchmark datasets demonstrate the effectiveness of our method, which outperforms existing models and achieves state-of-the-art performance.
Yijun Liu 0004, Feifei Dai, Xiaoyan Gu 0001, Minghui Zhai, Bo Li 0063, Meiou Zhang
LREC/COLING2
2024 GL-NER: Generation-Aware Large Language Models for Few-Shot Named Entity Recognition
Xingyu Zhu 0017, Feifei Dai, Xiaoyan Gu 0001, Bo Li 0063, Meiou Zhang, Weiping Wang 0005
ICANN (7)2
2024 Online Caching With Switching Cost and Operational Long-Term Constraints: An Online Learning Approach
abstract
The design of effective online caching policies is an increasingly important problem for content distribution networks, online recommender systems, and edge computing services, etc. Exiting literature usually tackles this problem through the lens of optimistic online learning and aims to achieve sublinear regret. In this paper, we focus on a non-trivial extension of classic online caching problem inspired by operational requirements of real-world systems including switching costs and long-term constraints. To tackle the challenges of switching costs and operational long-term constraints in the online caching, we introduce the Block-structured Follow-the-Regularized-Leader (B-FTRL) caching policy. Our approach incorporates a block structure that divides time into blocks to minimize caching switching costs. The theoretical analysis shows that B-FTRL achieves a utility regret bound of $O\left( {{T^{\frac{{2a - b + 1}}{{1 + a}}}} + {T^{\frac{b}{{1 + a}}}}} \right)$ and switching costs bound of $O\left( {{T^{\frac{1}{{1 + a}}}}} \right)$, where a and b are tunable algorithm parameters. By carefully selecting the values of a and b, we are able to limit the total regret to O(T2/3) while satisfying the operational long-term constraints in expectation. Additionally, we provide high-probability constraint violation bounds of $O\left( {\sqrt T } \right)$. The performance of the proposed algorithm is evaluated with detailed trace-driven numerical tests.
Zifan Jia, Qingsong Liu 0001, Xiaoyan Gu 0001, Haihui Fan, Feifei Dai, Bo Li 0063, Weiping Wang 0005
ICASSP5
2024 FUR-API: Dataset and Baselines Toward Realistic API Anomaly Detection
abstract
The Application Program Interface (API) security is crucial for data security as it ensures the safety and authority of data exchange between different applications. However, the absence of high-quality datasets significantly impedes the development of API anomaly detection. This paper presents a benchmark dataset and baselines for realistic API anomaly detection involving few-shot and unknown-risk scenarios. The dataset is synthesized using an Iterative data Generation approach with Dual-channel Filtering (IGDF). By leveraging large language models and dual-channel filtering models, we can iteratively generate and filter data, yielding a high-quality dataset. Moreover, we have developed baselines and conducted extensive experiments on the proposed dataset. The results indicate that few-shot and unknown-risk API anomaly detection remains a challenging task and still requires further research. All details and resources are released at https://github.com/yijunL/FUR-API.
Yijun Liu 0004, Honglan Yu, Feifei Dai, Xiaoyan Gu 0001, Chenxu Cui, Bo Li 0063, Weiping Wang 0005
ICASSP3
2024 Identifying Misaligned Features for Cross-Domain Cold-Start Recommendation
Mingda Qian, Feifei Dai, Xiaoyan Gu 0001, Haihui Fan, Bo Li 0063
ICONIP (5)2
2024 Fine-Grained Common Knowledge Learning for Domain Adaptive Few-Shot Relation Extraction
Minghui Zhai, Feifei Dai, Xiaoyan Gu 0001, Chuanrong Li, Bo Li 0063, Weiping Wang 0005
ICONIP (6)2
2023 Learning Pair-Centric Representation for Link Sign Prediction with Subgraph
abstract
Signed graphs are prevalent data structures containing both positive and negative links. Recently, the fundamental network analysis task on signed graphs, namely link sign prediction, has received careful attention. Existing methods learn two target node representations independently, and the sign between these two nodes is predicted based on similarity. However, such a paradigm is node-centric that cannot distinguish node pairs with distinct contexts, thus lowering the prediction performance. Learning pair-centric representation is therefore a rewarding way to be aware of differences between pairs. There is no study yet on how to build such an appropriate representation that can effectively infer the sign between the target node pair. In this paper, we provide a new perspective to conduct link sign prediction within the paradigm of subgraph classification and propose a novel Subgraph-based link Sign Prediction (SSP) model. Technically, SSP uses importance-based sampling to extract an informative subgraph around each target node pair. For each subgraph, an innovative node labeling scheme is designed to encode its structural and signed information for representation learning. To further utilize the subgraph representation for imbalanced sign classification, SSP employs self-pruning contrastive learning to gain balanced representations. Extensive experiments on real-world datasets demonstrate that SSP consistently and significantly outperforms all the state-of-the-art baselines.
Jushuo Chen, Feifei Dai, Xiaoyan Gu 0001, Haihui Fan, Bo Li 0063, Weiping Wang 0005
CIKM2
2023 Powering Fine-Tuning: Learning Compatible and Class-Sensitive Representations for Domain Adaption Few-shot Relation Extraction
Yijun Liu 0004, Feifei Dai, Xiaoyan Gu 0001, Haihui Fan, Bo Li 0063, Weiping Wang 0005
DASFAA (4)2
2023 ERPG: Enhancing Entity Representations with Prompt Guidance for Complex Named Entity Recognition
abstract
Recently, sequence generation methods are widely used in complex named entity recognition. By selecting high-related tokens to generate complex named entities, these methods obtain several achievements. However, due to lack of guidance in learning output format and ignoring labels in obtaining features, sequence generation methods suffer invalid output and inaccurate recognition. To solve that, we propose an Enhancing Entity Representation method with Prompt Guidance (ERPG). Specifically, in order to reduce invalid output, we design the candidate entity generation module that generate candidate entities and their labels as expected. Besides, to accurately recognize candidate entities, we propose candidate entity refine module, which obtain distinguishable candidate entity representations and filter them accurately. Based on that, our method finally outperforms baselines by 1.20, 1.62 and 0.69 F1 scores in ACE2004, GENIA and CADEC corpora, which proves the effectiveness in complex named entity recognition.
Xingyu Zhu 0017, Feifei Dai, Xiaoyan Gu 0001, Haihui Fan, Bo Li 0063, Weiping Wang 0005
ICME2
2023 Learning Discriminative Semantic and Multi-view Context for Domain Adaptive Few-Shot Relation Extraction
Minghui Zhai, Feifei Dai, Xiaoyan Gu 0001, Haihui Fan, Bo Li 0063
ICONIP (15)2
2023 Universal Domain Adaptive Network Embedding for Node Classification
abstract
Cross-network node classification aims to leverage the abundant knowledge from a labeled source network to help classify the node in an unlabeled target network. However, existing methods assume that label sets are identical across domains, which is easily violated in practice. Hence, we attempt to integrate network embedding with universal domain adaptation, which transfers valuable knowledge across domains without assumption on the label sets, to assist in node classification. Nonetheless, the complex network relationships between nodes increase the difficulty of this universal domain adaptive node classification task. In this work, we propose a novel Universal Domain Adaptive Network Embedding (UDANE) framework, which learns transferable node representations across networks to succeed in such a task. Technically, we first adopt the cross-network node embedding component to model comprehensive node information of both networks. Then we employ the inter-domain adaptive alignment component to exploit and relate knowledge across domains, learning domain-invariant representation for knowledge transfer. In addition, the intra-domain contrastive alignment component is proposed to learn discriminative representations beneficial for classification by sufficiently utilizing unlabeled data in the target domain. Extensive experiments have been conducted on real-world datasets, demonstrating that the proposed UDANE model outperforms the state-of-the-art baselines by a large margin.
Jushuo Chen, Feifei Dai, Xiaoyan Gu 0001, Bo Li 0063, Weiping Wang 0005
ACM Multimedia2
2022 Flexible Order Aware Sequential Recommendation
abstract
Sequential recommendations can dynamically model user interests, which has great value since users' interests may change rapidly with time. Traditional sequential recommendation methods assume that the user behaviors are rigidly ordered and sequentially dependent. However, some user behaviors have flexible orders, meaning the behaviors may occur in any order and are not sequentially dependent. Therefore, traditional methods may capture inaccurate user interests based on wrong dependencies. Motivated by this, several methods identify flexible orders by continuity or similarity. However, these methods fail to comprehensively understand the nature of flexible orders since continuity or similarity do not determine order flexibilities. Therefore, these methods may misidentify flexible orders, leading to inappropriate recommendations. To address these issues, we propose a Flexible Order aware Sequential Recommendation (FOSR) method to identify flexible orders comprehensively. We argue that orders' flexibilities are highly related to the frequencies of item pair co-occurrences. In light of this, FOSR employs a probabilistic based flexible order evaluation module to simulate item pair frequencies and infer accurate order flexibilities. The frequency labeling module extracts labels from the real item pair frequencies to guide the order flexibility measurement. Given the measured order flexibilities, we develop a flexible order aware self-attention module to model dependencies from flexible orders comprehensively and learn dynamic user interests effectively. Extensive experiments on four benchmark datasets show that our model outperforms various state-of-the-art sequential recommendation methods.
Mingda Qian, Xiaoyan Gu 0001, Lingyang Chu, Feifei Dai, Haihui Fan, Bo Li 0063
ICMR4
2021 Combining Meta-path Instances into Layer-Wise Graphs for Recommendation
Mingda Qian, Bo Li 0063, Xiaoyan Gu 0001, Feifei Dai, Weiping Wang 0005
DASFAA (3)5
2021 Attention-Based Multi-view Feature Fusion for Cross-Domain Recommendation
Feifei Dai, Xiaoyan Gu 0001, Bo Li 0063, Mingda Qian, Weiping Wang 0005
ICANN (1)1
2021 Heterogeneous Side Information-based Iterative Guidance Model for Recommendation
abstract
Heterogeneous side information has been widely used in recommender systems to alleviate the data sparsity problem. However, the heterogeneous side information in existing methods provides insufficient guidance for predicting user preferences as its effect is inevitably weakened during utilization. Furthermore, most existing methods cannot effectively utilize the heterogeneous side information to understand users and items. They often neglect the interrelation among various types of heterogeneous side information of a user or an item. As a result, it is difficult for existing methods to comprehensively understand users and items so that the recommender system recommends inappropriate items to users. To overcome the above drawbacks, we propose an interrelation learning-based recommendation method with iterative heterogeneous side information guidance (ILIG). ILIG includes two modules: 1) Iterative Heterogeneous Side Information Guidance Module. It uses heterogeneous side information to iteratively guide the prediction of user preferences, which effectively enhances the effect of the heterogeneous side information. 2) Interrelation Learning-based Portrait Construction Module. It captures the interrelation among various types of heterogeneous side information to comprehensively learn the representations of users and items. To demonstrate the effectiveness of ILIG, we conduct extensive experiments on Movielens-100K, Movielens-1M, and BookCrossing datasets. The experimental results show that ILIG outperforms the state-of-the-art recommender systems.
Feifei Dai, Xiaoyan Gu 0001, Mingda Qian, Bo Li 0063, Weiping Wang 0005
ICMR1
2020 Privacy-Preserving Cloud Establishment and Data Dissemination Scheme for Vehicular Cloud
abstract
Vehicular cloud (VC) extends cloud computing to vehicles participating in vehicular ad hoc networks, aiming to provide computing and storage services at low cost to vehicles, improve traffic efficiency and safety, ensure real-time services, etc. Due to the highly dynamic nature of VC, it is challenging to efficiently form a dynamic VC securely and anonymously or to securely deliver messages to the dynamic VC without potentially violating the privacy of cloud users. In this paper, we present a concrete secure and privacy-preserving communication scheme for VC establishment and data dissemination. Our scheme allows a group of vehicles that are geographically close to each other to form a VC securely, anonymously and dynamically. This allows vehicle resources to be integrated and shared securely. Once a VC is formed, any cloud user may deliver messages to be securely and anonymously processed in the VC.
Lei Zhang 0009, Kim-Kwang Raymond Choo, Yuanfei Zhang, Feifei Dai
IEEE Trans. Dependable Secur. Comput.5
2018 Secure intelligent traffic light control using fog computing
Jian Liu 0007, Jiangtao Li 0003, Lei Zhang 0009, Feifei Dai, Yuanfei Zhang, Jian Shen 0001
Future Gener. Comput. Syst.4