Chenxu Cui

dblp:382/2139 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0007-5228-7989ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval-augmented generation
1.922026
Incentivizing Retrieval-Augmented Generation via Inner Adaptive Context Selection · SIGIR 2026
CIRAG: Retrieval-Augmented Language Model with Collective Intelligence · SIGIR 2025
Information retrieval
retrieval models
1.922026
Incentivizing Retrieval-Augmented Generation via Inner Adaptive Context Selection · SIGIR 2026
CIRAG: Retrieval-Augmented Language Model with Collective Intelligence · SIGIR 2025
Information retrieval › question answering
multi-hop question answering
0.912025
CIRAG: Retrieval-Augmented Language Model with Collective Intelligence · SIGIR 2025
Information retrieval
question answering and dialogue systems
0.912025
CIRAG: Retrieval-Augmented Language Model with Collective Intelligence · SIGIR 2025

Methods — techniques the papers use, named apart from their topics

re-ranking · 0.9query expansion · 0.9
YearPublicationVenuePosition
2026 Incentivizing Retrieval-Augmented Generation via Inner Adaptive Context Selection
abstract
Retrieval-Augmented Generation (RAG) techniques have emerged as a promising direction to merge the non-parametric knowledge into Large Language Models (LLMs), thereby alleviating factual errors, hallucinations and outdated knowledge. Existing RAG methods, which append multiple retrieved documents or passages to the input of LLMs, will inevitably increase the context length, resulting in not only significant computational overhead and inference latency, but also performance degradation. Although reranking or compression modules have been introduced to address these challenges, they overlook the contextual preferences of the generative LLMs itself and may inadvertently discard information that is crucial for generation accuracy. To this end, we introduce InnerRAG, which incentivizes RAG via Inner Adaptive Context Selection. InnerRAG is a novel paradigm that empowers LLMs to autonomously select relevant context during generation. Our proposed InnerRAG endows the model to accurately identify the documents that are most helpful for generation from long contexts. By endowing the model with this capability, InnerRAG facilitates more effective exploitation of external knowledge without being misled by disturbed information, leading to substantial improvements in generation quality while maintaining computational efficiency. Extensive experiments across multiple benchmarks and human evaluations demonstrate that our method consistently outperforms state-of-the-art RAG baselines. Moreover, our framework is orthogonal and complementary to in-context RAG approaches, offering further performance improvements when combined.
Chenxu Cui, Haihui Fan, Sa Zhu, Feifei Dai, Bo Li 0063
SIGIR1
2025 Retrieval-Augmented Image Captioning via Synthesized Entity-Aware Knowledge Representations
abstract
Retrieval-Augmented Image Captioning enhances the model's understanding of real-world images by retrieving external knowledge. Existing methods mainly use original captions or isolated entities related to the query image to help generate captions. However, these methods make the model either imitate the caption style or fail to capture the relationship between entities, resulting in a lack of diversity or inaccuracy in the generated captions. To address these issues, we propose SEAR, a novel framework that utilizes external Synthesized Entity-Aware knowledge Representations to improve captioning performance. Specifically, SEAR clusters images based on scene-level and entity-level features, and synthesizes each clustered images into representative images as retrieval indexes, and simultaneously utilizes a large model to extract and supplement structured knowledge graphs from the corresponding cluster captions. Furthermore, we design a knowledge-graph pruner to prune the knowledge graph by retaining the most relevant subgraphs to the query image. By undertaking these steps in an integrated manner, SEAR enables the model to acquire non-redundant and structured information for generating captions and avoid data-related privacy issues. Extensive experiments on MSCOCO, Flickr30k, and NoCaps demonstrate the effectiveness of our method both in-domain and out-of-domain, outperforming existing lightweight RAIC methods and remaining competitive with heavyweight models.
Chenxu Cui, Jinchao Zhang 0002, Haihui Fan, Haotian Jin, Bo Li 0063
CIKM2
2025 CIRAG: Retrieval-Augmented Language Model with Collective Intelligence
abstract
Retrieval-augmented generation (RAG) paradigms can integrate external knowledge to enhance and validate the output of Large Language Models (LLMs) thereby mitigating generative hallucinations and broadening the model's knowledge scope. Despite advancements, existing RAG methods still suffer from uncertainty of prediction during the multi-round retrieval-generation process, and a lack of the ability to balance the adequacy and redundancy of retrieved information. To address these challenges, we propose CIRAG, an approach that combines the RAG process with collective intelligence. Inspired by the crowd of wisdom, CIRAG simulates individual independent decision-making and information aggregation within a crowd. Specifically, CIRAG first enhances retrieval diversity by expanding queries based on extracted entities, then combines frequency-based and semantic-based reranking to form a multi granularity fusion reranking thereby assessing better relevance, and integrate multiple information sources for accurate content generation. By undertaking these steps in an integrated manner, CIRAG enables the model to acquire comprehensive and non-redundant information for generating responses. We conduct extensive experiments with HotPotQA and 2WikiMultihopQA datasets, popular benchmark for retrieval-based, multi-step question-answering. Experimental results show that our approach surpasses existing advanced RAG framework while providing high portability in query expansion as well as strong comprehensiveness exhibited in the collective intelligence.
Chenxu Cui, Haihui Fan, Jinchao Zhang 0002, Bo Li 0063, Weiping Wang 0005
SIGIR1
2024 FUR-API: Dataset and Baselines Toward Realistic API Anomaly Detection
abstract
The Application Program Interface (API) security is crucial for data security as it ensures the safety and authority of data exchange between different applications. However, the absence of high-quality datasets significantly impedes the development of API anomaly detection. This paper presents a benchmark dataset and baselines for realistic API anomaly detection involving few-shot and unknown-risk scenarios. The dataset is synthesized using an Iterative data Generation approach with Dual-channel Filtering (IGDF). By leveraging large language models and dual-channel filtering models, we can iteratively generate and filter data, yielding a high-quality dataset. Moreover, we have developed baselines and conducted extensive experiments on the proposed dataset. The results indicate that few-shot and unknown-risk API anomaly detection remains a challenging task and still requires further research. All details and resources are released at https://github.com/yijunL/FUR-API.
Yijun Liu 0004, Honglan Yu, Feifei Dai, Xiaoyan Gu 0001, Chenxu Cui, Bo Li 0063, Weiping Wang 0005
ICASSP5