EDBT 2026 Demo / reviewers in the wild / expert
Chenxu Cui
dblp:382/2139
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0007-5228-7989ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval-augmented generation |
1.9 | 2 | 2026 | Incentivizing Retrieval-Augmented Generation via Inner Adaptive Context Selection · SIGIR 2026 CIRAG: Retrieval-Augmented Language Model with Collective Intelligence · SIGIR 2025 |
Information retrieval
retrieval models |
1.9 | 2 | 2026 | Incentivizing Retrieval-Augmented Generation via Inner Adaptive Context Selection · SIGIR 2026 CIRAG: Retrieval-Augmented Language Model with Collective Intelligence · SIGIR 2025 |
Information retrieval › question answering
multi-hop question answering |
0.9 | 1 | 2025 | CIRAG: Retrieval-Augmented Language Model with Collective Intelligence · SIGIR 2025 |
Information retrieval
question answering and dialogue systems |
0.9 | 1 | 2025 | CIRAG: Retrieval-Augmented Language Model with Collective Intelligence · SIGIR 2025 |
Methods — techniques the papers use, named apart from their topics
re-ranking · 0.9query expansion · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Incentivizing Retrieval-Augmented Generation via Inner Adaptive Context SelectionabstractRetrieval-Augmented Generation (RAG) techniques have emerged as a promising direction to merge the non-parametric knowledge into Large Language Models (LLMs), thereby alleviating factual errors, hallucinations and outdated knowledge. Existing RAG methods, which append multiple retrieved documents or passages to the input of LLMs, will inevitably increase the context length, resulting in not only significant computational overhead and inference latency, but also performance degradation. Although reranking or compression modules have been introduced to address these challenges, they overlook the contextual preferences of the generative LLMs itself and may inadvertently discard information that is crucial for generation accuracy. To this end, we introduce InnerRAG, which incentivizes RAG via Inner Adaptive Context Selection. InnerRAG is a novel paradigm that empowers LLMs to autonomously select relevant context during generation. Our proposed InnerRAG endows the model to accurately identify the documents that are most helpful for generation from long contexts. By endowing the model with this capability, InnerRAG facilitates more effective exploitation of external knowledge without being misled by disturbed information, leading to substantial improvements in generation quality while maintaining computational efficiency. Extensive experiments across multiple benchmarks and human evaluations demonstrate that our method consistently outperforms state-of-the-art RAG baselines. Moreover, our framework is orthogonal and complementary to in-context RAG approaches, offering further performance improvements when combined. Chenxu Cui, Haihui Fan, Sa Zhu, Feifei Dai, Bo Li 0063 |
SIGIR | 1 |
| 2025 | Retrieval-Augmented Image Captioning via Synthesized Entity-Aware Knowledge RepresentationsabstractRetrieval-Augmented Image Captioning enhances the model's understanding of real-world images by retrieving external knowledge. Existing methods mainly use original captions or isolated entities related to the query image to help generate captions. However, these methods make the model either imitate the caption style or fail to capture the relationship between entities, resulting in a lack of diversity or inaccuracy in the generated captions. To address these issues, we propose SEAR, a novel framework that utilizes external Synthesized Entity-Aware knowledge Representations to improve captioning performance. Specifically, SEAR clusters images based on scene-level and entity-level features, and synthesizes each clustered images into representative images as retrieval indexes, and simultaneously utilizes a large model to extract and supplement structured knowledge graphs from the corresponding cluster captions. Furthermore, we design a knowledge-graph pruner to prune the knowledge graph by retaining the most relevant subgraphs to the query image. By undertaking these steps in an integrated manner, SEAR enables the model to acquire non-redundant and structured information for generating captions and avoid data-related privacy issues. Extensive experiments on MSCOCO, Flickr30k, and NoCaps demonstrate the effectiveness of our method both in-domain and out-of-domain, outperforming existing lightweight RAIC methods and remaining competitive with heavyweight models. Chenxu Cui, Jinchao Zhang 0002, Haihui Fan, Haotian Jin, Bo Li 0063 |
CIKM | 2 |
| 2025 | CIRAG: Retrieval-Augmented Language Model with Collective IntelligenceabstractRetrieval-augmented generation (RAG) paradigms can integrate external knowledge to enhance and validate the output of Large Language Models (LLMs) thereby mitigating generative hallucinations and broadening the model's knowledge scope. Despite advancements, existing RAG methods still suffer from uncertainty of prediction during the multi-round retrieval-generation process, and a lack of the ability to balance the adequacy and redundancy of retrieved information. To address these challenges, we propose CIRAG, an approach that combines the RAG process with collective intelligence. Inspired by the crowd of wisdom, CIRAG simulates individual independent decision-making and information aggregation within a crowd. Specifically, CIRAG first enhances retrieval diversity by expanding queries based on extracted entities, then combines frequency-based and semantic-based reranking to form a multi granularity fusion reranking thereby assessing better relevance, and integrate multiple information sources for accurate content generation. By undertaking these steps in an integrated manner, CIRAG enables the model to acquire comprehensive and non-redundant information for generating responses. We conduct extensive experiments with HotPotQA and 2WikiMultihopQA datasets, popular benchmark for retrieval-based, multi-step question-answering. Experimental results show that our approach surpasses existing advanced RAG framework while providing high portability in query expansion as well as strong comprehensiveness exhibited in the collective intelligence. Chenxu Cui, Haihui Fan, Jinchao Zhang 0002, Bo Li 0063, Weiping Wang 0005 |
SIGIR | 1 |
| 2024 | FUR-API: Dataset and Baselines Toward Realistic API Anomaly DetectionabstractThe Application Program Interface (API) security is crucial for data security as it ensures the safety and authority of data exchange between different applications. However, the absence of high-quality datasets significantly impedes the development of API anomaly detection. This paper presents a benchmark dataset and baselines for realistic API anomaly detection involving few-shot and unknown-risk scenarios. The dataset is synthesized using an Iterative data Generation approach with Dual-channel Filtering (IGDF). By leveraging large language models and dual-channel filtering models, we can iteratively generate and filter data, yielding a high-quality dataset. Moreover, we have developed baselines and conducted extensive experiments on the proposed dataset. The results indicate that few-shot and unknown-risk API anomaly detection remains a challenging task and still requires further research. All details and resources are released at https://github.com/yijunL/FUR-API. Yijun Liu 0004, Honglan Yu, Feifei Dai, Xiaoyan Gu 0001, Chenxu Cui, Bo Li 0063, Weiping Wang 0005 |
ICASSP | 5 |