EDBT 2026 Demo / reviewers in the wild / expert
Yuanjie Lyu
dblp:331/3722
· DBLP profile ↗
10ranked-venue papers
6as first author
10since 2021 · last 2026
0009-0001-2628-334XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based RewardsabstractSmall LLMs often struggle to match the agentic capabilities of large, costly models.While reinforcement learning can help, progress has been limited by two structural bottlenecks: existing open-source agentic training data are narrow in task variety and easily solved; real-world APIs lack diversity and are unstable for large-scale reinforcement learning rollout processes.We address these challenges with SYNTHAGENT, a framework that jointly synthesizes diverse tooluse training data and simulates complete environments.Specifically, a strong teacher model creates novel tasks and tool ecosystems, then rewrites them into intentionally underspecified instructions.This compels agents to actively query users for missing details.When handling synthetic tasks, an LLM-based user simulator provides user-private information, while a mock tool system delivers stable tool responses.For rewards, task-level rubrics are constructed based on required subgoals, user-agent interactions, and forbidden behaviors.Across 14 challenging datasets in math, search, and tool use, models trained on our synthetic data achieve substantial gains, with small models showing performance comparable to some larger baselines in certain domains. 1 Yuanjie Lyu, Chengyu Wang 0014, Tong Xu 0001 |
ACL (1) | 1 |
| 2026 | Towards Effective Long-Video Event Prediction via Multi-level Event Semantics Mining
Yuanjie Lyu, Penggang Qin, Tong Xu 0001 |
MMM (1) | 2 |
| 2026 | TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation FrameworkabstractRetrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models’ (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning. This tradeoff prioritizes accuracy over efficiency. To address this issue, this work proposes TeaRAG, a T oken- e fficient a gentic RAG framework capable of compressing both retrieval content and reasoning steps. (1) First, the retrieved content is compressed by augmenting chunk-based semantic retrieval with a graph retrieval using concise triplets. A knowledge association graph is then built from semantic similarity and co-occurrence. Finally, Personalized PageRank is leveraged to highlight key knowledge within this graph, reducing the number of tokens per retrieval. (2) Besides, to reduce reasoning steps, Iterative Process-aware Direct Preference Optimization (IP-DPO) is proposed. Specifically, our reward function evaluates the knowledge sufficiency by a knowledge matching mechanism, while penalizing excessive reasoning steps. This design can produce high-quality preference-pair datasets, supporting iterative DPO to improve reasoning conciseness. Across six datasets, TeaRAG improves the average Exact Match by \(4\%\) and \(2\%\) while reducing output tokens by \(61\%\) and \(59\%\) on Llama3-8B-Instruct and Qwen2.5-14B-Instruct, respectively. Code is available at https://github.com/Applied-Machine-Learning-Lab/TeaRAG . Chao Zhang 0096, Yuhao Wang 0006, Derong Xu, Yuanjie Lyu, Shuochen Liu, Tong Xu 0001, Xiangyu Zhao 0001, Yan Gao 0017, Yao Hu 0002, Enhong Chen |
ACM Trans. Inf. Syst. | 5 |
| 2025 | Think Wider, Detect Sharper: Reinforced Reference Coverage for Document-Level Self-Contradiction DetectionabstractDetecting self-contradictions within documents is a challenging task for ensuring textual coherence and reliability.While large language models (LLMs) have advanced in many natural language understanding tasks, document-level self-contradiction detection (DSCD) remains insufficiently studied.Recent approaches leveraging Chain-of-Thought (CoT) prompting aim to enhance reasoning and interpretability; however, they only gain marginal improvement and often introduce inconsistencies across repeated responses.We observe that such inconsistency arises from incomplete reasoning chains that fail to include all relevant contradictory sentences consistently.To address this, we propose a two-stage method that combines supervised fine-tuning (SFT) and reinforcement learning (RL) to enhance DSCD performance.In the SFT phase, a teacher model helps the model learn reasoning patterns, while RL further refines its reasoning ability.Our method incorporates a task-specific reward function to expand the model's reasoning scope, boosting both accuracy and consistency.On the Con-traDoc benchmark, our approach significantly boosts Llama 3.1-8B-Instruct's accuracy from 38.5% to 51.1%, and consistency from 59.6% to 76.2%. 1 Yuanjie Lyu, Shuochen Liu, Chao Zhang 0096, Junhui Lv, Tong Xu 0001 |
EMNLP | 2 |
| 2025 | Generating Event-Oriented Attribution for Movies via Two-Stage Prefix-Enhanced Multimodal LLM
Yuanjie Lyu, Tong Xu 0001, Zihan Niu, Jing Ke |
KSEM (4) | 1 |
| 2025 | CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language ModelsabstractRetrieval-augmented generation (RAG) is a technique that enhances the capabilities of large language models (LLMs) by incorporating external knowledge sources. This method addresses common LLM limitations, including outdated information and the tendency to produce inaccurate “hallucinated” content. However, evaluating RAG systems is a challenge. Most benchmarks focus primarily on question-answering applications, neglecting other potential scenarios where RAG could be beneficial. Accordingly, in the experiments, these benchmarks often assess only the LLM components of the RAG pipeline or the retriever in knowledge-intensive scenarios, overlooking the impact of external knowledge base construction and the retrieval component on the entire RAG pipeline in non-knowledge-intensive scenarios. To address these issues, this article constructs a large-scale and more comprehensive benchmark and evaluates all the components of RAG systems in various RAG application scenarios. Specifically, we refer to the CRUD actions that describe interactions between users and knowledge bases and also categorize the range of RAG applications into four distinct types—create, read, update, and delete (CRUD). “Create” refers to scenarios requiring the generation of original, varied content. “Read” involves responding to intricate questions in knowledge-intensive situations. “Update” focuses on revising and rectifying inaccuracies or inconsistencies in pre-existing texts. “Delete” pertains to the task of summarizing extensive texts into more concise forms. For each of these CRUD categories, we have developed different datasets to evaluate the performance of RAG systems. We also analyze the effects of various components of the RAG system, such as the retriever, context length, knowledge base construction, and LLM. Finally, we provide useful insights for optimizing the RAG technology for different scenarios. The source code is available at GitHub: https://github.com/IAAR-Shanghai/CRUD_RAG . Yuanjie Lyu, Simin Niu, Feiyu Xiong, Bo Tang 0018, Wenjin Wang 0003, Hao Wu 0022, Huanyong Liu, Tong Xu 0001, Enhong Chen |
ACM Trans. Inf. Syst. | 1 |
| 2024 | Retrieve-Plan-Generation: An Iterative Planning and Answering Framework for Knowledge-Intensive LLM GenerationabstractDespite the significant progress of large language models (LLMs) in various tasks, they often produce factual errors due to their limited internal knowledge.Retrieval-Augmented Generation (RAG), which enhances LLMs with external knowledge sources, offers a promising solution.However, these methods can be misled by irrelevant paragraphs in retrieved documents.Due to the inherent uncertainty in LLM generation, inputting the entire document may introduce off-topic information, causing the model to deviate from the central topic and affecting the relevance of the generated content.To address these issues, we propose the Retrieve-Plan-Generation (RPG) framework.RPG generates plan tokens to guide subsequent generation in the plan stage.In the answer stage, the model selects relevant fine-grained paragraphs based on the plan and uses them for further answer generation.This plan-answer process is repeated iteratively until completion, enhancing generation relevance by focusing on specific topics.To implement this framework efficiently, we utilize a simple but effective multi-task prompt-tuning method, enabling the existing LLMs to handle both planning and answering.We comprehensively compare RPG with baselines across 5 knowledge-intensive generation tasks, demonstrating the effectiveness of our approach.1 Yuanjie Lyu, Zihan Niu, Zheyong Xie, Chao Zhang 0096, Tong Xu 0001, Yang Wang 0001, Enhong Chen |
EMNLP | 1 |
| 2024 | Enhancing Complex Question Answering via LLM Pseudo-Document and Adaptive Retrieval
Zhi Zheng 0008, Yuanjie Lyu, Tong Xu 0001 |
WISE (1) | 3 |
| 2024 | InteractNet: Social Interaction Recognition for Semantic-rich VideosabstractThe overwhelming surge of online video platforms has raised an urgent need for social interaction recognition techniques. Compared with simple short-term actions, long-term social interactions in semantic-rich videos could reflect more complicated semantics such as character relationships or emotions, which will better support various downstream applications, e.g., story summarization and fine-grained clip retrieval. However, considering the longer duration of social interactions with severe mutual overlap, involving multiple characters, dynamic scenes, and multi-modal cues, among other factors, traditional solutions for short-term action recognition may probably fail in this task. To address these challenges, in this article, we propose a hierarchical graph-based system, named InteractNet, to recognize social interactions in a multi-modal perspective. Specifically, our approach first generates a semantic graph for each sampled frame with integrating multi-modal cues and then learns the node representations as short-term interaction patterns via an adapted GCN module. Along this line, global interaction representations are accumulated through a sub-clip identification module, effectively filtering out irrelevant information and resolving temporal overlaps between interactions. In the end, the association among simultaneous interactions will be captured and modelled by constructing a global-level character-pair graph to predict the final social interactions. Comprehensive experiments on publicly available datasets demonstrate the effectiveness of our approach compared with state-of-the-art baseline methods. Yuanjie Lyu, Penggang Qin, Tong Xu 0001, Chen Zhu 0003, Enhong Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Faithful Abstractive Summarization via Fact-aware Consistency-constrained TransformerabstractAbstractive summarization is a classic task in Natural Language Generation (NLG), which aims to produce a concise summary of the original document. Recently, great efforts have been made on sequence-to-sequence neural networks to generate abstractive sum- maries with a high level of fluency. However, prior arts mainly focus on the optimization of token-level likelihood, while the rich semantic information in documents has been largely ignored. In this way, the summarization results could be vulnerable to hallucinations, i.e., the semantic-level inconsistency between a summary and corresponding original document. To deal with this challenge, in this paper, we propose a novel fact-aware abstractive summarization model, named Entity-Relation Pointer Generator Network (ERPGN). Specially, we attempt to formalize the facts in original document as a factual knowledge graph, and then generate the high-quality summary via directly modeling consistency between summary and the factual knowledge graph. To that end, we first leverage two pointer net- work structures to capture the fact in original documents. Then, to enhance the traditional token-level likelihood loss, we design two extra semantic-level losses to measure the disagreement between a summary and facts from its original document. Extensive experi- ments on public datasets demonstrate that our ERPGN framework could outperform both classic abstractive summarization models and the state-of-the-art fact-aware baseline methods, with significant improvement in terms of faithfulness. Yuanjie Lyu, Chen Zhu 0003, Tong Xu 0001, Zikai Yin, Enhong Chen |
CIKM | 1 |