EDBT 2026 Demo / reviewers in the wild / expert
Jiejun Tan
dblp:355/5723
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0001-8106-4780ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 100% | |
| Artificial intelligence
4 papers |
Language models and text generation · 76% Knowledge representation and reasoning · 18% Reinforcement learning · 7% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
retrieval-augmented generation |
2.5 | 3 | 2025 | HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems · WWW 2025 OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain · EMNLP 2025 Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs · ACL (1) 2024 |
Information retrieval
retrieval-augmented generation |
1.9 | 2 | 2026 | HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches · AAAI 2026 OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain · EMNLP 2025 |
Information retrieval › distributed information retrieval
multi-source retrieval |
1.0 | 1 | 2026 | HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches · AAAI 2026 |
Information retrieval › document processing › document analysis
document representation |
0.9 | 1 | 2025 | HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems · WWW 2025 |
Natural language and speech › Language models and text generation › retrieval-augmented generation
adaptive retrieval |
0.8 | 1 | 2024 | Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs · ACL (1) 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge acquisition |
0.8 | 1 | 2024 | Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs · ACL (1) 2024 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.3 | 1 | 2026 | HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
knowledge refinement · 2.0hierarchical reinforcement learning · 2.0block-tree pruning · 1.7automatic evaluation · 1.7adversarial attack · 1.7HTML cleaning · 1.7proxy model · 0.8heuristic answers · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web SearchesabstractRecently, large reasoning models have demonstrated strong mathematical and coding abilities, and deep search leverages their reasoning capabilities in challenging information retrieval tasks. Existing deep search works are generally limited to a single knowledge source, either local or the Web. However, enterprises often require private deep search systems that can leverage search tools over both local and the Web corpus. Simply training an agent equipped with multiple search tools using flat reinforcement learning (RL) is a straightforward idea, but it has problems such as low training data efficiency and poor mastery of complex tools. To address the above issue, we propose a hierarchical agentic deep search framework, HierSearch, trained with hierarchical RL. At the low level, a local deep search agent and a Web deep search agent are trained to retrieve evidence from their corresponding domains. At the high level, a planner agent coordinates low-level agents and provides the final answer. Moreover, to prevent direct answer copying and error propagation, we design a knowledge refiner that filters out hallucinations and irrelevant evidence returned by low-level agents. Experiments show that HierSearch achieves better performance compared to flat RL, and outperforms various deep search and multi-source retrieval-augmented generation baselines in six benchmarks across general, finance, and medical domains. Jiejun Tan, Zhicheng Dou, Jiehan Cheng, Lifeng Liu, Ji-Rong Wen |
AAAI | 1 |
| 2025 | OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domainabstractretrieval-augmented large language models via transferable adversarial attacks. Shuting Wang 0002, Jiejun Tan, Zhicheng Dou, Ji-Rong Wen |
EMNLP | 2 |
| 2025 | HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG SystemsabstractRetrieval-Augmented Generation (RAG) has been shown to improve knowledge capabilities and alleviate the hallucination problem of LLMs. The Web is a major source of external knowledge used in RAG systems, and many commercial RAG systems have used Web search engines as their major retrieval systems. Typically, such RAG systems retrieve search results, download HTML sources of the results, and then extract plain texts from the HTML sources. Plain text documents or chunks are fed into the LLMs to augment the generation. However, much of the structural and semantic information inherent in HTML, such as headings and table structures, is lost during this plain-text-based RAG process. To alleviate this problem, we propose HtmlRAG, which uses HTML instead of plain text as the format of retrieved knowledge in RAG. We believe HTML is better than plain text in modeling knowledge in external documents, and most LLMs possess robust capacities to understand HTML. However, utilizing HTML presents new challenges. HTML contains additional content such as tags, JavaScript, and CSS specifications, which bring extra input tokens and noise to the RAG system. To address this issue, we propose HTML cleaning, compression, and a two-step block-tree-based pruning strategy, to shorten the HTML while minimizing the loss of information. Experiments on six QA datasets confirm the superiority of using HTML in RAG systems. Our code and datasets are available at https://github.com/plageon/HtmlRAG. Jiejun Tan, Zhicheng Dou, Wen Wang 0016, Weipeng Chen, Ji-Rong Wen |
WWW | 1 |
| 2024 | Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMsabstractThe integration of large language models (LLMs) and search engines represents a significant evolution in knowledge acquisition methodologies.However, determining the knowledge that an LLM already possesses and the knowledge that requires the help of a search engine remains an unresolved issue.Most existing methods solve this problem through the results of preliminary answers or reasoning done by the LLM itself, but this incurs excessively high computational costs.This paper introduces a novel collaborative approach, namely SlimPLM, that detects missing knowledge in LLMs with a slim proxy model, to enhance the LLM's knowledge acquisition process.We employ a proxy model which has far fewer parameters, and take its answers as heuristic answers.Heuristic answers are then utilized to predict the knowledge required to answer the user question, as well as the known and unknown knowledge within the LLM.We only conduct retrieval for the missing knowledge in questions that the LLM does not know.Extensive experimental results on five datasets with two LLMs demonstrate a notable improvement in the end-to-end performance of LLMs in question-answering tasks, achieving or surpassing current state-of-the-art models with lower LLM inference costs. 1 Jiejun Tan, Zhicheng Dou, Yutao Zhu 0001, Peidong Guo, Ji-Rong Wen |
ACL (1) | 1 |