Jiejun Tan

dblp:355/5723 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0001-8106-4780ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%
Artificial intelligence
4 papers
Language models and text generation · 76% Knowledge representation and reasoning · 18% Reinforcement learning · 7%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
retrieval-augmented generation
2.532025
HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems · WWW 2025
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain · EMNLP 2025
Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs · ACL (1) 2024
Information retrieval
retrieval-augmented generation
1.922026
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches · AAAI 2026
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain · EMNLP 2025
Information retrieval › distributed information retrieval
multi-source retrieval
1.012026
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches · AAAI 2026
Information retrieval › document processing › document analysis
document representation
0.912025
HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems · WWW 2025
Natural language and speech › Language models and text generation › retrieval-augmented generation
adaptive retrieval
0.812024
Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs · ACL (1) 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge acquisition
0.812024
Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs · ACL (1) 2024
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.312026
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches · AAAI 2026

Methods — techniques the papers use, named apart from their topics

knowledge refinement · 2.0hierarchical reinforcement learning · 2.0block-tree pruning · 1.7automatic evaluation · 1.7adversarial attack · 1.7HTML cleaning · 1.7proxy model · 0.8heuristic answers · 0.8
YearPublicationVenuePosition
2026 HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
abstract
Recently, large reasoning models have demonstrated strong mathematical and coding abilities, and deep search leverages their reasoning capabilities in challenging information retrieval tasks. Existing deep search works are generally limited to a single knowledge source, either local or the Web. However, enterprises often require private deep search systems that can leverage search tools over both local and the Web corpus. Simply training an agent equipped with multiple search tools using flat reinforcement learning (RL) is a straightforward idea, but it has problems such as low training data efficiency and poor mastery of complex tools. To address the above issue, we propose a hierarchical agentic deep search framework, HierSearch, trained with hierarchical RL. At the low level, a local deep search agent and a Web deep search agent are trained to retrieve evidence from their corresponding domains. At the high level, a planner agent coordinates low-level agents and provides the final answer. Moreover, to prevent direct answer copying and error propagation, we design a knowledge refiner that filters out hallucinations and irrelevant evidence returned by low-level agents. Experiments show that HierSearch achieves better performance compared to flat RL, and outperforms various deep search and multi-source retrieval-augmented generation baselines in six benchmarks across general, finance, and medical domains.
Jiejun Tan, Zhicheng Dou, Jiehan Cheng, Lifeng Liu, Ji-Rong Wen
AAAI1
2025 OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
abstract
retrieval-augmented large language models via transferable adversarial attacks.
Shuting Wang 0002, Jiejun Tan, Zhicheng Dou, Ji-Rong Wen
EMNLP2
2025 HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems
abstract
Retrieval-Augmented Generation (RAG) has been shown to improve knowledge capabilities and alleviate the hallucination problem of LLMs. The Web is a major source of external knowledge used in RAG systems, and many commercial RAG systems have used Web search engines as their major retrieval systems. Typically, such RAG systems retrieve search results, download HTML sources of the results, and then extract plain texts from the HTML sources. Plain text documents or chunks are fed into the LLMs to augment the generation. However, much of the structural and semantic information inherent in HTML, such as headings and table structures, is lost during this plain-text-based RAG process. To alleviate this problem, we propose HtmlRAG, which uses HTML instead of plain text as the format of retrieved knowledge in RAG. We believe HTML is better than plain text in modeling knowledge in external documents, and most LLMs possess robust capacities to understand HTML. However, utilizing HTML presents new challenges. HTML contains additional content such as tags, JavaScript, and CSS specifications, which bring extra input tokens and noise to the RAG system. To address this issue, we propose HTML cleaning, compression, and a two-step block-tree-based pruning strategy, to shorten the HTML while minimizing the loss of information. Experiments on six QA datasets confirm the superiority of using HTML in RAG systems. Our code and datasets are available at https://github.com/plageon/HtmlRAG.
Jiejun Tan, Zhicheng Dou, Wen Wang 0016, Weipeng Chen, Ji-Rong Wen
WWW1
2024 Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs
abstract
The integration of large language models (LLMs) and search engines represents a significant evolution in knowledge acquisition methodologies.However, determining the knowledge that an LLM already possesses and the knowledge that requires the help of a search engine remains an unresolved issue.Most existing methods solve this problem through the results of preliminary answers or reasoning done by the LLM itself, but this incurs excessively high computational costs.This paper introduces a novel collaborative approach, namely SlimPLM, that detects missing knowledge in LLMs with a slim proxy model, to enhance the LLM's knowledge acquisition process.We employ a proxy model which has far fewer parameters, and take its answers as heuristic answers.Heuristic answers are then utilized to predict the knowledge required to answer the user question, as well as the known and unknown knowledge within the LLM.We only conduct retrieval for the missing knowledge in questions that the LLM does not know.Extensive experimental results on five datasets with two LLMs demonstrate a notable improvement in the end-to-end performance of LLMs in question-answering tasks, achieving or surpassing current state-of-the-art models with lower LLM inference costs. 1
Jiejun Tan, Zhicheng Dou, Yutao Zhu 0001, Peidong Guo, Ji-Rong Wen
ACL (1)1