VLDB 2026 Research / reviewers in the wild / expert
Cehao Yang
dblp:380/7640
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0003-9196-7072ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Knowledge graphs · 57% Information retrieval · 43% | |
| Artificial intelligence
4 papers |
Language models and text generation · 58% Knowledge representation and reasoning · 21% Question answering and dialogue systems · 21% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
retrieval-augmented generation |
1.9 | 2 | 2026 | Financial Wind Tunnel: A Retrieval-Augmented Market Simulator · WWW 2026 Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation · ICLR 2025 |
Computational finance and economics › financial modeling
market simulation |
1.0 | 1 | 2026 | Financial Wind Tunnel: A Retrieval-Augmented Market Simulator · WWW 2026 |
Information retrieval › fact-checking
evidence verification |
1.0 | 1 | 2026 | LLM-Oriented Information Retrieval: A Denoising-First Perspective · SIGIR 2026 |
Information retrieval
retrieval-augmented generation |
1.0 | 1 | 2026 | LLM-Oriented Information Retrieval: A Denoising-First Perspective · SIGIR 2026 |
Natural language and speech › Question answering and dialogue systems
knowledge-intensive question answering |
0.9 | 1 | 2025 | Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation · ICLR 2025 |
Knowledge graphs › link prediction
inductive link prediction |
0.9 | 1 | 2025 | Context-aware Inductive Knowledge Graph Completion with Latent Type Constraints and Subgraph Reasoning · AAAI 2025 |
Knowledge graphs
knowledge graph reasoning |
0.9 | 1 | 2025 | Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation · ICLR 2025 |
Knowledge graphs
link prediction |
0.9 | 1 | 2025 | Context-aware Inductive Knowledge Graph Completion with Latent Type Constraints and Subgraph Reasoning · AAAI 2025 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2026 | LLM-Oriented Information Retrieval: A Denoising-First Perspective · SIGIR 2026 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.3 | 1 | 2025 | Context-aware Inductive Knowledge Graph Completion with Latent Type Constraints and Subgraph Reasoning · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
retrieval-augmented generation · 2.0supervised fine-tuning · 1.7subgraph reasoning · 1.7prompt engineering · 1.7knowledge-guided retrieval · 1.7iterative retrieval · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-Oriented Information Retrieval: A Denoising-First PerspectiveabstractModern information retrieval (IR) is no longer consumed primarily by humans but increasingly by large language models (LLMs) via retrieval-augmented generation (RAG) and agentic search. Unlike human users, LLMs are constrained by limited attention budgets and are uniquely vulnerable to noise; misleading or irrelevant information is no longer just a nuisance, but a direct cause of hallucinations and reasoning failures. In this perspective paper, we argue that denoising-maximizing usable evidence density and verifiability within a context window-is becoming the primary bottleneck across the full information access pipeline. We conceptualize this paradigm shift through a four-stage framework of IR challenges: from inaccessible to undiscoverable, to misaligned, and finally to unverifiable. Furthermore, we provide a pipeline-organized taxonomy of signal-to-noise optimization techniques, spanning indexing, retrieval, context engineering, verification, and agentic workflow. We also present research works on information denoising in domains that rely heavily on retrieval such as lifelong assistant, coding agent, deep research, and multimodal understanding. Lu Dai 0001, Fanpu Cao, Ziyang Rao, Cehao Yang, Hao Liu 0026, Hui Xiong 0001 |
SIGIR | 5 |
| 2026 | Financial Wind Tunnel: A Retrieval-Augmented Market Simulator
Bokai Cao, Xueyuan Lin, Yiyan Qi, Chengjin Xu, Cehao Yang, Jian Guo 0016 |
WWW | 5 |
| 2025 | Context-aware Inductive Knowledge Graph Completion with Latent Type Constraints and Subgraph ReasoningabstractInductive knowledge graph completion (KGC) aims to predict missing triples with unseen entities. Recent works focus on modeling reasoning paths between the head and tail entity as direct supporting evidence. However, these methods depend heavily on the existence and quality of reasoning paths, which limits their general applicability in different scenarios. In addition, we observe that latent type constraints and neighboring facts inherent in KGs are also vital in inferring missing triples. To effectively utilize all useful information in KGs, we introduce CATS, a novel context-aware inductive KGC solution. With sufficient guidance from proper prompts and supervised fine-tuning, CATS activates the strong semantic understanding and reasoning capabilities of large language models to assess the existence of query triples, which consist of two modules. First, the type-aware reasoning module evaluates whether the candidate entity matches the latent entity type as required by the query relation. Then, the subgraph reasoning module selects relevant reasoning paths and neighboring facts, and evaluates their correlation to the query triple. Experiment results on three widely used datasets demonstrate that CATS significantly outperforms state-of-the-art methods in 16 out of 18 transductive, inductive, and few-shot settings with an average absolute MRR improvement of 7.2%. Muzhi Li 0001, Cehao Yang, Chengjin Xu, Zixing Song, Xuhui Jiang, Jian Guo 0016, Ho-fung Leung, Irwin King |
AAAI | 2 |
| 2025 | Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented GenerationabstractRetrieval-augmented generation (RAG) has improved large language models (LLMs) by using knowledge retrieval to overcome knowledge deficiencies. However, current RAG methods often fall short of ensuring the depth and completeness of retrieved information, which is necessary for complex reasoning tasks. In this work, we introduce Think-on-Graph 2.0 (ToG-2), a hybrid RAG framework that iteratively retrieves information from both unstructured and structured knowledge sources in a tight-coupling manner. Specifically, ToG-2 leverages knowledge graphs (KGs) to link documents via entities, facilitating deep and knowledge-guided context retrieval. Simultaneously, it utilizes documents as entity contexts to achieve precise and efficient graph retrieval.
ToG-2 alternates between graph retrieval and context retrieval to search for in-depth clues relevant to the question, enabling LLMs to generate answers.
We conduct a series of well-designed experiments to highlight the following advantages of ToG-2: 1) ToG-2 tightly couples the processes of context retrieval and graph retrieval, deepening context retrieval via the KG while enabling reliable graph retrieval based on contexts; 2) it achieves deep and faithful reasoning in LLMs through an iterative knowledge retrieval process of collaboration between contexts and the KG; and 3) ToG-2 is training-free and plug-and-play compatible with various LLMs. Extensive experiments demonstrate that ToG-2 achieves overall state-of-the-art (SOTA) performance on 6 out of 7 knowledge-intensive datasets with GPT-3.5, and can elevate the performance of smaller models (e.g., LLAMA-2-13B) to the level of GPT-3.5’s direct reasoning. The source code is available on https://anonymous.4open.science/r/ToG2. Shengjie Ma, Chengjin Xu, Xuhui Jiang, Muzhi Li 0001, Huaren Qu, Cehao Yang, Jiaxin Mao, Jian Guo 0016 |
ICLR | 6 |
| 2025 | Retrieval, Reasoning, Re-ranking: A Context-Enriched Framework for Knowledge Graph CompletionabstractMuzhi Li, Cehao Yang, Chengjin Xu, Xuhui Jiang, Yiyan Qi, Jian Guo, Ho-fung Leung, Irwin King. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Muzhi Li 0001, Cehao Yang, Chengjin Xu, Xuhui Jiang, Yiyan Qi, Jian Guo 0016, Ho-fung Leung, Irwin King |
NAACL (Long Papers) | 2 |