VLDB 2026 Research / reviewers in the wild / expert
Zhaoheng Huang
dblp:331/3371
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 86% Efficient and distributed learning · 11% Question answering and dialogue systems · 3% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 100% | |
| Network and information security
1 paper |
Digital forensics and information hiding · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
hallucination detection |
2.0 | 2 | 2026 | RLSeek: Evidence-Grounded Reasoning for RAG Hallucination Detection · ACL (1) 2026 Evaluating the Factuality of Large Language Models Using Multiple Plug-and-Play Fact Sources · AAAI 2026 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
1.9 | 2 | 2026 | RLSeek: Evidence-Grounded Reasoning for RAG Hallucination Detection · ACL (1) 2026 One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models · AAAI 2025 |
Natural language and speech › Language models and text generation › retrieval-augmented generation
evidence-grounded reasoning |
1.0 | 1 | 2026 | RLSeek: Evidence-Grounded Reasoning for RAG Hallucination Detection · ACL (1) 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
factuality evaluation |
1.0 | 1 | 2026 | Evaluating the Factuality of Large Language Models Using Multiple Plug-and-Play Fact Sources · AAAI 2026 |
Information retrieval
fact-checking |
1.0 | 1 | 2026 | Evaluating the Factuality of Large Language Models Using Multiple Plug-and-Play Fact Sources · AAAI 2026 |
Information retrieval
retrieval-augmented generation |
1.0 | 1 | 2026 | LLM-Generated Text May Harm Your Retrieval! A Robust Detection Strategy for Retrieval-Augmented Generation · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.9 | 1 | 2025 | One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models · AAAI 2025 |
Natural language and speech › Language models and text generation
prompt tuning |
0.9 | 1 | 2025 | One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models · AAAI 2025 |
Information retrieval › ranking › text ranking
document ranking |
0.9 | 1 | 2025 | CAGS: Context-Aware Document Ranking With Contrastive Graph Sampling · IEEE Trans. Knowl. Data Eng. 2025 |
Information retrieval › interactive information retrieval
session search |
0.9 | 1 | 2025 | CAGS: Context-Aware Document Ranking With Contrastive Graph Sampling · IEEE Trans. Knowl. Data Eng. 2025 |
Digital forensics and information hiding › synthetic media detection
machine-generated text detection |
0.9 | 1 | 2025 | Enhancing LLM Text Detection with Retrieved Contexts and Logits Distribution Consistency · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
web search · 2.0retrieval-augmented evaluation · 2.0large language model · 2.0retrieval-augmented generation · 1.9data augmentation · 1.9reinforcement learning · 1.0fine-tuning · 1.0prompt tuning · 0.9logits distribution consistency · 0.9graph neural network · 0.9embedding tuning · 0.9contrastive graph sampling · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating the Factuality of Large Language Models Using Multiple Plug-and-Play Fact SourcesabstractLarge language models (LLMs) often produce factually inaccurate content, or hallucinations, which undermines their reliability. Existing factuality evaluation systems usually rely on a single predefined fact source, making them task-specific and hard to extend. We present UFO, a unified framework for factuality evaluation that supports multiple plug-and-play fact sources. UFO integrates human-written evidence, web search results, and LLM knowledge within a single evaluation pipeline, and allows users to flexibly select, reorder, and even define customized sources. The system is accessible through both a Python interface and a web-based demo, offering interactive claim-level verification and visualization. Experiments show that UFO system achieves moderate consistency with human annotations. Overall, UFO serves as a transparent and extensible platform for benchmarking fact sources, comparing LLMs, and enabling real-world fact-checking applications across diverse domains. Zhaoheng Huang, Yutao Zhu 0001, Ji-Rong Wen, Zhicheng Dou |
AAAI | 1 |
| 2026 | RLSeek: Evidence-Grounded Reasoning for RAG Hallucination DetectionabstractZhaoheng Huang, Dacheng Wen, Yutao Zhu, Xiaoying Lian, Yushi Liang, Kai Hao, Nan Li, Liangjie Zhang, Qi Zhang, Ji-Rong Wen, Zhicheng Dou, Fangzhao Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhaoheng Huang, Dacheng Wen, Yutao Zhu 0001, Xiaoying Lian, Yushi Liang, Kai Hao, Liangjie Zhang, Qi Zhang 0001, Ji-Rong Wen, Zhicheng Dou, Fangzhao Wu |
ACL (1) | 1 |
| 2026 | LLM-Generated Text May Harm Your Retrieval! A Robust Detection Strategy for Retrieval-Augmented GenerationabstractRetrieval-augmented generation (RAG) effectively enhances the accuracy and timeliness of large language models (LLMs) by incorporating external knowledge retrieved from external sources.However, with the increasing prevalence of LLM-generated content, external corpora used by RAG systems may become contaminated with LLM-generated texts.Such contamination compromises the reliability and quality of retrieved results, ultimately leading to a degradation in RAG performance, and raises concerns about the diminishing presence of human texts and the "Spiral of Silence" effect.A natural solution is to incorporate LLM text detectors into the RAG pipeline to filter out LLM-generated texts from the retrieved results.However, their effective use in RAG remains under-explored.In this paper, we explore the usage paradigms of LLM text detectors for RAG and highlight key limitations of off-the-shelf or directly fine-tuned detectors.To this end, we propose a RAG-aware data augmentation strategy that aligns detector training with realistic contamination patterns.Our approach synthesizes training data from both LLM and human texts under diverse generation modes.Experiments show that our method mitigates performance degradation and improves the long-term stability of RAG systems. Zhaoheng Huang, Yutao Zhu 0001, Ji-Rong Wen, Zhicheng Dou |
ACL (1) | 1 |
| 2026 | DA-TSK-PLR-FS: Domain Adaptive Takagi-Sugeno-Kang Fuzzy System via Pseudolabel Refinement for CCTA-Based Vulnerable Coronary Plaques RecognitionabstractArtificial intelligence has shown great promise in noninvasive recognition of vulnerable coronary plaques. However, practical data issues in multicenter studies, such as inconsistent data distribution and insufficient or missing data labels, could significantly affect the recognition accuracy. Unsupervised domain adaptation (UDA) can be introduced to address this challenge, but several limits still remain. First, many existing UDA models are black boxes, hindering healthcare professionals' ability to interpret and trust the model's decision. Second, some methods use pseudolabel to enhance performance, but often overlook the quality assessment of these pseudolabels, potentially leading to negative knowledge transfer. To this end, based on the interpretable Takagi–Sugeno–Kang fuzzy system (TSK-FS), a novel domain adaptive method is proposed to improve model generalizability for vulnerable coronary plaques recognition in multicenter data. First of all, TSK-FS is employed to construct a shared fuzzy feature space for the source domain and the target domain, aiming to better align data distribution. To make full use of the information of unlabeled target domain data and further reduce the negative knowledge transfer, the enhanced pseudolabel learning mechanism is further introduced by combining the graph-based random walking and label filtering. Moreover, Multicenter data of 910 patients with suspected or diagnosed coronary artery disease were collected from three hospitals for experiments. Experimental results demonstrate that the proposed DA-TSK-PLR-FS achieves the promising generalizability across multicenter datasets Yuanpeng Zhang 0001, Wei Zhang 0221, Zhaoheng Huang, Saikit Lam, Shitong Wang 0001, Jing Cai 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2025 | One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language ModelsabstractRetrieval-augmented generation (RAG) is a promising way to improve large language models (LLMs) for generating more factual, accurate, and up-to-date content. Existing methods either optimize prompts to guide LLMs in leveraging retrieved information or directly fine-tune LLMs to adapt to RAG scenarios. Although fine-tuning can yield better performance, it often compromises the LLMs' general generation capabilities by modifying their parameters. This limitation poses challenges in practical applications, especially when LLMs are already deployed, as parameter adjustments may affect their original functionality. To address this, we propose a novel method that involves learning scalable and pluggable virtual tokens for RAG. By maintaining the LLMs' original parameters and fine-tuning only the embeddings of these pluggable tokens, our approach not only enhances LLMs' performance but also preserves their general generation capabilities. Furthermore, we design several training strategies to improve the scalability, flexibility, and generalizability of our method. Comprehensive experiments across 12 question-answering tasks demonstrate the superiority of our approach. Yutao Zhu 0001, Zhaoheng Huang, Zhicheng Dou, Ji-Rong Wen |
AAAI | 2 |
| 2025 | Enhancing LLM Text Detection with Retrieved Contexts and Logits Distribution ConsistencyabstractLarge language models (LLMs) can generate fluent text, raising concerns about misuse in online comments and academic writing, leading to issues like corpus pollution and copyright infringement.Existing LLM text detection methods often rely on features from the logit distribution of the input text.However, the distinction between the LLM-generated and human-written texts may rely on only a few tokens due to the short length or insufficient information in some texts, leading to minimal and hard-to-detect differences in logit distributions.To address this, we propose HALO, an LLM-based detection method that leverages external text corpora to evaluate the difference in the logit distribution of input text under retrieved human-written and LLM-rewritten contexts.HALO also complements basic detection features and can serve as a plug-and-play module to enhance existing detection methods.Extensive experiments on five public datasets with three widely-used source LLMs show that our proposed detection method achieves state-ofthe-art performance in AUROC, both in crossdomain and domain-specific scenarios. Zhaoheng Huang, Yutao Zhu 0001, Ji-Rong Wen, Zhicheng Dou |
EMNLP | 1 |
| 2025 | CAGS: Context-Aware Document Ranking With Contrastive Graph SamplingabstractIn search sessions, a series of interactions in the context has been proven to be advantageous in capturing users’ search intents. Existing studies show that designing pre-training tasks and data augmentation strategies for session search improves the robustness and generalizability of the model. However, such data augmentation strategies only focus on changing the original session structure to learn a better representation. Ignoring information from outside the session, users’ diverse and complex intents cannot be learned well by simply reordering and deleting historical behaviors, proving that such strategies are limited and inadequate. In order to solve the problem of insufficient modeling under complex user intents, we propose exploiting information outside the original session. More specifically, in this paper, we sample queries and documents from the global click-on and follow-up session graph, alter an original session with these samples, and construct a new session that shares a similar user intent with the original one. Specifically, we design four data augmentation strategies based on session graphs in view of both one-hop and multi-hop structures to sample intent-associated query/document nodes. Experiments conducted on three large-scale public datasets demonstrate that our model outperforms the existing ad-hoc and context-aware document ranking models. Zhaoheng Huang, Yutao Zhu 0001, Zhicheng Dou, Ji-Rong Wen |
IEEE Trans. Knowl. Data Eng. | 1 |