VLDB 2026 Research / reviewers in the wild / expert
Huixuan Zhang
dblp:351/0476
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 48% Trustworthy machine learning · 22% Generative modeling · 15% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › calibration
confidence calibration |
1.0 | 1 | 2026 | DecoCal: Decoding with Calibration in Diffusion Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › decoding
diffusion language model decoding |
1.0 | 1 | 2026 | DecoCal: Decoding with Calibration in Diffusion Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation
hallucination detection |
0.9 | 1 | 2025 | ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs · ACL (1) 2025 |
Natural language and speech › Language models and text generation
large language model safety |
0.9 | 1 | 2025 | DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language Models · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › natural language semantics › semantic interpretation
semantic attachment |
0.9 | 1 | 2025 | R-Bind: Unified Enhancement of Attribute and Relation Binding in Text-to-Image Diffusion Models · EMNLP 2025 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.9 | 1 | 2025 | R-Bind: Unified Enhancement of Attribute and Relation Binding in Text-to-Image Diffusion Models · EMNLP 2025 |
Security and privacy of machine learning
adversarial attack |
0.9 | 1 | 2025 | DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language Models · EMNLP 2025 |
Security and privacy of machine learning › adversarial attack
jailbreak attack |
0.9 | 1 | 2025 | DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language Models · EMNLP 2025 |
Machine learning › Trustworthy machine learning › interpretability › explainable AI › interpretable deep learning
residual stream analysis |
0.3 | 1 | 2025 | ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
monte carlo tree search · 1.7dialogue-aware search · 1.7consistency check · 1.0confidence calibration · 1.0probing · 0.9information contribution metric · 0.9inference-time optimization · 0.9attention map adjustment · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DecoCal: Decoding with Calibration in Diffusion Large Language ModelsabstractDiffusion Large Language Models (DLLMs) generate text via iterative masked-token denoising, supporting parallel prediction and bidirectional context modeling.Despite these advantages, decoding remains challenging: many tokens appear predictable early, yet single-step predictions are often unstable, exhibiting temporal oscillations or overconfidence, making it difficult to determine which tokens can be safely committed.To address these challenges, we propose DecoCal 1 , a Decoding framework that explicitly performs Calibration of tokenlevel confidence across diffusion steps and leverages the calibrated results to guide decoding decisions.Specifically, DecoCal aggregates historical predictions to maintain calibrated confidence, triggering unmasking only when a token is sufficiently stable, while a remasking mechanism allows revision of premature commitments.This calibration-based design enables early decoding of reliably converged tokens while deferring or correcting unstable ones, balancing reliability and speed.Experiments on multiple DLLMs and benchmarks show that DecoCal improves generation accuracy compared to existing strategies.Our results highlight the importance of temporal calibration in unlocking the full potential of diffusion-based language generation.𝐩 !(#) 𝐪 !%& (#) 𝐪 !(#) 𝑥 (') [MASK] 𝑥 (() [MASK] 𝐾𝐿 a. Confidence Check b.Consistency Check Calibration Sequence [MASK] 𝑥 (') 𝑥 (&) 𝑥 (() 𝑥 ()) [MASK] Unmask Remask 𝑥 (') [MASK] 𝑥 (() [MASK] [MASK] 𝑥 (') 𝑥 (&) [MASK] 𝑥 ()) [MASK] Huixuan Zhang, Xiaojun Wan 0001 |
ACL (1) | 2 |
| 2026 | EAMA: Entity-Aware Multimodal Alignment Based Approach for News Image CaptioningabstractNews image captioning requires model to generate an informative caption rich in entities, with the news image and the associated news article. Current MLLMs still bear limitations in handling entity information in news image captioning tasks. Besides, generating high-quality news image captions requires a tradeoff between sufficiency and conciseness of textual input information. To explore the potential of MLLMs, we propose an Entity-Aware Multimodal Alignment (EAMA) based approach for News Image Captioning. Our approach first aligns the MLLM with two extra alignment tasks: Entity-Aware Sentence Selection task and Entity Selection task, together with News Image Captioning task. The aligned MLLM will utilize the additional entity-related information extracted by itself to supplement the textual input while generating news image captions. Our approach achieves better results than all previous models on two mainstream news image captioning datasets. Junzhe Zhang 0004, Huixuan Zhang, Xunjian Yin, Xiaojun Wan 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMsabstractLarge language models (LLMs) excel at various natural language processing tasks, but their tendency to generate hallucinations undermines their reliability. Existing hallucination detection methods leveraging hidden states predominantly focus on static and isolated representations, overlooking their dynamic evolution across layers, which limits efficacy. To address this limitation, we shift the focus to the hidden state update process and introduce a novel metric, the ICR Score (Information Contribution to Residual Stream), which quantifies the contribution of modules to the hidden states’ update. We empirically validate that the ICR Score is effective and reliable in distinguishing hallucinations. Building on these insights, we propose a hallucination detection method, the ICR Probe, which captures the cross-layer evolution of hidden states. Experimental results show that the ICR Probe achieves superior performance with significantly fewer parameters. Furthermore, ablation studies and case analyses offer deeper insights into the underlying mechanism of this method, improving its interpretability. Zhenliang Zhang 0003, Xinyu Hu 0001, Huixuan Zhang, Junzhe Zhang 0004, Xiaojun Wan 0001 |
ACL (1) | 3 |
| 2025 | Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
Zhenliang Zhang 0003, Junzhe Zhang 0004, Xinyu Hu 0001, Huixuan Zhang, Xiaojun Wan 0001 |
CIKM | 4 |
| 2025 | C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination EvaluationabstractDespite the rapid advancement of large language models, they remain highly susceptible to generating hallucinations, which significantly hinders their widespread application. Hallucination research requires dynamic and fine-grained evaluation. However, most existing hallucination benchmarks (especially in Chinese language) rely on human annotations, making automatical and cost-effective hallucination evaluation challenging. To address this, we introduce HaluAgent, an agentic framework that automatically constructs fine-grained question-answering (QA) dataset based on some knowledge documents. Our experiments demonstrate that the manually designed rules and prompt optimization can improve the quality of generated data. Using HaluAgent, we construct C-FAITH, a Chinese QA hallucination benchmark created from 1,399 knowledge documents obtained from web scraping, totaling 60,702 entries. We comprehensively evaluate 16 mainstream LLMs with our proposed C-FAITH, providing detailed experimental results and analysis. Xu Zhang 0077, Zhifei Liu, Huixuan Zhang, Junzhe Zhang 0004, Xiaojun Wan 0001 |
CIKM | 4 |
| 2025 | R-Bind: Unified Enhancement of Attribute and Relation Binding in Text-to-Image Diffusion ModelsabstractText-to-image models frequently fail to achieve perfect alignment with textual prompts, particularly in maintaining proper semantic binding between semantic elements in the given prompt. Existing approaches typically require costly retraining or focus on only correctly generating the attributes of entities (entity-attribute binding), ignoring the cruciality of correctly generating the relations between entities (entity-relation-entity binding), resulting in unsatisfactory semantic binding performance. In this work, we propose a novel training-free method R-Bind that simultaneously improves both entity-attribute and entity-relation-entity binding. Our method introduces three inference-time optimization losses that adjust attention maps during generation. Comprehensive evaluations across multiple datasets demonstrate our approach’s effectiveness, validity, and flexibility in enhancing semantic binding without additional training. Huixuan Zhang, Xiaojun Wan 0001 |
EMNLP | 1 |
| 2025 | DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language ModelsabstractWhile large language models (LLMs) demonstrate remarkable capabilities across a wide range of tasks, they remain vulnerable to generating outputs that are potentially harmful.Red teaming, which involves crafting adversarial inputs to expose vulnerabilities, is a widely adopted approach for evaluating the robustness of these models.Prior studies have indicated that LLMs are susceptible to vulnerabilities exposed through multi-turn interactions as opposed to single-turn scenarios.Nevertheless, existing methods for multi-turn attacks mainly utilize a predefined dialogue pattern, limiting their effectiveness in realistic situations.Effective attacks require adaptive dialogue strategies that respond dynamically to the initial user prompt and the evolving context of the conversation.To address these limitations, we propose DAMON, a novel multi-turn jailbreak attack method.DAMON leverages Monte Carlo Tree Search (MCTS) to systematically explore multiturn conversational spaces, efficiently identifying sub-instruction sequences that induce harmful responses.We evaluate DAMON's efficacy across five LLMs and three datasets.Our experimental results show that DAMON can effectively induce undesired behaviors. Xu Zhang 0077, Xunjian Yin, Dinghao Jing, Huixuan Zhang, Xinyu Hu 0001, Xiaojun Wan 0001 |
EMNLP | 4 |
| 2025 | How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion ModelsabstractWith the rapid development of text-to-vision generation diffusion models, classifier-free guidance has emerged as the most prevalent method for conditioning. However, this approach inherently requires twice as many steps for model forwarding compared to unconditional generation, resulting in significantly higher costs. While previous study has introduced the concept of adaptive guidance, it lacks solid analysis and empirical results, making previous method unable to be applied to general diffusion models. In this work, we present another perspective of applying adaptive guidance and propose Step AG, which is a simple, universally applicable adaptive guidance strategy. Our evaluations focus on both image quality and image-text alignment. whose results indicate that restricting classifier-free guidance to the first several denoising steps is sufficient for generating high-quality, well-conditioned images, achieving an average speedup of 20% to 30%. Such improvement is consistent across different settings such as inference steps, and various models including video generation models, highlighting the superiority of our method. Huixuan Zhang, Xiaojun Wan 0001 |
MMAsia | 1 |
| 2024 | Image Matters: A New Dataset and Empirical Study for Multimodal Hyperbole DetectionabstractHyperbole, or exaggeration, is a common linguistic phenomenon. The detection of hyperbole is an important part of understanding human expression. There have been several studies on hyperbole detection, but most of which focus on text modality only. However, with the development of social media, people can create hyperbolic expressions with various modalities, including text, images, videos, etc. In this paper, we focus on multimodal hyperbole detection. We create a multimodal detection dataset from Weibo (a Chinese social media) and carry out some studies on it. We treat the text and image from a piece of weibo as two modalities and explore the role of text and image for hyperbole detection. Different pre-trained multimodal encoders are also evaluated on this downstream task to show their performance. Besides, since this dataset is constructed from five different keywords, we also evaluate the cross-domain performance of different models. These studies can serve as a benchmark and point out the direction of further study on multimodal hyperbole detection. Huixuan Zhang, Xiaojun Wan 0001 |
LREC/COLING | 1 |