Guodong Du 0008

dblp:437/3455 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 70% Vision and language · 23% Language models and text generation · 7%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › language model interpretability
attention head analysis
1.012026
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination · ACL (1) 2026
Machine learning › Trustworthy machine learning
hallucination
1.012026
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination · ACL (1) 2026
Machine learning › Trustworthy machine learning › hallucination
multimodal hallucination mitigation
1.012026
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination · ACL (1) 2026
Computer vision › Vision and language
vision-language model
1.012026
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model
0.312026
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

logit lens · 1.0attention head intervention · 1.0
YearPublicationVenuePosition
2026 Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
abstract
Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal tasks, yet their reliability is persistently undermined by hallucinations-generating text that contradicts visual input.Recent studies often attribute these errors to inadequate visual attention.In this work, we analyze the attention mechanisms via the logit lens, uncovering a distinct anomaly we term Vocabulary Hijacking.We discover that specific visual tokens, defined as Inert Tokens, disproportionately attract attention.Crucially, when their intermediate hidden states are projected into the vocabulary space, they consistently decode to a fixed set of unrelated words (termed Hijacking Anchors) across layers, revealing a rigid semantic collapse.Leveraging this semantic rigidity, we propose Hijacking Anchor-Based Identification (HABI), a robust strategy to accurately localize these Inert Tokens.To quantify the impact of this phenomenon, we introduce the Non-Hijacked Visual Attention Ratio (NHAR), a novel metric designed to identify attention heads that remain resilient to hijacking and are critical for factual accuracy.Building on these insights, we propose Hijacking-Aware Visual Attention Enhancement (HAVAE), a trainingfree intervention that selectively strengthens the focus of these identified heads on salient visual content.Extensive experiments across multiple benchmarks demonstrate that HAVAE significantly mitigates hallucinations with no additional computational overhead, while preserving the model's general capabilities.
Yangneng Chen, Weijun Yao, Xilai Ma, Guodong Du 0008, Wenya Wang 0001
ACL (1)5