VLDB 2026 Research / reviewers in the wild / expert
Chenglong Lu
dblp:395/4214
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0002-8262-5076ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% | |
| Artificial intelligence
1 paper |
Language models and text generation · 61% Vision and language · 39% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › decoding
decoding strategy |
1.0 | 1 | 2026 | VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language Models · AAAI 2026 |
Natural language and speech › Language models and text generation
hallucination mitigation |
1.0 | 1 | 2026 | VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language Models · AAAI 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 1 | 2026 | VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language Models · AAAI 2026 |
Knowledge graphs › knowledge graph alignment
entity alignment |
0.9 | 1 | 2025 | Breaking the Noise Barrier: LLM-Guided Semantic Filtering and Enhancement for Multi-Modal Entity Alignment · EMNLP 2025 |
Knowledge graphs
knowledge graph alignment |
0.9 | 1 | 2025 | Breaking the Noise Barrier: LLM-Guided Semantic Filtering and Enhancement for Multi-Modal Entity Alignment · EMNLP 2025 |
Knowledge graphs
knowledge graph construction |
0.9 | 1 | 2025 | Breaking the Noise Barrier: LLM-Guided Semantic Filtering and Enhancement for Multi-Modal Entity Alignment · EMNLP 2025 |
Knowledge graphs › knowledge graph alignment › entity alignment
multi-modal entity alignment |
0.9 | 1 | 2025 | Breaking the Noise Barrier: LLM-Guided Semantic Filtering and Enhancement for Multi-Modal Entity Alignment · EMNLP 2025 |
Computer vision › Vision and language
cross-modal consistency |
0.3 | 1 | 2026 | VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language Models · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
reward agent · 1.0reinforcement learning · 1.0image confidence constraints · 1.0caption model · 1.0semantic filtering · 0.9large language model · 0.9attribute summarization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs) demonstrate strong capabilities in multimodal understanding, reasoning, and interaction but still face the fundamental limitation of hallucinations, where they generate erroneous or fabricated information. Most existing research induces hallucinations by manually perturbing visual or instruction inputs, then uses output differences or model-generated descriptions as references to mitigate hallucinations and improve responsevisual consistency. However, these methods are constrained by model capabilities and prone to hallucination propagation. We propose Visual Clue Guided Decoding (VCGD), a novel decoding strategy that introduces an auxiliary Caption Model to generate precise visual clues during decoding for guiding model generation. It further incorporates image confidence constraints to critically suppress hallucination propagation during generation, thereby significantly improving content reliability and visual consistency. Specifically, VCGD leverages high-quality visual descriptions to guide MLLMs in correcting perceptual biases while generating answers. Furthermore, we introduce a Reinforcement Learning-based training paradigm for the Caption Model, in which a Reward Agent provides feedback on the quality of visual clues, further enhancing the accuracy of visual information. Extensive experiments across multiple benchmark datasets and state-of-the-art MLLMs demonstrate that VCGD significantly reduces hallucination rates and improves cross-modal consistency. Our method exhibits strong generalizability and scalability, offering an effective decoding enhancement strategy that can be seamlessly integrated into existing multimodal frameworks. Fu Zhang 0001, Chenglong Lu, Jingwei Cheng |
AAAI | 4 |
| 2025 | RRHF-V: Ranking Responses to Mitigate Hallucinations in Multimodal Large Language Models with Human FeedbackabstractMultimodal large language models (MLLMs) demonstrate strong capabilities in multimodal understanding, reasoning, and interaction but still face the fundamental limitation of hallucinations, where they generate erroneous or fabricated information. To mitigate hallucinations, existing methods annotate pair-responses (one non-hallucination vs one hallucination) using manual methods or GPT-4V, and train alignment algorithms to improve the correspondence between images and text. More critically, an image description often involve multiple dimensions (e.g., object attributes, posture, and spatial relationships), making it challenging for the model to comprehensively learn multidimensional information from pair-responses. To this end, in this paper, we propose RRHFV, which is the first using rank-responses (one non-hallucination vs multiple ranking hallucinations) to mitigate multimodal hallucinations. Instead of using pair-responses to train the model, RRHF-V expands the number of hallucinatory responses, so that the responses with different scores in a rank-response enable the model to learn rich semantic information across various dimensions of the image. Further, we propose a scene graph-based approach to automatically construct rank-responses in a cost-effective and automatic manner. We also design a novel training objective based on rank loss and margin loss to balance the differences between hallucinatory responses within a rankresponse, thereby improving the model’s image comprehension. Experiments on two MLLMs of different sizes and four widely used benchmarks demonstrate that RRHF-V is effective in mitigating hallucinations and outperforms the DPO method based on pair-responses. Fu Zhang 0001, Jinghao Lin, Chenglong Lu, Jingwei Cheng |
COLING | 4 |
| 2025 | Breaking the Noise Barrier: LLM-Guided Semantic Filtering and Enhancement for Multi-Modal Entity AlignmentabstractMulti-modal entity alignment (MMEA) aims to identify equivalent entities between two multimodal knowledge graphs (MMKGs).Existing methods have made substantial advancements in enhancing multi-modal fusion.However, the intrinsic noise within modalities, such as the inconsistency in visual modality and redundant attributes, has not been thoroughly investigated.Excessive noise not only weakens semantic representation but also increases the risk of overfitting in attention-based fusion methods.To address this, we propose LGEA (LLM-Guided Entity Alignment), a novel LLM-guided MMEA framework that prioritizes noise reduction before fusion.Specifically, LGEA introduces two key strategies: (1) fine-grained visual filtering to remove irrelevant images at the semantic level, and (2) contextual summarization of attribute information to enhance entity semantics.To our knowledge, we are the first work to apply LLMs for both visual filtering and attribute-level semantic enhancement in MMEA.Experiments on multiple benchmarks, including the noisy FBYG dataset, show that LGEA sets a new state-of-the-art (SOTA) in robust multi-modal alignment, highlighting the potential of noiseaware strategies as a promising direction for future MMEA research 1 . Chenglong Lu, Chenxiao Li, Jingwei Cheng, Yongquan Ji, Fu Zhang 0001 |
EMNLP | 1 |
| 2025 | Reinforcement Learning-based Token Pruning in Vision Transformers: A Markov Game ApproachabstractVision Transformers (ViTs) have computational costs scaling quadratically with the number of tokens, calling for effective token pruning policies. Most existing policies are handcrafted, lacking adaptivity to varying inputs. Moreover, they fail to consider the sequential nature of token pruning across multiple layers. In this work, for the first time (as far as we know), we exploit Reinforcement Learning (RL) to data-adaptively learn a pruning policy. Formulating token pruning as a sequential decision-making problem, we model it as a Markov Game and utilize Multi-Agent Proximal Policy Optimization (MAPPO) where each agent makes an individualized pruning decision for a single token. We also develop reward functions that enable simultaneous collaboration and competition of these agents to balance efficiency and accuracy. On the well-known ImageNet-1k dataset, our method improves the inference speed by up to 44% while incurring only a negligible accuracy drop of 0.4%. The source code is available at https://github.com/daashuai/rl4evit. Chenglong Lu |
ICME | 1 |