VLDB 2026 Research / reviewers in the wild / expert
Yue Zhang 0096
dblp:47/722-96
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0007-1441-3163ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Vision and language · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
multimodal evaluation |
0.9 | 1 | 2025 | Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization · AAAI 2025 |
Computer vision › Vision and language › visual reasoning
visual entailment |
0.9 | 1 | 2025 | Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
reward-driven optimization · 0.9multimodal large language model · 0.9contrastive learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven OptimizationabstractWe introduce a new task called Defeasible Visual Entailment (DVE), where the goal is to allow the modification of the entailment relationship between an image premise and a text hypothesis based on an additional update. While this concept is well-established in Natural Language Inference, it remains unexplored in visual entailment. At a high level, DVE enables models to refine their initial interpretations, leading to improved accuracy and reliability in various applications such as detecting misleading information in images, enhancing visual question answering, and refining decision-making processes in autonomous systems. Existing metrics do not adequately capture the change in the entailment relationship brought by updates. To address this, we propose a novel inference-aware evaluator designed to capture changes in entailment strength induced by updates, using pairwise contrastive learning and categorical information learning. Additionally, we introduce a reward-driven update optimization method to further enhance the quality of updates generated by multimodal models. Experimental results demonstrate the effectiveness of our proposed evaluator and optimization method. Yue Zhang 0096, Liqiang Jing, Vibhav Gogate |
AAAI | 1 |
| 2025 | Can Large Vision-Language Models Understand Multimodal Sarcasm?abstractSarcasm is a complex linguistic phenomenon that involves a disparity between literal and intended meanings, making it challenging for sentiment analysis and other emotion-sensitive tasks. While traditional sarcasm detection methods primarily focus on text, recent approaches have incorporated multimodal information. However, the application of Large Visual Language Models (LVLMs) in Multimodal Sarcasm Analysis (MSA) remains underexplored. In this paper, we evaluate LVLMs in MSA tasks, specifically focusing on Multimodal Sarcasm Detection and Multimodal Sarcasm Explanation. Through comprehensive experiments, we identify key limitations, such as insufficient visual understanding and a lack of conceptual knowledge. To address these issues, we propose a training-free framework that integrates in-depth object extraction and external conceptual knowledge to improve the model's ability to interpret and explain sarcasm in multimodal contexts. The experimental results on multiple models show the effectiveness of our proposed framework. The code is available at https://github.com/cp-cp/LVLM-MSA. Yue Zhang 0096, Liqiang Jing |
CIKM | 2 |
| 2025 | Tutorial Proposal: Hallucinations in Large Language Models and Large Vision-Language ModelsabstractMultimodal Large Language Models (MLLMs), Large Vision-Language Models (LVLMs), and Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, including text-based reasoning and multimodal content generation. However, these models frequently generate hallucinations-factually incorrect or misleading content-that pose significant challenges, particularly in high-stakes domains such as healthcare, law, and finance. This tutorial provides a comprehensive exploration of hallucinations in MLLMs, LVLMs, and LLMs, examining their causes, detection methods, and mitigation strategies. We discuss different types of hallucination evaluation and benchmarking and explore state-of-the-art techniques for hallucination mitigation. Liqiang Jing, Yue Zhang 0096, Xinya Du |
ICMR | 2 |