VLDB 2026 Research / reviewers in the wild / expert
Zihao Yi
dblp:371/4658
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 41% Language models and text generation · 23% Question answering and dialogue systems · 18% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
image captioning |
0.9 | 1 | 2025 | Zero-Shot Image Captioning with Multi-type Entity Representations · AAAI 2025 |
Computer vision › Vision and language › image captioning › low-shot image captioning
zero-shot image captioning |
0.9 | 1 | 2025 | Zero-Shot Image Captioning with Multi-type Entity Representations · AAAI 2025 |
Information retrieval › search engines › semantic search
entity retrieval |
0.9 | 1 | 2025 | Zero-Shot Image Captioning with Multi-type Entity Representations · AAAI 2025 |
Natural language and speech › Question answering and dialogue systems › open-domain dialogue
emotional support conversation |
0.8 | 1 | 2024 | Dynamic Demonstration Retrieval and Cognitive Understanding for Emotional Support Conversation · SIGIR 2024 |
Information retrieval › retrieval augmentation
demonstration retrieval |
0.8 | 1 | 2024 | Dynamic Demonstration Retrieval and Cognitive Understanding for Emotional Support Conversation · SIGIR 2024 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.3 | 1 | 2026 | Attention Basin: Why Contextual Position Matters in Large Language Models · ACL (1) 2026 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.3 | 1 | 2025 | Zero-Shot Image Captioning with Multi-type Entity Representations · AAAI 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.2 | 1 | 2024 | Dynamic Demonstration Retrieval and Cognitive Understanding for Emotional Support Conversation · SIGIR 2024 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 1.7GPT-2 · 1.7CLIP · 1.7retrieval · 1.5persona information · 1.5in-context learning · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Attention Basin: Why Contextual Position Matters in Large Language ModelsabstractZihao Yi, Zhenqing Ling, Delong Zeng, Haohao Luo, Zhe Xu, Wei Liu, Jian Luan, Wanxia Cao, Ying Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zihao Yi, Zhenqing Ling, Delong Zeng, Haohao Luo, Zhe Xu 0009, Wei Liu 0302, Jian Luan 0001, Wanxia Cao, Ying Shen 0001 |
ACL (1) | 1 |
| 2026 | Sequential pattern transformer (SPT): a generative and interpretable framework for predicting disease trajectoriesabstractThe effective integration of artificial intelligence into clinical workflows requires models that go beyond simple prediction to generate comprehensive, explainable, and actionable disease trajectories. Addressing the limitations of opaque deep learning architectures and the noise inherent in electronic health records, we introduce the sequential pattern transformer (SPT), a novel framework that synergizes sequential pattern mining with generative transformer modeling. Using four years of inpatient data from 258,460 type 2 diabetes patients, we applied the PrefixSpan algorithm to distill noisy diagnostic histories into a curated vocabulary of 95,630 statistically validated disease progression patterns. A decoder-only transformer was trained exclusively on these evidence-based sequences to learn the temporal dynamics of disease evolution. This pattern-guided approach shifts the modeling paradigm from classification to probabilistic trajectory generation. The model achieved a robust 85.78% Top-5 accuracy, significantly outperforming a standard LSTM baseline (71.47%). Beyond predictive accuracy, the framework constructs a dynamic Disease Atlas, a branching tree structure that visualizes likely future pathways, augmented by multi-level explainable AI (XAI) including learned clinical clusters, SHAP-based feature attribution, and counterfactual simulations. Crucially, this methodology is domain-agnostic and capable of efficient fine-tuning, making it a transferable solution for adapting to diverse clinical conditions and local hospital settings. SPT thus offers a transparent, robust, and scalable framework for mapping the complex temporal dynamics of disease, bridging the gap between high-performance AI and interpretable clinical application. Mohammad Assadi Shalmani, Masoud Khani, Amirsajjad Taleban, Zihao Yi, Jennifer T. Fink, Christopher E. Weber, Qiang Lu 0005, Jake Luo |
Neural Comput. Appl. | 4 |
| 2025 | Zero-Shot Image Captioning with Multi-type Entity RepresentationsabstractAs data and computational resources continue to expand, incorporating a variety of knowledge during the pre-training phase enhances large models, providing them with strong zero-shot capabilities. Due to the alignment of modal features by visual language models, zero-shot image captioning no longer necessitates pre-training on paired image-text labeled data, enabling accurate text description generation for images not encountered before. While recent research focuses on methods utilizing entity retrieval as anchors to bridge the gap between different modalities, these approaches often fall short of thoroughly analyzing the impact of entity retrieval recall on the zero-shot generation capabilities. To address this issue, we propose MERCap, a zero-shot image captioning method employing Multi-type Entity representation Retrieval. More specifically, we first approximate image representation using the CLIP representation of text and Gaussian noise to address the modality gap. Then, we train a GPT-2 decoder to reconstruct text using entities as hard prompts and CLIP representations as soft prompts. Additionally, we construct a domain-specific entity set, assigning multiple representations to each entity and refining their representation vectors through contrastive learning. During inference, we retrieve entities and input them into the decoder to generate corresponding captions. Extensive experiments validate that our approach is efficient, achieving a new state-of-the-art level in cross-domain captioning and demonstrating strong competitiveness in in-domain captioning compared to existing methods. Delong Zeng, Ying Shen 0001, Man Lin, Zihao Yi, Jiarui Ouyang |
AAAI | 4 |
| 2025 | Intent-driven In-context Learning for Few-shot Dialogue State TrackingabstractDialogue state tracking (DST) plays an essential role in task-oriented dialogue systems. However, user’s input may contain implicit information, posing significant challenges for DST tasks. Additionally, DST data includes complex information, which not only contains a large amount of noise unrelated to the current turn, but also makes constructing DST datasets expensive. To address these challenges, we introduce Intent-driven In-context Learning for Few-shot DST (IDIC-DST). By extracting user’s intent, we propose an Intent-driven Dialogue Information Augmentation module to augment the dialogue information, which can track dialogue states more effectively. Moreover, we mask noisy information from DST data and rewrite user’s input in the Intent-driven Examples Retrieval module, where we retrieve similar examples. We then utilize a pre-trained large language model to update the dialogue state using the augmented dialogue information and examples. Experimental results demonstrate that IDIC-DST achieves state-of-the-art performance in few-shot settings on MultiWOZ 2.1 and MultiWOZ 2.4 datasets. Zihao Yi, Zhe Xu 0009, Ying Shen 0001 |
ICASSP | 1 |
| 2024 | Dynamic Demonstration Retrieval and Cognitive Understanding for Emotional Support ConversationabstractEmotional Support Conversation (ESC) systems are pivotal in providing empathetic interactions, aiding users through negative emotional states by understanding and addressing their unique experiences. In this paper, we tackle two key challenges in ESC: enhancing contextually relevant and empathetic response generation through dynamic demonstration retrieval, and advancing cognitive understanding to grasp implicit mental states comprehensively. We introduce Dynamic Demonstration Retrieval and Cognitive-Aspect Situation Understanding (D2RCU), a novel approach that synergizes these elements to improve the quality of support provided in ESCs. By leveraging in-context learning and persona information, we introduce an innovative retrieval mechanism that selects informative and personalized demonstration pairs. We also propose a cognitive understanding module that utilizes four cognitive relationships from the ATOMIC knowledge source to deepen situational awareness of help-seekers' mental states. Our supportive decoder integrates information from diverse knowledge sources, underpinning response generation that is both empathetic and cognitively aware. The effectiveness of D2RCU is demonstrated through extensive automatic and human evaluations, revealing substantial improvements over numerous state-of-the-art models, with up to 13.79% enhancement in overall performance of ten metrics. Our codes are available for public access to facilitate further research and development. Zhe Xu 0009, Daoyuan Chen, Jiayi Kuang, Zihao Yi, Yaliang Li, Ying Shen 0001 |
SIGIR | 4 |