EDBT 2026 Demo / reviewers in the wild / expert
Yunyan Zhang
dblp:58/463
· DBLP profile ↗
21ranked-venue papers
4as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm DetectionabstractMultimodal sarcasm detection (MSD) aims to identify sarcasm polarity from diverse modalities (i.e., image–text pairs), a task that has received increasing attention. While significant progress has been made, existing approaches still face two major issues: lack of explainability and weak generalizability. In this paper, we introduce a new large vision–language model (LVLM) dubbed S³-MSD for explainable and generalizable MSD through three key components. For explainability, we develop (1) a self-training paradigm that automatically bootstraps answers with explanations, and (2) a self-calibrating mechanism that rectifies flawed explanations. For generalizability, we design (3) a self-focusing module that amplifies visual semantic entities through preference optimization, thereby mitigating textual over-reliance. Experimental results on both in-distribution and out-of-distribution (OOD) benchmarks demonstrate that S³-MSD consistently outperforms state-of-the-art methods in detection performance. Furthermore, the proposed S³-MSD provides persuasive explanations, as verified by both quantitative metrics and human evaluations. Zhihong Zhu 0001, Fan Zhang 0111, Yunyan Zhang, Jinghan Sun, Guimin Hu, Hao Wu 0094, Yuyan Chen, Xian Wu 0001 |
AAAI | 3 |
| 2026 | CMID: Towards Medical Visual Question Answering via Contrastive Mutual Information DecodingabstractMedical Visual Question Answering (Med-VQA) aims to generate accurate answers for clinical questions grounded in medical images, which has attracted increasing research attention due to its potential to streamline diagnostics and reduce clinical burden. Recent advances in Large Vision-Language Models (LVLMs) have shown great promise for Med-VQA, but still suffer from two inference-time issues: (1) attention shift, where the LVLM over-relies on textual priors; and (2) attention dispersion, where it fails to focus on critical diagnostic regions. To tackle these issues, we propose Contrastive Mutual Information Decoding (CMID), a training-free inference-time intervention grounded in information theory for Med-VQA. Concretely, CMID first identifies the Principal Focus Area (PFA) from decoder attention maps, then constructs focus-preserving and focus-excluding views to derive dual contrastive signals that simultaneously amplify salient visual cues and suppress background noise. Crucially, these corrective signals are adaptively scaled by a reliability-gated self-correction mechanism, based on the distributional shift induced by the PFA. Extensive experiments on three Med-VQA benchmarks demonstrate the effectiveness of CMID. Further analyses showcase its robust generalizability across diverse medical architectures and tasks. Zhihong Zhu 0001, Yunyan Zhang, Fan Zhang 0111, Xian Wu 0001 |
AAAI | 2 |
| 2025 | CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMabstractLarge Language Models (LLMs) have demonstrated significant potential in medical diagnostics and clinical decision-making.While benchmarks such as MedQA and PubMedQA have advanced the evaluation of qualitative reasoning, existing medical NLP benchmarks still face two limitations: the absence of a Chinese benchmark for medical calculation tasks, and the lack of fine-grained evaluation of intermediate reasoning.In this paper, we introduce CMedCalc-Bench, a new benchmark designed for Chinese medical calculation.CMedCalc-Bench covers 69 calculators across 12 clinical departments, featuring over 1,000 real-world patient cases.Building on this, we design a fine-grained evaluation framework that disentangles clinical entity extraction from numerical computation, enabling systematic diagnosis of model deficiencies.Experiments across four model families, including medicalspecialized and reasoning-focused, provide an assessment of their strengths and limitations on Chinese medical calculation.Furthermore, explorations on faithful reasoning and the demonstration effect offer early insights into advancing safe and reliable clinical computation. † Equal contribution. Yunyan Zhang, Zhihong Zhu 0001, Xian Wu 0001 |
EMNLP | 1 |
| 2024 | Alignment before Awareness: Towards Visual Question Localized-Answering in Robotic Surgery via Optimal Transport and Answer SemanticsabstractThe visual question localized-answering (VQLA) system has garnered increasing attention due to its potential as a knowledgeable assistant in surgical education. Apart from providing text-based answers, VQLA can also pinpoint the specific region of interest for better surgical scene understanding. Although recent Transformer-based models for VQLA have obtained promising results, they (1) conduct vanilla text-to-image cross attention, leading to unidirectional and coarse-grained alignment; (2) ignore exploiting the semantics of answers to further boost performance. In this paper, we propose a novel model termed OTAS, which first introduces optimal transport to achieve bidirectional and fine-grained alignment between images and questions, enabling more precise localization. Besides, OTAS incorporates a set of learnable candidate answer embeddings to query the probability of each answer class for a given image-question pair. Through Transformer attention, the candidate answer embeddings interact with the fused features of the image-question pair to make the answer decision. Extensive experiments on two widely-used benchmark datasets demonstrate the superiority of our model over state-of-the-art methods. Zhihong Zhu 0001, Yunyan Zhang, Xuxin Cheng, Zhiqi Huang 0001, Derong Xu, Xian Wu 0001, Yefeng Zheng 0001 |
LREC/COLING | 2 |
| 2024 | DGLF: A Dual Graph-based Learning Framework for Multi-modal Sarcasm DetectionabstractZhihong Zhu, Kefan Shen, Zhaorun Chen, Yunyan Zhang, Yuyan Chen, Xiaoqi Jiao, Zhongwei Wan, Shaorong Xie, Wei Liu, Xian Wu, Yefeng Zheng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhihong Zhu 0001, Kefan Shen, Zhaorun Chen, Yunyan Zhang, Yuyan Chen, Xiaoqi Jiao, Zhongwei Wan, Shaorong Xie, Wei Liu 0027, Xian Wu 0001, Yefeng Zheng 0001 |
EMNLP | 4 |
| 2024 | TFCD: Towards Multi-modal Sarcasm Detection via Training-Free Counterfactual Debiasing
Zhihong Zhu 0001, Xianwei Zhuang, Yunyan Zhang, Derong Xu, Guimin Hu, Xian Wu 0001, Yefeng Zheng 0001 |
IJCAI | 3 |
| 2024 | Enhancing New Multiple Sclerosis Lesion Segmentation via Self-supervised Pre-training and Synthetic Lesion Integration
Peyman Tahghighi, Yunyan Zhang, Roberto Souza 0001, Amin Komeili |
MICCAI (8) | 2 |
| 2024 | Multivariate Cooperative Game for Image-Report Pairs: Hierarchical Semantic Alignment for Medical Report Generation
Zhihong Zhu 0001, Xuxin Cheng, Yunyan Zhang, Zhaorun Chen, Qingqing Long, Hongxiang Li 0004, Zhiqi Huang 0001, Xian Wu 0001, Yefeng Zheng 0001 |
MICCAI (3) | 3 |
| 2024 | Aspects are Anchors: Towards Multimodal Aspect-based Sentiment Analysis via Aspect-driven Alignment and RefinementabstractGiven coupled sentence image pairs, Multimodal Aspect-based Sentiment Analysis (MABSA) aims to detect aspect terms and predict their sentiment polarity. While existing methods have made great efforts in aligning images and text for improved MABSA performance, they still struggle to effectively mitigate the challenge of the noisy correspondence problem (NCP): the text description is often not well-aligned with the visual content. To alleviate NCP, in this paper, we introduce Aspect-driven Alignment and Refinement (ADAR), which is a two-stage coarse-to-fine alignment framework. In the first stage, ADAR devises a novel Coarse-to-fine Aspect-driven Alignment Module, which introduces Optimal Transport (OT) to learn the coarse-grained alignment between visual and textual features. Then the adaptive filter bin is applied to remove the irrelevant image regions at a fine-grained level; In the second stage, ADAR introduces an Aspect-driven Refinement Module to further refine the cross-modality feature representation. Extensive experiments on two benchmark datasets demonstrate the superiority of our model over state-of-the-art performance in the MABSA task. Zhanpeng Chen, Zhihong Zhu 0001, Wanshi Xu, Yunyan Zhang, Xian Wu 0001, Yefeng Zheng 0001 |
ACM Multimedia | 4 |
| 2024 | InMu-Net: Advancing Multi-modal Intent Detection via Information Bottleneck and Multi-sensory ProcessingabstractMulti-modal intent detection (MID) aims to comprehend users' intentions through diverse modalities, which has received widespread attention in dialogue systems. Despite the promising advancements in complex fusion mechanisms or architecture designs, challenges remain due to: (1) various noise and redundancy in both visual and audio modalities and (2) long-tailed distributions of intent categories. In this paper, to tackle the above two issues, we propose InMu-Net, a simple yet effective framework for MID from the Information bottleneck and Multi-sensory processing perspective. Our contributions lie in three aspects. First, we devise a denoising bottleneck module to filter out the intent-irrelevant information in the fused feature; Second, we introduce a saliency preservation loss to prevent the dropping of intent-relevant information; Ultimately, kurtosis regulation is introduced to maintain representation smoothness during the filtering process, mitigating the adverse impact of the long tail distribution. Comprehensive experiments on two MID benchmark datasets demonstrate the effectiveness of InMu-Net and its vital components. Impressively, a series of analyses reveal our denoising potential and robustness in low-resource, modality corruption, cross-architecture and cross-task scenarios. Zhihong Zhu 0001, Xuxin Cheng, Zhaorun Chen, Yuyan Chen, Yunyan Zhang, Xian Wu 0001, Yefeng Zheng 0001 |
ACM Multimedia | 5 |
| 2024 | MedJourney: Benchmark and Evaluation of Large Language Models over Patient Clinical JourneyabstractLarge language models (LLMs) have demonstrated remarkable capabilities in language understanding and generation, leading to their widespread adoption across various fields. Among these, the medical field is particularly well-suited for LLM applications, as many medical tasks can be enhanced by LLMs. Despite the existence of benchmarks for evaluating LLMs in medical question-answering and exams, there remains a notable gap in assessing LLMs' performance in supporting patients throughout their entire hospital visit journey in real-world clinical practice. In this paper, we address this gap by dividing a typical patient's clinical journey into four stages: planning, access, delivery and ongoing care. For each stage, we introduce multiple tasks and corresponding datasets, resulting in a comprehensive benchmark comprising 12 datasets, of which five are newly introduced, and seven are constructed from existing datasets. This proposed benchmark facilitates a thorough evaluation of LLMs' effectiveness across the entire patient journey, providing insights into their practical application in clinical settings. Additionally, we evaluate three categories of LLMs against this benchmark: 1) proprietary LLM services such as GPT-4; 2) public LLMs like QWen; and 3) specialized medical LLMs, like HuatuoGPT2. Through this extensive evaluation, we aim to provide a better understanding of LLMs' performance in the medical domain, ultimately contributing to their more effective deployment in healthcare settings. Xian Wu 0001, Yutian Zhao, Yunyan Zhang, Jiageng Wu, Zhihong Zhu 0001, Zhenxi Lin, Jie Yang 0039, Yefeng Zheng 0001 |
NeurIPS | 3 |
| 2024 | Acquiring New Knowledge Without Losing Old Ones for Effective Continual Dialogue Policy LearningabstractDialogue policy learning is the core decision-making module of a task-oriented dialogue system. Its primary objective is to assist users to achieve their goals effectively in as few turns as possible. A practical dialogue-policy agent must be able to expand its knowledge to handle new scenarios efficiently without affecting its performance. Nevertheless, when adapting to new tasks, existing dialogue-policy agents often fail to retain their existing (old) knowledge. To overcome this predicament, we propose a novel continual dialogue-policy model which tackles the issues of “not forgetting the old” and “acquiring the new” from three different aspects: (1) For effective old-task preservation, we introduce the forgetting preventor which uses a behavior cloning technique to force the agent to take actions consistent with the replayed experience to retain the policy trained on historic tasks. (2) For new-task acquisition, we introduce the adaption accelerator which employs an invariant risk minimization mechanism to produce a stable policy predictor to avoid spurious corrections in training data. (3) For reducing the storage cost of the replayed experience, we introduce a replay manager which helps regularly clean up the old data. The effectiveness of the proposed model is evaluated both theoretically and experimentally and demonstrated favorable results. Yunyan Zhang, Yifan Yang 0008, Yefeng Zheng 0001, Kam-Fai Wong |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Cross-Modal Contrastive Learning for Event Extraction
Shuo Wang 0008, Meizhi Ju, Yunyan Zhang, Yefeng Zheng 0001, Meng Wang 0001, Guilin Qi |
DASFAA (3) | 3 |
| 2022 | BNU: A Balance-Normalization-Uncertainty Model for Incremental Event DetectionabstractEvent detection is challenging in real-world application since new events continually occur and old events still exist which may result in repeated labeling for old events. Therefore, incremental event detection is essential where a model continuously learns new events and meanwhile prevents performance from degrading on old events. Although existing incremental event detection models achieve impressive performance, they face the data imbalance problem between old classes and new classes, and have the knowledge transfer problem which cannot adequately utilize the knowledge provided by the previous model and data. To this end, we propose a Balance-Normalization-Uncertainty (BNU) model to address above problems. Specifically, in order to mitigate the adverse effects of data imbalance, we incorporate a balanced fine-tuning stage and a cosine normalization module. Meanwhile, we consider aleatoric uncertainty to preserve previous knowledge while training for new events. Experimental results show that our proposed method resolves the above challenges effectively and achieves consistent and significant performance on ACE and TAC KBP datasets. Jia Li 0012, Yunyan Zhang, Yifan Yang 0008, Zhicheng An, Yefeng Zheng 0001 |
ICASSP | 2 |
| 2022 | CLINER: Clinical Interrogation Named Entity Recognition
Tianyang Cao, Yifan Yang 0008, Yunyan Zhang, Xi Chen 0003, Baobao Chang, Zhifang Sui, Ruihui Zhao, Yefeng Zheng 0001, Bang Liu 0003 |
KSEM (2) | 4 |
| 2021 | PRGC: Potential Relation and Global Correspondence Based Joint Relational Triple ExtractionabstractHengyi Zheng, Rui Wen, Xi Chen, Yifan Yang, Yunyan Zhang, Ziheng Zhang, Ningyu Zhang, Bin Qin, Xu Ming, Yefeng Zheng. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hengyi Zheng, Rui Wen 0001, Xi Chen 0003, Yifan Yang 0008, Yunyan Zhang, Ningyu Zhang 0001, Xu Ming, Yefeng Zheng 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | Integrate syntax information for target-oriented opinion words extraction with target-specific graph convolutional network
Feng Li 0030, Zequn Zhang, Guangluan Xu, Yang Wang 0056, Yunyan Zhang |
Neurocomputing | 7 |
| 2021 | Adaptive Spatiotemporal Graph Convolutional Networks for Motor Imagery ClassificationabstractClassification of electroencephalogram-based motor imagery (MI-EEG) tasks is crucial in brain computer interfaces (BCI). In view of the characteristics of non-stationarity, time-variability and individual diversity of EEG signals, a novel framework based on graph neural network is proposed for MI-EEG classification. First, an adaptive graph convolutional layer (AGCL) is constructed, by which the electrode channel information are integrated dynamically. We further propose an adaptive spatiotemporal graph convolutional network (ASTGCN), which fully exploits the characteristics of EEG signals in time domain and the channel correlations in spatial domain simultaneously. We execute the experiments using EEG signals recorded at motor imagery scenarios, where twenty-five healthy subjects performed MI movements of the right hand and feet to generate motor commands. Experimental results reveal that the proposed method outperforms state-of-the-art methods in terms of both classification quality and robustness. The advantages of ASTGCN include high accuracy, high efficiency, and robustness to cross-trial and cross-subject variations, making it an ideal candidate for long-term MI-EEG applications. Biao Sun 0003, Han Zhang 0035, Zexu Wu, Yunyan Zhang, Ting Li 0012 |
IEEE Signal Process. Lett. | 4 |
| 2019 | Empower event detection with bi-directional neural language model
Yunyan Zhang, Guangluan Xu, Yang Wang 0056, Lei Wang 0077, Tinglei Huang 0001 |
Knowl. Based Syst. | 1 |
| 2006 | A Novel MRI Texture Analysis of Demyelination and Inflammation in Relapsing-Remitting Experimental Allergic Encephalomyelitis
Yunyan Zhang, Jennifer Wells, Richard Buist, James Peeling, V. Wee Yong, Joseph Ross Mitchell |
MICCAI (1) | 1 |
| 2003 | Texture Analysis of MR Images of Minocycline Treated MS Patients
Yunyan Zhang, Hongmei Zhu, Ricardo José Ferrari, Xingchang Wei, Michael Eliasziw, Luanne M. Metz, Joseph Ross Mitchell |
MICCAI (1) | 1 |