EDBT 2026 Demo / reviewers in the wild / expert
Siyuan Yan
dblp:206/1562
· DBLP profile ↗
18ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dr.V : A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-Grained Spatial-Temporal Grounding
Meng Luo 0010, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Jinxiang Lai, Tianlong Wu, Xinya Du, Siyuan Yan, Jiebo Luo 0001, William Yang Wang, Hao Fei 0001, Mong-Li Lee, Wynne Hsu |
Int. J. Comput. Vis. | 10 |
| 2026 | Learning-Based Temporal Sequence of Constrained Handling Selection for Constrained Multi-Objective Evolutionary OptimizationabstractConstraint-handling techniques and genetic operators are two crucial components in constrained multi-objective evolutionary algorithms (CMOEAs). Recent research in most of CMOEAs has primarily focused on adaptive designs of these components to address various constrained multi-objective optimization problems (CMOPs). However, the evolutionary process of solving a CMOP can involve various characteristics, such as continuity, discreteness, degeneracy, or some combination thereof, necessitating the tailored selection of constraint-handling techniques and genetic operators across different generations. This study conceptualizes these selections as a temporal sequence of constrained handling selection, where the time means the generation number. We argue that discovering the systematic patterns within the sequence based on the historical data of applying different selections significantly improves the performance of CMOEAs in finding Pareto optimal solutions. Based on this conceptualization, we propose a CMOEA with a deep reinforcement learning model for solving CMOPs. Specifically, the deep reinforcement learning model dynamically refines the selection of constraint-handling techniques and genetic operators for upcoming generations by learning from the performance of previous selections, thereby enhancing the predictive accuracy for subsequent selections. Experiments are conducted to validate the performance of the proposed algorithm against nine CMOEAs on thirty-seven benchmark problems and an unmanned aerial vehicle path planning problem. Experimental results show that the proposed algorithm substantially outperforms the compared algorithms regarding the obtained Pareto optimal solutions. Additionally, the results verify that discovering the systematic patterns within the sequence for CMOEAs has a positive impact on solving CMOPs in terms of objective optimization and constraint satisfaction. Chaoda Peng, Siyuan Yan, Cankun Zhong, Qiong Huang 0001, Chunguo Wu, Han Huang 0002 |
IEEE Trans. Evol. Comput. | 2 |
| 2025 | WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image ClassificationabstractMultimodal Large Language Models (MLLMs) have shown promise in visual-textual reasoning, with Multimodal Chain-of-Thought (MCoT) prompting significantly enhancing interpretability.However, existing MCoT methods rely on rationale-rich datasets and largely focus on inter-object reasoning, overlooking the intraobject understanding crucial for image classification.To address this gap, we propose WISE, a Weak-supervIsion-guided Step-bystep Explanation method that augments any image classification dataset with MCoTs by reformulating the concept-based representations from Concept Bottleneck Models (CBMs) into concise, interpretable reasoning chains under weak supervision.Experiments across ten datasets show that our generated MCoTs not only improve interpretability by 37% but also lead to gains in classification accuracy when used to fine-tune MLLMs 1 .Our work bridges concept-based interpretability and generative MCoT reasoning, providing a generalizable framework for enhancing MLLMs in fine-grained visual understanding. Yiwen Jiang, Deval Mehta 0001, Siyuan Yan, Yaling Shen, ZongYuan Ge |
EMNLP | 3 |
| 2025 | OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingabstractSurgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challenging due to this complexity and the limited availability of annotated data. To address the gap, we propose OphCLIP, a hierarchical retrieval-augmented vision-language pretraining framework specifically designed for ophthalmic surgical workflow understanding. OphCLIP leverages the OphVL dataset we constructed, a large-scale and comprehensive collection of over 375K hierarchically structured video-text pairs with tens of thousands of different combinations of attributes (surgeries, phases/operations/actions, instruments, medications, as well as more advanced aspects like the causes of eye diseases, surgical objectives, and postoperative recovery recommendations, etc). These hierarchical video-text correspondences enable OphCLIP to learn both fine-grained and long-term visual representations by aligning short video clips with detailed narrative descriptions and full videos with structured titles, capturing intricate surgical details and high-level procedural insights, respectively. Our OphCLIP also designs a retrieval-augmented pretraining framework to leverage the underexplored large-scale silent surgical procedure videos, automatically retrieving semantically relevant content to enhance the representation learning of narrative videos. Evaluation across 11 datasets for phase recognition and multi-instrument identification shows OphCLIP's robust generalization and superior performance. Kun Yuan 0004, Yaling Shen, Xiaohao Xu, Wei Li 0320, Zhongxing Xu, Zelin Peng, Siyuan Yan, Vinkle Srivastav, Diping Song, Tianbin Li, Danli Shi, Jin Ye 0002, Nicolas Padoy, Nassir Navab, Junjun He, ZongYuan Ge |
ICCV | 11 |
| 2025 | Derm1M: A Million-Scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for DermatologyabstractThe emergence of vision-language models has transformed medical AI, enabling unprecedented advances in diagnostic capability and clinical applications. However, progress in dermatology has lagged behind other medical domains due to the lack of standard image-text pairs. Existing dermatological datasets are limited in both scale and depth, offering only single-label annotations across a narrow range of diseases instead of rich textual descriptions, and lacking the crucial clinical context needed for real-world applications. To address these limitations, we present Derm1M, the first large-scale vision-language dataset for dermatology, comprising 1,029,761 image-text pairs. Built from diverse educational resources and structured around a standard ontology collaboratively developed by experts, Derm1M provides comprehensive coverage for over 390 skin conditions across four hierarchical levels and 130 clinical concepts with rich contextual information such as medical history, symptoms, and skin tone. To demonstrate Derm1M potential in advancing both AI research and clinical application, we pretrained a series of CLIP-like models, collectively called DermLIP, on this dataset. The DermLIP family significantly outperforms state-of-the-art foundation models on eight diverse datasets across multiple tasks, including zero-shot skin disease classification, clinical and artifacts concept identification, few-shot/full-shot learning, and cross-modal retrieval. Our dataset and code will be publicly available at https://github.com/SiyuanYan1/Derm1M upon acceptance. Siyuan Yan, Yiwen Jiang, Xieji Li, Hao Fei 0001, Philipp Tschandl, Harald Kittler, ZongYuan Ge |
ICCV | 1 |
| 2025 | Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model
Siyuan Yan, Jason J. Ong, ZongYuan Ge, Lei Zhang 0095 |
MICCAI (16) | 3 |
| 2025 | MAKE: Multi-Aspect Knowledge-Enhanced Vision-Language Pretraining for Zero-Shot Dermatological Assessment
Siyuan Yan, Xieji Li, Yiwen Jiang, ZongYuan Ge |
MICCAI (5) | 1 |
| 2025 | ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion ModelabstractThe automatic generation of diverse and human-like facial reactions in dyadic dialogue remains a critical challenge for human-computer interaction systems. Existing methods fail to model the stochasticity and dynamics inherent in real human reactions. To address this, we propose ReactDiff, a novel temporal diffusion framework for generating diverse facial reactions that are appropriate for responding to any given dialogue context. Our key insight is that plausible human reactions demonstrate smoothness, and coherence over time, and conform to constraints imposed by human facial anatomy. To achieve this, ReactDiff incorporates two vital priors (spatio-temporal facial kinematics) into the diffusion process: i) temporal facial behavioral kinematics and ii) facial action unit dependencies. These two constraints guide the model toward realistic human reaction manifolds, avoiding visually unrealistic jitters, unstable transitions, unnatural expressions, and other artifacts. Extensive experiments on the REACT2024 dataset demonstrate that our approach not only achieves state-of-the-art reaction quality but also excels in diversity and reaction appropriateness. Our code is publicly available at https://github.com/lingjivoo/ReactDiff. Siyang Song, Siyuan Yan, ZongYuan Ge |
ACM Multimedia | 3 |
| 2025 | Prompt-Driven Latent Domain Generalization for Medical Image ClassificationabstractDeep learning models for medical image analysis easily suffer from distribution shifts caused by dataset artifact bias, camera variations, differences in the imaging station, etc., leading to unreliable diagnoses in real-world clinical settings. Domain generalization (DG) methods, which aim to train models on multiple domains to perform well on unseen domains, offer a promising direction to solve the problem. However, existing DG methods assume domain labels of each image are available and accurate, which is typically feasible for only a limited number of medical datasets. To address these challenges, we propose a unified DG framework for medical image classification without relying on domain labels, called Prompt-driven Latent Domain Generalization (PLDG). PLDG consists of unsupervised domain discovery and prompt learning. This framework first discovers pseudo domain labels by clustering the bias-associated style features, then leverages collaborative domain prompts to guide a Vision Transformer to learn knowledge from discovered diverse domains. To facilitate cross-domain knowledge learning between different prompts, we introduce a domain prompt generator that enables knowledge sharing between domain prompts and a shared prompt. A domain mixup strategy is additionally employed for more flexible decision margins and mitigates the risk of incorrect domain assignments. Extensive experiments on three medical image classification tasks and one debiasing task demonstrate that our method can achieve comparable or even superior performance than conventional DG algorithms without relying on domain labels. Our code is publicly available at https://github.com/SiyuanYan1/PLDG/tree/main. Siyuan Yan, Chi Liu 0002, Lie Ju, Dwarikanath Mahapatra, Brigid Betz-Stablein, Victoria Mar, Monika Janda, H. Peter Soyer, ZongYuan Ge |
IEEE Trans. Medical Imaging | 1 |
| 2024 | OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding
Peng Xia 0005, Lin Wang 0027, Siyuan Yan, Zhongxing Xu, Yimin Luo, Kaimin Song, Jürgen Leitner, Xuelian Cheng, Chi Liu 0002, Kaijing Zhou, ZongYuan Ge |
ECCV (4) | 4 |
| 2024 | Speciesism and Preference of Human-Artificial Intelligence Interaction: A Study on Medical Artificial IntelligenceabstractArtificial intelligence (AI) has revolutionized the medical industry in the decade. It is critical to integrate human–computer interaction into daily clinic service and further increase the public acceptance of medical AI. Based on self-categorization theory, our research draws on speciesism as a vital cognitive factor to examine how patients’ speciesism affects their acceptance of medical AI in different roles. The study adopted a positivist research paradigm by examining 249 samples of data collected during COVID-19 in China. The results indicate that patients with higher speciesism tend to have lower acceptance of medical AI in an independent role but higher acceptance in an assistive role. Furthermore, we verified the mediating effect of human–computer trust and the positive moderating role of human uniqueness perception. This article expands the practicality of speciesism from human–animal relationships into human–AI relationships and contributes to human–computer interaction from the perspective of medical AI acceptance. Weiwei Huo, Jingjing Qu, Siyuan Yan, Jinyi Yan |
Int. J. Hum. Comput. Interact. | 5 |
| 2023 | Towards Trustable Skin Cancer Diagnosis via Rewriting Model's DecisionabstractDeep neural networks have demonstrated promising performance on image recognition tasks. However, they may heavily rely on confounding factors, using irrelevant artifacts or bias within the dataset as the cue to improve performance. When a model performs decision-making based on these spurious correlations, it can become untrustable and lead to catastrophic outcomes when deployed in the realworld scene. In this paper, we explore and try to solve this problem in the context of skin cancer diagnosis. We introduce a human-in-the-loop framework in the model training process such that users can observe and correct the model's decision logic when confounding behaviors happen. Specifically, our method can automatically discover confounding factors by analyzing the co-occurrence behavior of the samples. It is capable of learning confounding concepts using easily obtained concept exemplars. By mapping the black-box model's feature representation onto an explainable concept space, human users can interpret the concept and intervene via first order-logic instruction. We systematically evaluate our method on our newly crafted, well-controlled skin lesion dataset and several public skin lesion datasets. Experiments show that our method can effectively detect and remove confounding factors from datasets without any prior knowledge about the category distribution and does not require fully annotated concept labels. We also show that our method enables the model to focus on clinical-related concepts, improving the model's performance and trustworthiness during model inference. Siyuan Yan, Dwarikanath Mahapatra, Shekhar Chandra, Monika Janda, H. Peter Soyer, ZongYuan Ge |
CVPR | 1 |
| 2023 | EPVT: Environment-Aware Prompt Vision Transformer for Domain Generalization in Skin Lesion Recognition
Siyuan Yan, Chi Liu 0002, Lie Ju, Dwarikanath Mahapatra, Victoria Mar, Monika Janda, H. Peter Soyer, ZongYuan Ge |
MICCAI (7) | 1 |
| 2023 | NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity UnderstandingabstractThe application of deep learning to nursing procedure activity understanding has the potential to greatly enhance the quality and safety of nurse-patient interactions. By utilizing the technique, we can facilitate training and education, improve quality control, and enable operational compliance monitoring. However, the development of automatic recognition systems in this field is currently hindered by the scarcity of appropriately labeled datasets. The existing video datasets pose several limitations: 1) these datasets are small-scale in size to support comprehensive investigations of nursing activity; 2) they primarily focus on single procedures, lacking expert-level annotations for various nursing procedures and action steps; and 3) they lack temporally localized annotations, which prevents the effective localization of targeted actions within longer video sequences. To mitigate these limitations, we propose NurViD, a large video dataset with expert-level annotation for nursing procedure activity understanding. NurViD consists of over 1.5k videos totaling 144 hours, making it approximately four times longer than the existing largest nursing activity datasets. Notably, it encompasses 51 distinct nursing procedures and 177 action steps, providing a much more comprehensive coverage compared to existing datasets that primarily focus on limited procedures. To evaluate the efficacy of current deep learning methods on nursing activity understanding, we establish three benchmarks on NurViD: procedure recognition on untrimmed videos, procedure and action recognition on trimmed videos, and action detection. Our benchmark and code will be available at https://github.com/minghu0830/NurViD-benchmark. Lin Wang 0027, Siyuan Yan, Don Ma, Qingli Ren, Peng Xia 0005, Wei Feng 0015, Peibo Duan, Lie Ju, ZongYuan Ge |
NeurIPS | 3 |
| 2022 | Transmission-Guided Bayesian Generative Model for Smoke SegmentationabstractSmoke segmentation is essential to precisely localize wildfire so that it can be extinguished in an early phase. Although deep neural networks have achieved promising results on image segmentation tasks, they are prone to be overconfident for smoke segmentation due to its non-rigid shape and transparent appearance. This is caused by both knowledge level uncertainty due to limited training data for accurate smoke segmentation and labeling level uncertainty representing the difficulty in labeling ground-truth. To effectively model the two types of uncertainty, we introduce a Bayesian generative model to simultaneously estimate the posterior distribution of model parameters and its predictions. Further, smoke images suffer from low contrast and ambiguity, inspired by physics-based image dehazing methods, we design a transmission-guided local coherence loss to guide the network to learn pair-wise relationships based on pixel distance and the transmission feature. To promote the development of this field, we also contribute a high-quality smoke segmentation dataset, SMOKE5K, consisting of 1,400 real and 4,000 synthetic images with pixel-wise annotation. Experimental results on benchmark testing datasets illustrate that our model achieves both accurate predictions and reliable uncertainty maps representing model ignorance about its prediction. Our code and dataset are publicly available at: https://github.com/redlessme/Transmission-BVM. Siyuan Yan, Jing Zhang 0052, Nick Barnes |
AAAI | 1 |
| 2021 | Stress Recognition in Thermal Videos Using Bi-directional Long-Term Recurrent Convolutional Neural Networks
Siyuan Yan, Abhijit Adhikary |
ICONIP (2) | 1 |
| 2020 | Spatiotemporal Tree Filtering for Enhancing Image Change DetectionabstractChange detection has received extensive attention because of its realistic significance and broad application fields. However, none of the existing change detection algorithms can handle all scenarios and tasks so far. Different from the most of contributions from the research community in recent years, this paper does not work on designing new change detection algorithms. We, instead, solve the problem from another perspective by enhancing the raw detection results after change detection. As a result, the proposed method is applicable to various kinds of change detection methods, and regardless of how the results are detected. In this paper, we propose Fast Spatiotemporal Tree Filter (FSTF), a purely unsupervised detection method, to enhance coarse binary detection masks obtained by different kinds of change detection methods. In detail, the proposed FSTF has adopted a volumetric structure to effectively synthesize spatiotemporal information of the same target from the current time and history frames to enhance detection. The computational complexity analyzed in the view of graph theory also show that the fast realization of FSTF is a linear time algorithm, which is capable of handling efficient on-line detection tasks. Finally, comprehensive experiments based on qualitative and quantitative analysis verify that FSTF-based change detection enhancement is superior to several other state-of-the-art methods including fully connected Conditional Random Field (CRF), joint bilateral filter, and guided filter. It is illustrated that FSTF is versatile enough to also improve saliency detection as well as semantic image segmentation. Dawei Li 0001, Siyuan Yan, Ming-Bo Zhao, Tommy W. S. Chow |
IEEE Trans. Image Process. | 2 |
| 2018 | Digital Predistortion for Spectrum Compliance in the Internet of Things
Siyuan Yan, Changhong Jiang, Lingmei Wang, Fu Li 0004 |
J. Electron. Test. | 1 |