EDBT 2026 Demo / reviewers in the wild / expert
Yiqun Sun
dblp:277/5039
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Partially Shared Concept Bottleneck ModelsabstractConcept Bottleneck Models (CBMs) enhance interpretability by introducing a layer of human-understandable concepts between inputs and predictions. While recent methods automate concept generation using Large Language Models (LLMs) and Vision-Language Models (VLMs), they still face three fundamental challenges: poor visual grounding, concept redundancy, and the absence of principled metrics to balance predictive accuracy and concept compactness. We introduce PS-CBM, a Partially Shared CBM framework that addresses these limitations through three core components: (1) a multimodal concept generator that integrates LLM-derived semantics with exemplar-based visual cues; (2) a Partially Shared Concept Strategy that merges concepts based on activation patterns to balance specificity and compactness; and (3) Concept-Efficient Accuracy (CEA), a post-hoc metric that jointly captures both predictive accuracy and concept compactness. Extensive experiments on eleven diverse datasets show that PS-CBM consistently outperforms state-of-the-art CBMs, improving classification accuracy by 1.0%–7.4% and CEA by 2.0%–9.5%, while requiring significantly fewer concepts. These results underscore PS-CBM’s effectiveness in achieving both high accuracy and strong interpretability. Delong Zhao, Yiqun Sun |
AAAI | 4 |
| 2026 | Uncertainty-Guided Iterative Contrastive Fusion for Reliable Survival Prediction in Rectal CancerabstractIntegrating multimodal radiological images and clinical data is critical for survival prediction in rectal cancer. However, existing methods often lack sufficient consideration of 1) modality heterogeneity (caused by rectal peristalsis, noise artifacts, and missing modalities) and 2) site heterogeneity (caused by different imaging protocols and patient populations). These factors hinder the model from capturing reliable cross-modal relationships and adapting to distribution shifts across clinical sites. In this work, we propose UICSurv, a novel multimodal Survival prediction framework highlighted by Uncertainty-guided Iterative Contrastive fusion, to capture robust cross-site multimodal interactions while leveraging sample-level uncertainty to enhance fusion reliability. Specifically, UICSurv initializes a shared multimodal embedding and iteratively refines it by fusing each heterogeneous modality via the cross-attention mechanism. In each iteration, a novel Survival Contrastive Learning (SCL) strategy is designed to progressively enhance both cross-site alignment and survival discriminability of the multimodal embedding space. Moreover, we design an EvidenceHit module, which employs temporally consistent evidential learning to jointly estimate survival probabilities and uncertainty. The estimated uncertainty further guides the embedding alignment by reducing the interference of unreliable samples. All components operate synergistically within UICSurv to reinforce reliable survival prediction in rectal cancer. Extensive experiments on multimodal datasets of rectal cancer (collected from three sites) demonstrate the superiority of our method both in survival prediction and uncertainty estimation. The code is available open-source: https://github.com/ScorpioBao/UICSurv. Qingsen Bao, Lei Chen 0011, Kaicong Sun, Yiqun Sun, Fu Xiao 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2026 | An Alignment and Imputation Network (AINet) for Breast Cancer Diagnosis With Multimodal Multi-View Ultrasound ImagesabstractRecently, numerous deep learning models have been proposed for breast cancer diagnosis using multimodal multi-view ultrasound images. However, their performance could be highly affected by overlooking interactions between different modalities and views. Moreover, existing methods struggle to handle cases where certain modalities or views are missing, which limits their clinical applications. To address these issues, we propose a novel Alignment and Imputation Network (AINet) by integrating 1) alignment and imputation pre-training, and 2) hierarchical fusion fine-tuning. Specifically, in the pre-training stage, cross-modal contrastive learning is employed to align features across different modalities, for effectively capturing inter-modal interactions. To simulate missing modality (view) scenarios, we randomly mask out features and then impute them by leveraging inter-modal and inter-view relationships. Following the clinical diagnosis procedure, the subsequent fine-tuning stage further incorporates modality-level and view-level fusion in a hierarchical manner. The proposed AINet is developed and evaluated on three datasets, comprising 15,223 subjects in total. Experimental results demonstrate that AINet significantly outperforms state-of-the-art methods, particularly in handling missing modalities (views). This highlights its robustness and potential for real-world clinical applications. Yonghao Li, Yiqun Sun, Yaling Chen, Shichong Zhou, Zhenhui Li, Xuejun Qian, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Balancing Relevance and Diversity in k-Maximum Inner Product Search
Yanhao Wang 0001, Yiqun Sun, Anthony K. H. Tung, Jun Yu 0002 |
VLDB J. | 3 |
| 2025 | Don't Reinvent the Wheel: Efficient Instruction-Following Text Embedding based on Guided Space TransformationabstractYingchaojie Feng, Yiqun Sun, Yandong Sun, Minfeng Zhu, Qiang Huang, Anthony Kum Hoe Tung, Wei Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yingchaojie Feng, Yiqun Sun, Yandong Sun, Minfeng Zhu 0001, Anthony K. H. Tung, Wei Chen 0001 |
ACL (1) | 2 |
| 2025 | PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-EncoderabstractSemantic Text Embedding is a fundamental NLP task that encodes textual content into vector representations, where proximity in the embedding space reflects semantic similarity.While existing embedding models excel at capturing general meaning, they often overlook ideological nuances, limiting their effectiveness in tasks that require an understanding of political bias.To address this gap, we introduce PRISM, the first framework designed to Produce inteRpretable polItical biaS eMbeddings.PRISM operates in two key stages: (1) Controversial Topic Bias Indicator Mining, which systematically extracts fine-grained political topics and their corresponding bias indicators from weakly labeled news data, and (2) Cross-Encoder Political Bias Embedding, which assigns structured bias scores to news articles based on their alignment with these indicators.This approach ensures that embeddings are explicitly tied to bias-revealing dimensions, enhancing both interpretability and predictive power.Through extensive experiments on two large-scale datasets, we demonstrate that PRISM outperforms stateof-the-art text embedding models in political bias classification while offering highly interpretable representations that facilitate diversified retrieval and ideological analysis. Yiqun Sun, Anthony K. H. Tung, Jun Yu 0002 |
ACL (1) | 1 |
| 2025 | Uncovering the Bigger Picture: Comprehensive Event Understanding Via Diverse News RetrievalabstractAccess to diverse perspectives is essential for understanding real-world events, yet most news retrieval systems prioritize textual relevance, leading to redundant results and limited viewpoint exposure.We propose NEWSCOPE, a two-stage framework for diverse news retrieval that enhances event coverage by explicitly modeling semantic variation at the sentence level.The first stage retrieves topically relevant content using dense retrieval, while the second stage applies sentence-level clustering and diversity-aware re-ranking to surface complementary information.To evaluate retrieval diversity, we introduce three interpretable metrics, namely Average Pairwise Distance, Positive Cluster Coverage, and Information Density Ratio, and construct two paragraph-level benchmarks: LocalNews and DSGlobal.Experiments show that NEWSCOPE consistently outperforms strong baselines, achieving significantly higher diversity without compromising relevance.Our results demonstrate the effectiveness of fine-grained, interpretable modeling in mitigating redundancy and promoting comprehensive event understanding. Yiqun Sun, Anthony K. H. Tung |
EMNLP | 3 |
| 2025 | A General Framework for Producing Interpretable Semantic Text EmbeddingsabstractSemantic text embedding is essential to many tasks in Natural Language Processing (NLP). While black-box models are capable of generating high-quality embeddings, their lack of interpretability limits their use in tasks that demand transparency. Recent approaches have improved interpretability by leveraging domain-expert-crafted or LLM-generated questions, but these methods rely heavily on expert input or well-prompt design, which restricts their generalizability and ability to generate discriminative questions across a wide range of tasks. To address these challenges, we introduce \algo{CQG-MBQA} (Contrastive Question Generation - Multi-task Binary Question Answering), a general framework for producing interpretable semantic text embeddings across diverse tasks. Our framework systematically generates highly discriminative, low cognitive load yes/no questions through the \algo{CQG} method and answers them efficiently with the \algo{MBQA} model, resulting in interpretable embeddings in a cost-effective manner. We validate the effectiveness and interpretability of \algo{CQG-MBQA} through extensive experiments and ablation studies, demonstrating that it delivers embedding quality comparable to many advanced black-box models while maintaining inherently interpretability. Additionally, \algo{CQG-MBQA} outperforms other interpretable text embedding methods across various downstream tasks. The source code is available at \url{https://github.com/dukesun99/CQG-MBQA}. Yiqun Sun, Anthony K. H. Tung, Jun Yu 0002 |
ICLR | 1 |
| 2025 | Predicting Alzheimer's Disease Progression Using a Regression-Based Survival Model with Longitudinal Data
Yiqun Sun, Jincheng Gu, Qinsen Bao, Feihong Liu, Dinggang Shen |
MICCAI (15) | 2 |
| 2025 | MAST-Pro: Dynamic Mixture-of-Experts for Adaptive Segmentation of Pan-Tumors with Knowledge-Driven Prompts
Runqi Meng, Sifan Song, Pengfei Jin, Yiqun Sun, Yujin Oh, Xiang Li 0001, Quanzheng Li, Dinggang Shen |
MICCAI (16) | 6 |
| 2025 | Assessment of hybrid kernel function in extreme support vector regression model for streamflow time series forecasting based on a bayesian estimator decomposition algorithm
Lei Xu 0046, Simin Qu, Hongshi Wu, Qiongfang Li, Yiqun Sun |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | DiversiNews: Enriching News Consumption with Relevant yet Diverse News Articles RetrievalabstractIn the digital age, where echo chambers on social media and news platforms increasingly shape public opinion, there is a growing need for tools that present news consumers with a broad spectrum of perspectives. To this end, we introduce DiversiNews, a novel system designed to diversify news consumption by providing readers with articles that are not only relevant to their interests but also offer a variety of viewpoints. DiversiNews leverages state-of-the-art semantic text encoding techniques and implements advanced Diversity-aware k -Maximum Inner Product Search (D k MIPS) algorithms. Our demonstration highlights the potential of DiversiNews to broaden users' exposure to different viewpoints, thereby countering the polarizing effect of digital echo chambers. We showcase how DiversiNews can enrich the news reading experience, supporting the development of a more informed and balanced public discourse in digital news consumption applications. Yiqun Sun, Yanhao Wang 0001, Anthony K. H. Tung |
Proc. VLDB Endow. | 1 |
| 2023 | Developing Large Pre-trained Model for Breast Tumor Segmentation from Ultrasound Images
Meiyu Li, Kaicong Sun, Yuning Gu, Kai Zhang 0039, Yiqun Sun, Zhenhui Li, Dinggang Shen |
MICCAI (7) | 5 |
| 2023 | Breast Fibroglandular Tissue Segmentation for Automated BPE Quantification With Iterative Cycle-Consistent Semi-Supervised LearningabstractBackground Parenchymal Enhancement (BPE) quantification in Dynamic Contrast-Enhanced Magnetic Resonance Imaging (DCE-MRI) plays a pivotal role in clinical breast cancer diagnosis and prognosis. However, the emerging deep learning-based breast fibroglandular tissue segmentation, a crucial step in automated BPE quantification, often suffers from limited training samples with accurate annotations. To address this challenge, we propose a novel iterative cycle-consistent semi-supervised framework to leverage segmentation performance by using a large amount of paired pre-/post-contrast images without annotations. Specifically, we design the reconstruction network, cascaded with the segmentation network, to learn a mapping from the pre-contrast images and segmentation predictions to the post-contrast images. Thus, we can implicitly use the reconstruction task to explore the inter-relationship between these two-phase images, which in return guides the segmentation task. Moreover, the reconstructed post-contrast images across multiple auto-context modeling-based iterations can be viewed as new augmentations, facilitating cycle-consistent constraints across each segmentation output. Extensive experiments on two datasets with various data distributions show great segmentation and BPE quantification accuracy compared with other state-of-the-art semi-supervised methods. Importantly, our method achieves 11.80 times of quantification accuracy improvement along with 10 times faster, compared with clinical physicians, demonstrating its potential for automated BPE quantification. The code is available at https://github.com/ZhangJD-ong/Iterative-Cycle-consistent-Semi-supervised-Learning-for-fibroglandular-tissue-segmentation. Zhiming Cui 0001, Luping Zhou, Yiqun Sun, Zhenhui Li, Zaiyi Liu, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2021 | A newly-designed fault diagnostic method for transformers via improved empirical wavelet transform and kernel extreme learning machine
Sijia Lu, Wei Gao 0023, Cui Hong, Yiqun Sun |
Adv. Eng. Informatics | 4 |