VLDB 2026 Research / reviewers in the wild / expert
Mingquan Lin
dblp:155/5729 · also Minquan Lin
· DBLP profile ↗
19ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-0862-6588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial ApplicationabstractXueqing Peng, Lingfei Qian, Yan Wang, Ruoyu Xiang, Yueru He, Yang Ren, Mingyang Jiang, Vincent Jim Zhang, Yuqing Guo, Jeff Zhao, Huan He, Yi Han, Yun Feng, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Xiaoyu Wang, Penglei Gao, Shengyuan Lin, Keyi Wang, Shanshan Yang, Yilun Zhao, Zhiwei Liu, Peng Lu, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen, Junichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xueqing Peng, Lingfei Qian, Yan Wang 0015, Ruoyu Xiang, Yueru He, Mingyang Jiang, Vincent Jim Zhang, Jeff Zhao, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Penglei Gao, Shengyuan Lin, Yilun Zhao 0001, Zhiwei Liu 0003, Peng Lu 0006, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen 0002, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E. Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen 0003, Jun'ichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie |
ACL (1) | 35 |
| 2026 | Retrieval-augmented in-context learning for multimodal large language models in disease classification
Zaifu Zhan, Shuang Zhou 0012, Xiaoshan Zhou, Yongkang Xiao, Yiran Song, Mingquan Lin, Rui Zhang 0028 |
J. Biomed. Informatics | 10 |
| 2025 | Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal DecodingabstractRecent advancements in multimodal large language models (MLLMs) have significantly improved performance in visual question answering. However, they often suffer from hallucinations. In this work, hallucinations are categorized into two main types: initial hallucinations and snowball hallucinations. We argue that adequate contextual information can be extracted directly from the token interaction process. Inspired by causal inference in the decoding strategy, we propose to leverage causal masks to establish information propagation between multimodal tokens. The hypothesis is that insufficient interaction between those tokens may lead the model to rely on outlier tokens, overlooking dense and rich contextual cues. Therefore, we propose to intervene in the propagation process by tackling outlier tokens to enhance in-context inference. With this goal, we present FarSight, a versatile plug-and-play decoding strategy to reduce attention interference from outlier tokens merely by optimizing the causal mask. The heart of our method is effective token propagation. We design an attention register structure within the upper triangular matrix of the causal mask, dynamically allocating attention to capture attention diverted to outlier tokens. Moreover, a positional awareness encoding method with a diminishing masking rate is proposed, allowing the model to attend to further preceding tokens, especially for video sequence tasks. With extensive experiments, FarSight demonstrates significant hallucination-mitigating performance across different MLLMs on both image and video benchmarks, proving its effectiveness. Zhongxing Xu, Zile Huang, Haochen Xue, Ziyang Chen 0003, Zelin Peng, Sijin Zhou, Wenxue Li 0003, Yulong Li 0002, Wenxuan Song, Shiyan Su, Wei Feng 0015, Jionglong Su, Mingquan Lin, Yifan Peng 0002, Xuelian Cheng, Muhammad Imran Razzak, ZongYuan Ge |
CVPR | 17 |
| 2025 | Towards Collaborative Fairness in Federated Learning Under Imbalanced Covariate ShiftabstractCollaborative fairness is a crucial challenge in federated learning. However, existing approaches often overlook a practical yet complex form of heterogeneity: imbalanced covariate shift. We provide a theoretical analysis of this setting, which motivates the design of FedAKD (Federated Asynchronous Knowledge Distillation) - a simple yet effective approach that balances accurate prediction with collaborative fairness. FedAKD consists of client and server updates. In the client update, we introduce a novel asynchronous knowledge distillation strategy based on our preliminary analysis, which reveals that while correctly predicted samples exhibit similar feature distributions across clients, incorrectly predicted samples show significant variability. This suggests that imbalanced covariate shift primarily arises from misclassified samples. Leveraging this insight, our approach first applies traditional knowledge distillation to update client models while keeping the global model fixed. Next, we select the correctly predicted high-confidence samples and update the global model using these samples, while keeping the client models fixed. The server update simply aggregates all client models. We further provide a theoretical proof of FedAKD's convergence. Experimental results on both public datasets (FashionMNIST and CIFAR10) and a real-world Electronic Health Records (EHR) dataset demonstrate that FedAKD significantly improves collaborative fairness, enhances predictive accuracy, and fosters client participation, even under highly heterogeneous data distributions. Tianrun Yu, Jiaqi Wang 0002, Haoyu Wang 0004, Mingquan Lin, Han Liu 0008, Nelson S. Yee, Fenglong Ma |
KDD (2) | 4 |
| 2025 | A multimodal approach for few-shot biomedical named entity recognition in low-resource languages
Leilei Su, Mingquan Lin, Yifan Peng 0002, Cong Sun 0004 |
J. Biomed. Informatics | 4 |
| 2025 | CXR-LT 2024: A MICCAI challenge on long-tailed, multi-label, and zero-shot disease classification from chest X-ray
Mingquan Lin, Gregory Holste, Song Wang 0026, Yiliang Zhou, Yishu Wei, Imon Banerjee, Pengyi Chen, Tianjie Dai, Yuexi Du, Nicha C. Dvornek, Yuyan Ge, Zuwei Guo, Shohei Hanaoka, Dongkyun Kim, Pablo Messina, Yang Lu 0009, Denis Parra, Donghyun Son, Alvaro Soto, Aisha Urooj Khan, René Vidal, Yosuke Yamagishi, Pingkun Yan, Zefan Yang, Ruichi Zhang, Yang Zhou 0019, Leo A. Celi, Ronald M. Summers, Zhiyong Lu, Hao Chen 0011, Adam E. Flanders, George Shih, Zhangyang Wang, Yifan Peng 0002 |
Medical Image Anal. | 1 |
| 2024 | Deep learning with noisy labels in medical prediction problems: a scoping reviewabstractOBJECTIVES: Medical research faces substantial challenges from noisy labels attributed to factors like inter-expert variability and machine-extracted labels. Despite this, the adoption of label noise management remains limited, and label noise is largely ignored. To this end, there is a critical need to conduct a scoping review focusing on the problem space. This scoping review aims to comprehensively review label noise management in deep learning-based medical prediction problems, which includes label noise detection, label noise handling, and evaluation. Research involving label uncertainty is also included. METHODS: Our scoping review follows the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. We searched 4 databases, including PubMed, IEEE Xplore, Google Scholar, and Semantic Scholar. Our search terms include "noisy label AND medical/healthcare/clinical," "uncertainty AND medical/healthcare/clinical," and "noise AND medical/healthcare/clinical." RESULTS: A total of 60 papers met inclusion criteria between 2016 and 2023. A series of practical questions in medical research are investigated. These include the sources of label noise, the impact of label noise, the detection of label noise, label noise handling techniques, and their evaluation. Categorization of both label noise detection methods and handling techniques are provided. DISCUSSION: From a methodological perspective, we observe that the medical community has been up to date with the broader deep-learning community, given that most techniques have been evaluated on medical data. We recommend considering label noise as a standard element in medical research, even if it is not dedicated to handling noisy labels. Initial experiments can start with easy-to-implement methods, such as noise-robust loss functions, weighting, and curriculum learning. Yishu Wei, Cong Sun 0004, Mingquan Lin, Hongmei Jiang, Yifan Peng 0002 |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | A survey of recent methods for addressing AI fairness and bias in biomedicineabstractOBJECTIVES: Artificial intelligence (AI) systems have the potential to revolutionize clinical practices, including improving diagnostic accuracy and surgical decision-making, while also reducing costs and manpower. However, it is important to recognize that these systems may perpetuate social inequities or demonstrate biases, such as those based on race or gender. Such biases can occur before, during, or after the development of AI models, making it critical to understand and address potential biases to enable the accurate and reliable application of AI models in clinical settings. To mitigate bias concerns during model development, we surveyed recent publications on different debiasing methods in the fields of biomedical natural language processing (NLP) or computer vision (CV). Then we discussed the methods, such as data perturbation and adversarial learning, that have been applied in the biomedical domain to address bias. METHODS: We performed our literature search on PubMed, ACM digital library, and IEEE Xplore of relevant articles published between January 2018 and December 2023 using multiple combinations of keywords. We then filtered the result of 10,041 articles automatically with loose constraints, and manually inspected the abstracts of the remaining 890 articles to identify the 55 articles included in this review. Additional articles in the references are also included in this review. We discuss each method and compare its strengths and weaknesses. Finally, we review other potential methods from the general domain that could be applied to biomedicine to address bias and improve fairness. RESULTS: The bias of AIs in biomedicine can originate from multiple sources such as insufficient data, sampling bias and the use of health-irrelevant features or race-adjusted algorithms. Existing debiasing methods that focus on algorithms can be categorized into distributional or algorithmic. Distributional methods include data augmentation, data perturbation, data reweighting methods, and federated learning. Algorithmic approaches include unsupervised representation learning, adversarial learning, disentangled representation learning, loss-based methods and causality-based methods. Yifan Yang 0006, Mingquan Lin, Han Zhao 0002, Yifan Peng 0002, Furong Huang, Zhiyong Lu |
J. Biomed. Informatics | 2 |
| 2024 | Towards long-tailed, multi-label disease classification from chest X-ray: Overview of the CXR-LT challenge
Gregory Holste, Yiliang Zhou, Song Wang 0026, Ajay Jaiswal, Mingquan Lin, Sherry Zhuge, Yuzhe Yang 0003, Dongkyun Kim, Trong-Hieu Nguyen Mau, Minh-Triet Tran, Jaehyup Jeong, Wongi Park, Jong Bin Ryu, Feng Hong 0004, Arsh Verma, Yosuke Yamagishi, Hyeryeong Seo, Myungjoo Kang, Leo A. Celi, Zhiyong Lu, Ronald M. Summers, George Shih, Zhangyang Wang, Yifan Peng 0002 |
Medical Image Anal. | 5 |
| 2024 | Multimodal Image Classification by Multiview Latent Pattern Extraction, Selection, and CorrelationabstractThe large amount of data available in the modern big data era opens new opportunities to expand our knowledge by integrating information from heterogeneous sources. Multiview learning has recently achieved tremendous success in deriving complementary information from multiple data modalities. This article proposes a framework called multiview latent space projection (MVLSP) to integrate features extracted from multiple sources in a discriminative way to facilitate binary and multiclass classifications. Our approach is associated with three innovations. First, most existing multiview learning algorithms promote pairwise consistency between two views and do not have a natural extension to applications with more than two views. MVLSP finds optimum mappings from a common latent space to match the feature space in each of the views. As the matching is performed on a view-by-view basis, the framework can be readily extended to multiview applications. Second, feature selection in the common latent space can be readily achieved by adding a class view, which matches the latent space representations of training samples with their corresponding labels. Then, high-order view correlations are extracted by considering feature-label correlations. Third, a technique is proposed to optimize the integration of different latent patterns based on their correlations. The experimental results on the prostate image dataset demonstrate the effectiveness of the proposed method. Jianghong Ma, Weixuan Kou, Mingquan Lin, Carmen C. M. Cho, Bernard Chiu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Multiclass and Multilabel Classifications by Consensus and Complementarity-Based Multiview Latent Space ProjectionabstractThe fusion of multiview data sets, in which features of each sample are categorized into distinct groups, is increasingly important in the big data era. Successful multiview learning approaches have mechanisms to enforce consensus and/or complementarity among views. This article introduces a framework called the consensus and complementarity-based multiview latent space projection (MVLSP-2C) that enforces both principles simultaneously. Consensus is established by extracting and representing information shared by all views in a shared latent space, whereas complementarity among views is achieved by the representation in view-specific spaces. As the diversity of the multiview feature representation benefits classification performance, MVLSP-2C minimizes the similarity between the shared and view-specific representations, thereby improving diversity. The driving principle of MVLSP-2C is that the latent space representation is obtained by optimally projecting it to match the original feature space representation on a view-by-view basis. Unlike pairwise consensus methods that enforce consistency between two views, matching on a view-by-view basis allows extensions to settings with more than two views. A related and important advantage of this per-view matching design is that a class view can be readily incorporated to learn a supervised representation that facilitates subsequent classification. As the class view is added without an assumption on the exclusivity of classes, MVLSP-2C is equally applicable to multiclass single-label and multilabel classifications. MVLSP-2C further optimizes the integration of latent variables based on their correlation. Extensive experiments in multiclass and multiview image datasets show that MVLSP-2C produces more accurate classification results as compared to state-of-the-art methods. Jianghong Ma, Weixuan Kou, Mingquan Lin, Carmen C. M. Cho, Bernard Chiu |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | Person re-identification via semi-supervised adaptive graph embedding
Mingquan Lin, Ming-Bo Zhao, Choujun Zhan, Bing Li 0007, Kwok Tai Chui |
Appl. Intell. | 2 |
| 2023 | A scoping review on multimodal deep learning in biomedical images and texts
Zhaoyi Sun, Mingquan Lin, Qingqing Zhu, Qianqian Xie, Fei Wang 0001, Zhiyong Lu, Yifan Peng 0002 |
J. Biomed. Informatics | 2 |
| 2023 | A new classification method for diagnosing COVID-19 pneumonia based on joint CNN features of chest X-ray images and parallel pyramid MLP-mixer module
Wenyu Xing, Ming-Bo Zhao, Mingquan Lin |
Neural Comput. Appl. | 4 |
| 2022 | Cascaded Triplanar Autoencoder M-Net for Fully Automatic Segmentation of Left Ventricle Myocardial Scar From Three-Dimensional Late Gadolinium-Enhanced MR ImagesabstractWhile three-dimensional (3D) late gadolinium-enhanced (LGE) magnetic resonance (MR) imaging provides good conspicuity of small myocardial lesions with short acquisition time, it poses a challenge for image analysis as a large number of axial images are required to be segmented. We developed a fully automatic convolutional neural network (CNN) called cascaded triplanar autoencoder M-Net (CTAEM-Net) to segment myocardial scar from 3D LGE MRI. Two sub-networks were cascaded to segment the left ventricle (LV) myocardium and then the scar within the pre-segmented LV myocardium. Each sub-network contains three autoencoder M-Nets (AEM-Nets) segmenting the axial, sagittal and coronal slices of the 3D LGE MR image, with the final segmentation determined by voting. The AEM-Net integrates three features: (1) multi-scale inputs, (2) deep supervision and (3) multi-tasking. The multi-scale inputs allow consideration of the global and local features in segmentation. Deep supervision provides direct supervision to deeper layers and facilitates CNN convergence. Multi-task learning reduces segmentation overfitting by acquiring additional information from autoencoder reconstruction, a task closely related to segmentation. The framework provides an accuracy of 86.43% and 90.18% for LV myocardium and scar segmentation, respectively, which are the highest among existing methods to our knowledge. The time required for CTAEM-Net to segment LV myocardium and the scar was 49.72 ± 9.69s and 120.25 ± 23.18s per MR volume, respectively. The accuracy and efficiency afforded by CTAEM-Net will make possible future large population studies. The generalizability of the framework was also demonstrated by its competitive performance in two publicly available datasets of different imaging modalities. Mingquan Lin, Mingjie Jiang, Ming-Bo Zhao, Eranga Ukwatta, James A. White, Bernard Chiu |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Multi-task deep learning-based survival analysis on the prognosis of late AMD using the longitudinal data in AREDS
Gregory C. Ghahramani, Matthew Brendel, Mingquan Lin, Qingyu Chen 0001, Tiarnan D. Keenan, Kun Chen 0002, Emily Y. Chew, Zhiyong Lu, Yifan Peng 0002, Fei Wang 0001 |
AMIA | 3 |
| 2021 | Using Radiomics as Prior Knowledge for Thorax Disease Classification and Localization in Chest X-rays
Yan Han 0001, Chongyan Chen, Liyan Tang, Mingquan Lin, Ajay Jaiswal, Song Wang 0026, Ahmed H. Tewfik, George Shih, Ying Ding 0001, Yifan Peng 0002 |
AMIA | 4 |
| 2020 | Deep-recursive residual network for image semantic segmentation
Yue Zhang 0004, Xianrui Li, Mingquan Lin, Bernard Chiu, Ming-Bo Zhao |
Neural Comput. Appl. | 3 |
| 2018 | Trace Ratio Criterion based Discriminative Feature Selection via l2, p-norm regularization for supervised learning
Ming-Bo Zhao, Mingquan Lin, Bernard Chiu, Zhao Zhang 0001, Xue-Song Tang |
Neurocomputing | 2 |