VLDB 2026 Research / reviewers in the wild / expert
Qiuhui Chen
dblp:78/1438
· DBLP profile ↗
13ranked-venue papers
11as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing 3D Medical Image Understanding With Pretraining Aided by 2D Multimodal Large Language ModelsabstractUnderstanding 3D medical image volumes is critical in the medical field, yet existing 3D medical convolution and transformer-based self-supervised learning (SSL) methods often lack deep semantic comprehension. Recent advancements in multimodal large language models (MLLMs) provide a promising approach to enhance image understanding through text descriptions. To leverage these 2D MLLMs for improved 3D medical image understanding, we propose Med3DInsight, a novel pretraining framework that integrates 3D image encoders with 2D MLLMs via a specially designed plane-slice-aware transformer module. Additionally, our model employs a partial optimal transport based alignment, demonstrating greater tolerance to noise introduced by potential noises in LLM-generated content. Med3DInsight introduces a new paradigm for scalable multimodal 3D medical representation learning without requiring human annotations. Extensive experiments demonstrate our state-of-the-art performance on two downstream tasks, i.e., segmentation and classification, across various public datasets with CT and MRI modalities, outperforming current SSL methods. Med3DInsight can be seamlessly integrated into existing 3D medical image understanding networks, potentially enhancing their performance. Qiuhui Chen, Xuancheng Yao, Huping Ye |
IEEE J. Biomed. Health Informatics | 1 |
| 2026 | HoloDx: Knowledge- and Data-Driven Multimodal Diagnosis of Alzheimer's DiseaseabstractAccurate diagnosis of Alzheimer's disease (AD) requires effectively integrating multimodal data and clinical expertise. However, existing methods often struggle to fully utilize multimodal information and lack structured mechanisms to incorporate dynamic domain knowledge. To address these limitations, we propose HoloDx, a knowledge- and data-driven framework that enhances AD diagnosis by aligning domain knowledge with multimodal clinical data. HoloDx incorporates a knowledge injection module with a knowledge-aware gated cross-attention, allowing the model to dynamically integrate domain-specific insights from both large language models (LLMs) and clinical expertise. A memory injection module with prototypical memory attention further enables consistency preservation across decision trajectories. Through synergistic operation of these components, HoloDx achieves enhanced interpretability while maintaining precise knowledge-data alignment. Evaluations on five AD datasets demonstrate that HoloDx outperforms state-of-the-art methods, achieving superior diagnostic accuracy and strong generalization across diverse cohorts. The source code is released at https://github.com/Qybc/HoloDx. Qiuhui Chen |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Enhancing 3D Medical Image Understanding with 2D Multimodal Large Language ModelsabstractUnderstanding medical image volumes is crucial in healthcare, yet most current models for classification and segmentation often focus narrowly on task-specific features without capturing the broader medical context. To address this, we introduce Med3DInsight, a pre-training framework that enhances 3D image understanding by leveraging 2D multimodal large language models (MLLMs) through a Plane-Slice-Aware Transformer (PSAT) module. Med3DInsight connects 3D image encoders with 2D MLLMs, enhancing representation learning for downstream tasks. Extensive experiments on CT and MRI datasets demonstrate that Med3DInsight achieves state-of-the-art performance, surpassing 19 baseline methods and proving effective across diverse imaging modalities and anatomical structures. This framework can be seamlessly integrated into existing 3D medical imaging networks, significantly boosting their performance and adaptability. Our source code is publicly available at https://github.com/Qybc/Med3DInsight. Qiuhui Chen, Xuancheng Yao, Huping Ye |
ICASSP | 1 |
| 2025 | Volumetric medical image segmentation via scribble annotations and shape priors
Qiuhui Chen, Haiying Lyu |
Mach. Vis. Appl. | 1 |
| 2024 | MedBLIP: Bootstrapping Language-Image Pretraining from 3D Medical Images and Texts
Qiuhui Chen |
ACCV (3) | 1 |
| 2024 | Alifuse: Aligning and Fusing Multimodal Medical Data for Computer-Aided DiagnosisabstractMedical data collected for diagnostic decisions are typically multimodal, providing comprehensive information on a subject. While computer-aided diagnosis systems can benefit from multimodal inputs, effectively fusing such data remains a challenging task and a key focus in medical research. In this paper, we propose a transformer-based framework, called Alifuse, for aligning and fusing multimodal medical data. Specifically, we convert medical images and both unstructured and structured clinical records into vision and language tokens, employing intramodal and intermodal attention mechanisms to learn unified representations of all imaging and non-imaging data for classification. Additionally, we integrate restoration modeling with contrastive learning frameworks, jointly learning the high-level semantic alignment between images and texts and the low-level understanding of one modality with the help of another. We apply Alifuse to classify Alzheimer’s disease, achieving state-of-the-art performance on five public datasets and outperforming eight baselines. The source code is available at https://github.com/Qybc/Alifuse. Qiuhui Chen |
BIBM | 1 |
| 2024 | SMART: Self-Weighted Multimodal Fusion for Diagnostics of Neurodegenerative DisordersabstractMultimodal medical data, such as brain scans and non-imaging clinical records like demographics and neuropsychology examinations, play an important role in diagnosing neurodegenerative disorders, e.g., Alzheimer's disease (AD) and Parkinson's disease (PD). However, the disease-relevant information is overwhelmed by the high-dimensional image scans and the massive non-imaging data, making it a challenging task to fuse multimodal medical inputs efficiently. Recent multimodal learning methods adopt deep encoders to extract features and simple concatenation or alignment techniques for feature fusion, which suffer the representation degeneration issue due to the vast irrelevant information. To address this challenge, we propose a deep self-weighted multimodal relevance weighting approach, which leverages clustering-based constrastive learning and eliminates the intra- and inter-modal irrelevancy. The learned relevance score is integrated as a gate with a multimodal attention transformer to provide an improved fusion for the final diagnosis. Our proposed model, called SMART (Self-weighted Multimodal Attention-and-Relevance gated Transformer), is extensively evaluated on three public AD/PD datasets and achieves state-of-the-art (SOTA) performance in the diagnostics of neurodegenerative disorders. Our source code is available at https://github.com/Qybc/SMART. Qiuhui Chen |
ACM Multimedia | 1 |
| 2024 | LongFormer: Longitudinal Transformer for Alzheimer's Disease Classification with Structural MRIsabstractStructural magnetic resonance imaging (sMRI), especially longitudinal sMRI, is often used to monitor and capture disease progression during the clinical diagnosis of Alzheimer's Disease (AD). However, current methods neglect AD’s progressive nature and have mostly relied on a single image for recognizing AD. In this paper, we consider the problem of leveraging the longitudinal MRIs of a subject for AD classification. To address the challenges of missing data, data demand, and subtle changes over time in learning longitudinal 3D MRIs, we propose a novel model LongFormer, which is a hybrid 3D CNN and transformer design to learn from image and longitudinal flow pairs. Our model can fully leverage all images in a dataset and effectively fuse spatiotemporal features for classification. We evaluate our model on three datasets, i.e., ADNI, OASIS, and AIBL, and compare it to eight baseline algorithms. Our proposed LongFormer achieves state-of-the-art performance in classifying AD and NC subjects from all three public datasets. Our source code is available online at https://github.com/Qybc/LongFormer. Qiuhui Chen |
WACV | 1 |
| 2022 | Scribble2D5: Weakly-Supervised Volumetric Image Segmentation via Scribble Annotations
Qiuhui Chen |
MICCAI (8) | 1 |
| 2018 | SHPD: Surveillance Human Pose Dataset and Performance Evaluation for Coarse-Grained Pose EstimationabstractPose estimation is highly valued in surveillance systems in the era of big data. However, current human pose datasets are limited in their coverage of the pose estimation challenges in outdoor surveillance scenarios. In this paper, we introduce a novel Surveillance Human Pose Dataset (SHPD). Unlike the existing fine-grained parts or key-points based human pose datasets, SHPD is built for two aims: 1) constructing a more specialized human pose benchmark for surveillance tasks, and 2) focusing on coarse-grained global-pose estimation for small scale human objects, which are the most common targets in practical outdoor surveillance applications. The collected images in SHPD are all from on-using surveillance cameras and capture people from a wide and balanced range of outdoor scenarios. A wide variety of surveillance human global poses and their corresponding rich attributes are also provided. Based on SHPD, performance evaluation of global-pose estimation using a few baseline deep-learning networks indicates that, there are ample room for improvement of the recognition accuracy. Qiuhui Chen |
ICIP | 1 |
| 2014 | Amplitudes of mono-component signals and the generalized sampling functions
Qiuhui Chen, Luoqing Li |
Signal Process. | 1 |
| 2010 | A Blind Watermarking Scheme Using New Nontensor Product Wavelet Filter BanksabstractAs an effective method for copyright protection of digital products against illegal usage, watermarking in wavelet domain has recently received considerable attention due to the desirable multiresolution property of wavelet transform. In general, images can be represented with different resolutions by the wavelet decomposition, analogous to the human visual system (HVS). Usually, human eyes are insensitive to image singularities revealed by different high frequency subbands of wavelet decomposed images. Hence, adding watermarks into these singularities will improve the imperceptibility that is a desired property of a watermarking scheme. That is, the capability for revealing singularities of images plays a key role in designing wavelet-based watermarking algorithms. Unfortunately, the existing wavelets have a limited ability in revealing singularities in different directions. This motivates us to construct new wavelet filter banks that can reveal singularities in all directions. In this paper, we utilize special symmetric matrices to construct the new nontensor product wavelet filter banks, which can capture the singularities in all directions. Empirical studies will show their advantages of revealing singularities in comparison with the existing wavelets. Based upon these new wavelet filter banks, we, therefore, propose a modified significant difference watermarking algorithm. Experimental results show its promising results. Xinge You, Yiu-Ming Cheung, Qiuhui Chen |
IEEE Trans. Image Process. | 4 |
| 2006 | Thinning Character Using Modulus Minima of Wavelet TransformabstractAn essential step in character recognition is to extract the skeleton characteristics of the character. In this paper, an efficient algorithm is proposed to extract visually satisfactory skeleton from printed and handwritten characters, which overcomes fundamental shortcomings of our previous skeletonization technique based on the maximum modulus symmetry of wavelet transform (WT). The proposed method is motivated from some desirable properties of the WT with constructed wavelet functions: namely, the local modulus minima of the WT are scale-independent at different level scales and are located at the medial axis of the symmetrical contours of character stroke. Thus the modulus minima of the WT are computed as the intrinsic skeletons of character strokes. To achieve faster implementation, a multiscale processing technique is employed. Thus major structures of the skeleton are extracted using the coarse scale, while fine structures are extracted using the fine scale. We have tested the algorithm on handwritten and printed character images. Experimental results show that the proposed algorithm is applicable to not only binary image but also gray-level image where it can be impractical to use other skeletonization techniques, such as thinning and distance transforms. Further, it can effectively remove unwanted artifacts and branches from the extracted skeletons at the intersections and junctions of character strokes and is robust against noises while most existing methods perform poorly. Xinge You, Qiuhui Chen, Bin Fang 0001, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |