EDBT 2026 Demo / reviewers in the wild / expert
Qian Zhou 0001
dblp:88/123-1
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-7964-8130ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GaitProtector: Impersonation-Driven Gait De-Identification via Training-Free Diffusion Latent Optimization
Huiran Duan, Qian Zhou 0001, Zhongliang Guo 0001, Junhao Dong 0001, Guoying Zhao 0001, Yingli Tian |
FG | 2 |
| 2026 | Region-aware metric learning for few-shot detection of counterfeit cigarettes from packaging images
Qian Zhou 0001, Huanrou Ding, Chengzhe Li, Hua Zou 0002 |
Expert Syst. Appl. | 1 |
| 2025 | MSFP-Net: Multi-Scale Fusion of SAM-Derived Priors for Medical Image SegmentationabstractThe Segment Anything Model (SAM) has demonstrated strong zero-shot segmentation performance on natural images and has gained increasing attention in medical image applications. However, most existing approaches fine-tune SAM or incorporate task-specific adapters to adapt it to medical data, resulting in high computational cost and strong dependence on large-scale annotated datasets. In this paper, we propose MSFP-Net, a medical image segmentation framework that uses the Segment Anything Model (SAM) as a frozen prior generator. Unlike existing methods that fine-tune SAM or insert task-specific adapters, we apply SAM before training to generate coarse segmentation masks. These masks serve as priors and are injected into a U-shaped segmentation network through multiple prior learning blocks within skip connections. Each prior learning block includes a self-update module that applies self-attention to refine the priors, and a multi-scale learning block that performs cross-attention at multiple resolutions to enhance encoder features under the guidance of SAM priors. A dynamic learning block further fuses the original encoder features and the multi-scale updated features using learnable weights, enabling the network to adaptively balance task-specific representations with SAM-guided priors. By injecting SAM priors through prior learning blocks, MSFP- Net avoids fine-tuning and achieves superior segmentation accuracy with reduced computational cost, outperforming state-of-the-art SAM-adapted methods (e.g., SAMUS) by up to 4.9% in Dice score and 9.0% in IoU across four public datasets. Qian Zhou 0001, Hua Zou 0002, Fei Luo 0004 |
BIBM | 2 |
| 2025 | Anatomy-Aware Adaptation of Pre-Trained Models for Medical Difference Visual Question AnsweringabstractMedical Difference Visual Question Answering (Med-Diff-VQA) is a challenging and clinically significant task that requires identifying and interpreting subtle anatomical differences between pairs of medical images, such as chest X-rays, in response to domain-specific questions. Unlike traditional Visual Question Answering (VQA) tasks, Med-DiffVQA is characterized by high visual similarity, limited data availability, and a strong requirement for anatomically precise reasoning. To tackle these challenges, we propose$\mathbf{A}^{\mathbf{2}} \mathbf{M}$-Diff, an Anatomy-Aware adaptation framework that leverages pretrained vision and language models to meet the specific demands of Med-Diff-VQA. Specifically,$\mathbf{A}^{\mathbf{2}}$M-Diff utilizes Medical Masked Autoencoders (MedMAE) and the Medical Segment Anything Model (MedSAM) to extract both global contextual and anatomyfocused visual features. These features are token-compressed, projected, and injected into a pre-trained LLaMA2 language model via prompt-guided multimodal alignment. To efficiently adapt the language model to the medical domain with minimal additional parameters, we adopt Low-Rank Adaptation (LoRA), which updates only a small subset of model parameters. Experimental results on the MIMIC-Diff-VQA dataset demonstrate that$\mathbf{A}^{\mathbf{2}}$M-Diff outperforms existing methods, achieving a BLEU4 score of 0.542, METEOR of 0.412, ROUGE-L of 0.734, and CIDEr of 2.162. These results validate the effectiveness of anatomy-aware representation and lightweight adaptation in finegrained medical reasoning. The code is publicly available at: https://github.com/liyiersan/Med-Diff-VQA. Qian Zhou 0001, Hua Zou 0002, Fei Luo 0004, Xiwen Bai |
BIBM | 1 |
| 2025 | UML: A Unified Multimodal Learning Framework for Cataract Postoperative Visual Acuity Prediction with Uncertain Missing ModalitiesabstractCataracts are the leading cause of blindness worldwide, with surgery as the only effective treatment. Accurate prediction of Best Corrected Visual Acuity (BCVA) is crucial for surgical planning. In this paper, we propose a novel Unified Multimodal Learning (UML) framework for BCVA prediction with uncertain missing modalities. Unlike existing methods that apply generic encoders and overlook critical image variability, UML leverages medical priors to enhance feature extraction through three modules: central concave region enhancement, OCT re-weighting, and multi-scale attention. To manage missing modality uncertainty, we design a missing modality mask fusion network using an attentional mask for unified feature fusion. Additionally, an auxiliary diagnostic text-image contrastive learning task is introduced to further refine image features. UML achieves state-of-the-art performance with a mean absolute error (MAE) of 0.0457 and 96.25% predictions fall within an error of ± 0.10 LogMAR. Codes are available at https://github.com/yty9941/Eyer-BCVA Qian Zhou 0001, Hua Zou 0002 |
ICASSP | 2 |
| 2025 | A Robust 3D CNN with Pyramidal Attention for Spatiotemporal Gait RecognitionabstractGait recognition has become an increasingly important biometric technique for identifying individuals from a distance without requiring their active cooperation. Since gait involves a sequence of motion patterns, effectively capturing temporal dynamics is essential for accurate recognition. Traditional methods that extract temporal features independently and fuse them at a later stage often fail to model the continuity and interdependence of motion across frames. To overcome this limitation, we propose a novel three-dimensional convolutional architecture named Robust Spatiotemporal 3D Convolutional Neural Network (RST3D), which jointly captures spatial and temporal correlations throughout gait sequences. The proposed architecture incorporates a comprehensive 3D convolutional block that operates along the temporal, height, and width dimensions, enabling the network to learn more expressive and coherent spatiotemporal representations. In addition, we introduce a Temporal Pyramidal Attention (TPA) block to enhance the network’s ability to model temporal dependencies by capturing discriminative motion patterns across multiple temporal scales. We evaluate our method on four large-scale gait recognition datasets: CASIA-B, OUMVLP, GREW, and Gait3D. Experimental results show that our approach consistently achieves superior performance compared to existing 3D CNN-based methods, particularly under challenging conditions such as view variation, clothing changes, and occlusion. Jianyu Chen 0008, Qian Zhou 0001, Qin Zou 0001, Chao Liang 0001, Zengmin Xu, Gang Wu 0010, Zhongyuan Wang 0001 |
MMAsia | 2 |
| 2025 | Multi-Modal Gait Recognition via Collaborative Feature Learning from Silhouettes and Skeletons
Jianyu Chen 0008, Zhongyuan Wang 0001, Qian Zhou 0001, Qin Zou 0001, Chao Liang 0001, Gang Wu 0010 |
PRCV (15) | 3 |
| 2025 | IMedSeg: Towards efficient interactive medical segmentation
Zidi Shi, Qian Zhou 0001, Hua Zou 0002 |
Neurocomputing | 3 |
| 2025 | ActiveFreq: Integrating Active Learning and Frequency Domain Analysis for Interactive Segmentation
Lijun Guo, Qian Zhou 0001, Zidi Shi, Hua Zou 0002, Gang Ke |
Knowl. Based Syst. | 2 |
| 2024 | Refining Intraocular Lens Power Calculation: A Multi-modal Framework Using Cross-Layer Attention and Effective Channel Attention
Qian Zhou 0001, Hua Zou 0002, Zhongyuan Wang 0001 |
MICCAI (1) | 1 |
| 2023 | RHViT: A Robust Hierarchical Transformer for 3D Multimodal Brain Tumor Segmentation Using Biased Masked Image Modeling Pre-trainingabstractAccurate brain tumor segmentation in medical image analysis is crucial for diagnosis and treatment planning. While computer-aided methods have shown promise, several challenges persist. Most existing methods struggle with smaller tumors, treating all regions uniformly. Additionally, they lack robustness when dealing with data corruption and handling missing modalities, common in clinical settings. In this paper, we present a robust hierarchical vision transformer (RHViT) for 3D multimodal brain tumor segmentation, employing an encoder-decoder structure. Our approach combines 3D convolutions and self-attention, offering efficient and effective training. 3D convolutions help capture local information and generate hierarchical features, improving tumor segmentation accuracy. To enhance robustness, we pre-train the encoder using masked image modeling (MIM). This pre-training equips the model to handle data corruption, resulting in improved segmentation even in challenging scenarios. Furthermore, we introduce a novel biased masking strategy during MIM to focus the model's attention on tumor regions. This facilitates better tumor representations and effective fusion of multimodal features. Importantly, our biased masking technique strengthens the model's resilience when dealing with incomplete multimodal data during testing, making it a practical choice. Extensive experiments confirm the superiority of our model over existing approaches. Qian Zhou 0001, Hua Zou 0002, Fei Luo 0004, Yishi Qiu |
BIBM | 1 |
| 2023 | Incomplete Multimodal Learning for Visual Acuity Prediction After Cataract Surgery Using Masked Self-Attention
Qian Zhou 0001, Hua Zou 0002 |
MICCAI (7) | 1 |
| 2022 | Long-Tailed Multi-label Retinal Diseases Recognition via Relational Learning and Knowledge Distillation
Qian Zhou 0001, Hua Zou 0002, Zhongyuan Wang 0001 |
MICCAI (2) | 1 |