EDBT 2026 Demo / reviewers in the wild / expert
Hua Zou 0002
dblp:66/3384-2
· DBLP profile ↗
23ranked-venue papers
0as first author
19since 2021 · last 2026
0000-0002-3641-2686ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Affordance-Aware Robotic Dexterous Grasping with Human-like PriorsabstractA dexterous hand capable of generalizable grasping objects is fundamental for the development of general-purpose embodied AI. However, previous methods focus narrowly on low-level grasp stability metrics, neglecting affordance-aware positioning and human-like poses which are crucial for downstream manipulation. To address these limitations, we propose AffordDex, a novel framework with two-stage training that learns a universal grasping policy with an inherent understanding of both motion priors and object affordances. In the first stage, a trajectory imitator is pre-trained on a large corpus of human hand motions to instill a strong prior for natural movement. In the second stage, a residual module is trained to adapt these general human-like motions to specific object instances. This refinement is critically guided by two components: our Negative Affordance-aware Segmentation (NAA) module, which identifies functionally inappropriate contact regions, and a privileged teacher-student distillation process that ensures the final vision-based policy is highly successful. Extensive experiments demonstrate that AffordDex not only achieves universal dexterous grasping but also remains remarkably human-like in posture and functionally appropriate in contact location. As a result, AffordDex significantly outperforms state-of-the-art baselines across seen objects, unseen instances, and even entirely novel categories. Linghao Zhuang, Xingyue Zhao, Yuming Jiang 0007, Jun Cen, Kexiang Wang, Jiayan Guo, Siteng Huang, Xin Li 0056, Deli Zhao, Hua Zou 0002 |
AAAI | 13 |
| 2026 | Region-aware metric learning for few-shot detection of counterfeit cigarettes from packaging images
Qian Zhou 0001, Huanrou Ding, Chengzhe Li, Hua Zou 0002 |
Expert Syst. Appl. | 4 |
| 2026 | DCAFNet: a lightweight network with dynamic context-aware fusion for image deblurring
Gang Ke, Sio-Long Lo, Hua Zou 0002 |
Multim. Syst. | 3 |
| 2026 | MSTDNet: Multi-scale traffic object detection network with smooth information perception
Jie Hua 0005, Zhongyuan Wang 0001, Hua Zou 0002, Gang Wu 0010, Jiayi Ma 0001 |
Pattern Recognit. | 4 |
| 2025 | MSFP-Net: Multi-Scale Fusion of SAM-Derived Priors for Medical Image SegmentationabstractThe Segment Anything Model (SAM) has demonstrated strong zero-shot segmentation performance on natural images and has gained increasing attention in medical image applications. However, most existing approaches fine-tune SAM or incorporate task-specific adapters to adapt it to medical data, resulting in high computational cost and strong dependence on large-scale annotated datasets. In this paper, we propose MSFP-Net, a medical image segmentation framework that uses the Segment Anything Model (SAM) as a frozen prior generator. Unlike existing methods that fine-tune SAM or insert task-specific adapters, we apply SAM before training to generate coarse segmentation masks. These masks serve as priors and are injected into a U-shaped segmentation network through multiple prior learning blocks within skip connections. Each prior learning block includes a self-update module that applies self-attention to refine the priors, and a multi-scale learning block that performs cross-attention at multiple resolutions to enhance encoder features under the guidance of SAM priors. A dynamic learning block further fuses the original encoder features and the multi-scale updated features using learnable weights, enabling the network to adaptively balance task-specific representations with SAM-guided priors. By injecting SAM priors through prior learning blocks, MSFP- Net avoids fine-tuning and achieves superior segmentation accuracy with reduced computational cost, outperforming state-of-the-art SAM-adapted methods (e.g., SAMUS) by up to 4.9% in Dice score and 9.0% in IoU across four public datasets. Qian Zhou 0001, Hua Zou 0002, Fei Luo 0004 |
BIBM | 3 |
| 2025 | Anatomy-Aware Adaptation of Pre-Trained Models for Medical Difference Visual Question AnsweringabstractMedical Difference Visual Question Answering (Med-Diff-VQA) is a challenging and clinically significant task that requires identifying and interpreting subtle anatomical differences between pairs of medical images, such as chest X-rays, in response to domain-specific questions. Unlike traditional Visual Question Answering (VQA) tasks, Med-DiffVQA is characterized by high visual similarity, limited data availability, and a strong requirement for anatomically precise reasoning. To tackle these challenges, we propose$\mathbf{A}^{\mathbf{2}} \mathbf{M}$-Diff, an Anatomy-Aware adaptation framework that leverages pretrained vision and language models to meet the specific demands of Med-Diff-VQA. Specifically,$\mathbf{A}^{\mathbf{2}}$M-Diff utilizes Medical Masked Autoencoders (MedMAE) and the Medical Segment Anything Model (MedSAM) to extract both global contextual and anatomyfocused visual features. These features are token-compressed, projected, and injected into a pre-trained LLaMA2 language model via prompt-guided multimodal alignment. To efficiently adapt the language model to the medical domain with minimal additional parameters, we adopt Low-Rank Adaptation (LoRA), which updates only a small subset of model parameters. Experimental results on the MIMIC-Diff-VQA dataset demonstrate that$\mathbf{A}^{\mathbf{2}}$M-Diff outperforms existing methods, achieving a BLEU4 score of 0.542, METEOR of 0.412, ROUGE-L of 0.734, and CIDEr of 2.162. These results validate the effectiveness of anatomy-aware representation and lightweight adaptation in finegrained medical reasoning. The code is publicly available at: https://github.com/liyiersan/Med-Diff-VQA. Qian Zhou 0001, Hua Zou 0002, Fei Luo 0004, Xiwen Bai |
BIBM | 3 |
| 2025 | UML: A Unified Multimodal Learning Framework for Cataract Postoperative Visual Acuity Prediction with Uncertain Missing ModalitiesabstractCataracts are the leading cause of blindness worldwide, with surgery as the only effective treatment. Accurate prediction of Best Corrected Visual Acuity (BCVA) is crucial for surgical planning. In this paper, we propose a novel Unified Multimodal Learning (UML) framework for BCVA prediction with uncertain missing modalities. Unlike existing methods that apply generic encoders and overlook critical image variability, UML leverages medical priors to enhance feature extraction through three modules: central concave region enhancement, OCT re-weighting, and multi-scale attention. To manage missing modality uncertainty, we design a missing modality mask fusion network using an attentional mask for unified feature fusion. Additionally, an auxiliary diagnostic text-image contrastive learning task is introduced to further refine image features. UML achieves state-of-the-art performance with a mean absolute error (MAE) of 0.0457 and 96.25% predictions fall within an error of ± 0.10 LogMAR. Codes are available at https://github.com/yty9941/Eyer-BCVA Qian Zhou 0001, Hua Zou 0002 |
ICASSP | 3 |
| 2025 | PhysSplat: Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting
Hao Wang 0218, Xingyue Zhao, Hao Fei 0001, Hongqiu Wang, Chengjiang Long, Hua Zou 0002 |
ICCV | 7 |
| 2025 | IMedSeg: Towards efficient interactive medical segmentation
Zidi Shi, Qian Zhou 0001, Hua Zou 0002 |
Neurocomputing | 4 |
| 2025 | ActiveFreq: Integrating Active Learning and Frequency Domain Analysis for Interactive Segmentation
Lijun Guo, Qian Zhou 0001, Zidi Shi, Hua Zou 0002, Gang Ke |
Knowl. Based Syst. | 4 |
| 2025 | DiffusionMOT: A Diffusion-Based Multiple Object TrackerabstractRecently, researchers have introduced diffusion models into multiple object tracking (MOT) tasks. However, existing diffusion-based MOT methods, such as DiffusionTrack, have significant limitations, including frequent ID switching, reduced performance when tracking nonlinear motion objects, and long inference time. To this end, we propose a more effective diffusion-based multiple object tracker named DiffusionMOT. In particular, we propose a mixed intersection over union (IoU) and Re-Identification (ReID) method for trajectory matching, which effectively reduces incorrect matches. Meanwhile, we propose a secondary calibration method for trajectory boxes, improving the accuracy of the generated detection boxes. Moreover, we introduce the parallel sampling technique from the field of image generation into object tracking and propose a parallel sampling module to enhance the model's inference speed while maintaining tracking accuracy. Furthermore, we design a pair-based two-stage matching (PTM) pipeline to more effectively utilize potential detection information. Extensive experiments on several public MOT benchmarks, including DanceTrack, SportsMOT, MOT20, and MOT17, demonstrate that our approach achieves state-of-the-art (SOTA) performance. The code and models are available at https://github.com/sad123-yx/DiffusionMOT. Yaxuan Hu 0001, Jie Hua 0005, Zhen Han 0002, Hua Zou 0002, Gang Wu 0010, Zhongyuan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | MFSegDiff: A Multi-Frequency Diffusion Model for Medical Image SegmentationabstractMedical image segmentation accurately identifies and delineates diagnostic regions, which is a crucial step in the early detection and accurate diagnosis of diseases. Diffusion models have demonstrated remarkable performance in preserving image details and structures by gradually adding noise followed by a reverse denoising process, making them widely explored in image segmentation tasks. In this study, we propose a novel approach for medical image segmentation utilizing diffusion models, termed MFSegDiff, which frames the segmentation task as an iterative denoising process. To address the inconsistency between image semantic features and noise embeddings, we introduce a Cross-Attention Alignment Module(CAAM). This module enhances the original image features and integrates noise and semantic information into the network through a linear attention mechanism. Additionally, we employ a Global-Local Multi-Frequency Module (GLMFM) to extract global contextual information from the spatial to the frequency domain by integrating multi-frequency features with local features. We evaluate the proposed on three datasets, including ISIC-2017, ISIC-2018, and ROSE. Results demonstrate strong generalization capabilities and achieve state-of-the-art segmentation performance, highlighting significant potential for clinical applications. Zidi Shi, Hua Zou 0002, Fei Luo 0004, Zhiyu Huo |
BIBM | 2 |
| 2024 | Refining Intraocular Lens Power Calculation: A Multi-modal Framework Using Cross-Layer Attention and Effective Channel Attention
Qian Zhou 0001, Hua Zou 0002, Zhongyuan Wang 0001 |
MICCAI (1) | 2 |
| 2024 | Synergistic registration of CT-MRI brain images and retinal images: A novel approach leveraging reinforcement learning and modified artificial rabbit optimization
Xiaolei Luo, Hua Zou 0002, Peng Gui, Dengyi Zhang |
Neurocomputing | 2 |
| 2023 | RHViT: A Robust Hierarchical Transformer for 3D Multimodal Brain Tumor Segmentation Using Biased Masked Image Modeling Pre-trainingabstractAccurate brain tumor segmentation in medical image analysis is crucial for diagnosis and treatment planning. While computer-aided methods have shown promise, several challenges persist. Most existing methods struggle with smaller tumors, treating all regions uniformly. Additionally, they lack robustness when dealing with data corruption and handling missing modalities, common in clinical settings. In this paper, we present a robust hierarchical vision transformer (RHViT) for 3D multimodal brain tumor segmentation, employing an encoder-decoder structure. Our approach combines 3D convolutions and self-attention, offering efficient and effective training. 3D convolutions help capture local information and generate hierarchical features, improving tumor segmentation accuracy. To enhance robustness, we pre-train the encoder using masked image modeling (MIM). This pre-training equips the model to handle data corruption, resulting in improved segmentation even in challenging scenarios. Furthermore, we introduce a novel biased masking strategy during MIM to focus the model's attention on tumor regions. This facilitates better tumor representations and effective fusion of multimodal features. Importantly, our biased masking technique strengthens the model's resilience when dealing with incomplete multimodal data during testing, making it a practical choice. Extensive experiments confirm the superiority of our model over existing approaches. Qian Zhou 0001, Hua Zou 0002, Fei Luo 0004, Yishi Qiu |
BIBM | 2 |
| 2023 | Incomplete Multimodal Learning for Visual Acuity Prediction After Cataract Surgery Using Masked Self-Attention
Qian Zhou 0001, Hua Zou 0002 |
MICCAI (7) | 2 |
| 2023 | HeadPose-Softmax: Head pose adaptive curriculum learning loss for deep face recognition
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jinsheng Xiao, Chao Liang 0001, Zhen Han 0002, Hua Zou 0002 |
Pattern Recognit. | 7 |
| 2022 | Long-Tailed Multi-label Retinal Diseases Recognition via Relational Learning and Knowledge Distillation
Qian Zhou 0001, Hua Zou 0002, Zhongyuan Wang 0001 |
MICCAI (2) | 2 |
| 2022 | A lightweight network for vehicle detection based on embedded system
Huanhuan Wu, Yuantao Hua, Hua Zou 0002, Gang Ke |
J. Supercomput. | 3 |
| 2019 | Shadow Inpainting and Removal Using Generative Adversarial Networks with Slice ConvolutionsabstractAbstract In this paper, we propose a two‐stage top‐down and bottom‐up Generative Adversarial Networks (TBGANs) for shadow inpainting and removal which uses a novel top‐down encoder and a bottom‐up decoder with slice convolutions. These slice convolutions can effectively extract and restore the long‐range spatial information for either down‐sampling or up‐sampling. Different from the previous shadow removal methods based on deep learning, we propose to inpaint shadow to handle the possible dark shadows to achieve a coarse shadow‐removal image at the first stage, and then further recover the details and enhance the color and texture details with a non‐local block to explore both local and global inter‐dependencies of pixels at the second stage. With such a two‐stage coarse‐to‐fine processing, the overall effect of shadow removal is greatly improved, and the effect of color retention in non‐shaded areas is significant. By comparing with a variety of mainstream shadow removal methods, we demonstrate that our proposed method outperforms the state‐of‐the‐art methods. Jinjiang Wei, Chengjiang Long, Hua Zou 0002, Chunxia Xiao |
Comput. Graph. Forum | 3 |
| 2017 | Illumination Decomposition for Photograph With Multiple Light SourcesabstractIllumination decomposition for a single photograph is an important and challenging problem in image editing operation. In this paper, we present a novel coarse-to-fine strategy to perform illumination decomposition for photograph with multiple light sources. We first reconstruct the lighting environment of the image using the estimated geometry structure of the scene. With the position of lights, we detect the shadow regions as well as the highlights in the projected image for each light. Then, using the illumination cues from shadows, we estimate the coarse illumination decomposed image emitted by each light source. Finally, we present a light-aware illumination optimization model, which efficiently produces the finer illumination decomposition results, as well as recover the texture detail under the shadow. We validate our approach on a number of examples, and our method effectively decomposes the input image into multiple components corresponding to different light sources. Ling Zhang 0017, Qingan Yan, Zheng Liu 0004, Hua Zou 0002, Chunxia Xiao |
IEEE Trans. Image Process. | 4 |
| 2016 | Predicting potential side effects of drugs by recommender methods and ensemble learning
Wen Zhang 0008, Hua Zou 0002, Longqiang Luo, Qianchao Liu, Weijian Wu, Wenyi Xiao |
Neurocomputing | 2 |
| 2011 | Prediction of conformational B-cell epitopes from 3D structures by random forest with a distance-based featureabstractBACKGROUND: Antigen-antibody interactions are key events in immune system, which provide important clues to the immune processes and responses. In Antigen-antibody interactions, the specific sites on the antigens that are directly bound by the B-cell produced antibodies are well known as B-cell epitopes. The identification of epitopes is a hot topic in bioinformatics because of their potential use in the epitope-based drug design. Although most B-cell epitopes are discontinuous (or conformational), insufficient effort has been put into the conformational epitope prediction, and the performance of existing methods is far from satisfaction. RESULTS: In order to develop the high-accuracy model, we focus on some possible aspects concerning the prediction performance, including the impact of interior residues, different contributions of adjacent residues, and the imbalanced data which contain much more non-epitope residues than epitope residues. In order to address above issues, we take following strategies. Firstly, a concept of 'thick surface patch' instead of 'surface patch' is introduced to describe the local spatial context of each surface residue, which considers the impact of interior residue. The comparison between the thick surface patch and the surface patch shows that interior residues contribute to the recognition of epitopes. Secondly, statistical significance of the distance distribution difference between non-epitope patches and epitope patches is observed, thus an adjacent residue distance feature is presented, which reflects the unequal contributions of adjacent residues to the location of binding sites. Thirdly, a bootstrapping and voting procedure is adopted to deal with the imbalanced dataset. Based on the above ideas, we propose a new method to identify the B-cell conformational epitopes from 3D structures by combining conventional features and the proposed feature, and the random forest (RF) algorithm is used as the classification engine. The experiments show that our method can predict conformational B-cell epitopes with high accuracy. Evaluated by leave-one-out cross validation (LOOCV), our method achieves the mean AUC value of 0.633 for the benchmark bound dataset, and the mean AUC value of 0.654 for the benchmark unbound dataset. When compared with the state-of-the-art prediction models in the independent test, our method demonstrates comparable or better performance. CONCLUSIONS: Our method is demonstrated to be effective for the prediction of conformational epitopes. Based on the study, we develop a tool to predict the conformational epitopes from 3D structures, available at http://code.google.com/p/my-project-bpredictor/downloads/list. Wen Zhang 0008, Yi Xiong 0002, Hua Zou 0002, Xinghuo Ye, Juan Liu 0007 |
BMC Bioinform. | 4 |