EDBT 2026 Demo / reviewers in the wild / expert
Meilu Zhu
dblp:210/4739
· DBLP profile ↗
22ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-5563-7282ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedPD++: Enhanced Federated Open-Set Recognition with Parameter DisentanglementabstractAbstract Federated Learning (FL) typically operates in a closed-set setting where all test classes are known during training, limiting its applicability in real-world scenarios where models must handle emerging unknown classes. This leads to misclassification of unseen categories as known ones. To address this limitation, we introduce Federated Open-Set Recognition (FedOSR), a novel paradigm enabling distributed clients to collaboratively train models that classify known classes while detecting and rejecting unknown ones. However, FedOSR presents unique challenges: the inter-set interference between learning closed-set and open-set knowledge within each client, and the intra-set inconsistency arising from data heterogeneity across clients. These challenges fundamentally complicate the federated aggregation process, as divergent optimization objectives and heterogeneous data distributions lead to parameter misalignment during model aggregation. In this work, we propose FedPD++ , a parameter disentanglement guided framework that systematically addresses both challenges through coordinated client-server mechanisms. On the client side, Local Parameter Disentanglement (LPD) decouples each OSR model into task-specific closed-set and open-set subnetworks to prevent inter-set interference. We introduce a Dynamic Path Integral (DPI) score that robustly identifies task-relevant parameters by leveraging path integral stability, coupled with an Adaptive Soft Masking (ASM) strategy that creates flexible subnetworks with adaptive thresholds rather than rigid binary partitions. On the server side, Global Divide-and-Conquer Aggregation (GDCA) tackles intra-set inconsistency by partitioning each subnetwork into shared and specific components, then aligning corresponding parts across clients using optimal transport to eliminate parameter misalignment. To ensure stable aggregation, we integrate Sequential Batch-Norm Alignment (SBA) that leverages temporal batch normalization statistics from multiple clients. Extensive experiments on open-set classification and segmentation tasks demonstrate that FedPD++ consistently achieves significant performance improvements over state-of-the-art methods. Code is available at: https://github.com/CUHK-AIM-Group/FedPD Chen Yang 0026, Meilu Zhu, Yifan Liu 0010, Yixuan Yuan |
Int. J. Comput. Vis. | 2 |
| 2025 | FedBM: Stealing knowledge from pre-trained language models for heterogeneous federated learning
Meilu Zhu, Qiushi Yang, Zhifan Gao, Yixuan Yuan, Jun Liu 0007 |
Medical Image Anal. | 1 |
| 2025 | ParetoSSL: Pareto Semi-Supervised Learning With Bias-Aware Gradient Preferences for Fruit Yield EstimationabstractFruit counting is a fundamental task for fruit yield estimation. Though semi-supervised counting methods have received increased attention in recent years, due to the high data utilization of unlabeled data, they suffer from two limitations. Firstly, difficult weight selection is a limitation, as these methods rely on manually selected fixed weights for both the supervised learning loss and the consistency learning loss, resulting in limited performance. Secondly, biased pseudo-labeling is another limitation, as they may predict biased pseudo-labels that result in small consistency learning losses, leading to training being dominated by supervised learning with large losses. To tackle these two limitations, in this paper, we propose a novel method named ParetoSSL to automatically derive weights of losses from the perspective of multi-task learning. Specifically, ParetoSSL formulates a multi-objective optimization problem for weight derivation by maximizing the similarity between weighted gradients of losses and a customized gradient preference vector, in which, the vector can guide weight derivation. Moreover, to relieve the effect of pseudo-label biases on consistency learning, we propose a bias-aware gradient preference vector. This vector considers gradient biases brought by the pseudo-label biases, which will down-weight the supervised learning loss while high-weighting the consistency learning loss. Meanwhile, to improve the robustness of ParetoSSL, an inequality equation regarding the norm of the gradients of the consistency learning loss is designed to control the range of gradient biases. Extensive experiments are conducted on the Clustered-Fruit dataset and Fruit-2019 dataset to evaluate the effectiveness of ParetoSSL on semi-supervised counting. Experimental results show that our ParetoSSL is superior to state-of-the-art methods. Note to Practitioners—This work is motivated by the emerging need for semi-supervised counting methods in fruit yield estimation. The difficulty of selecting loss weights for training semi-supervised counting algorithms is exacerbated by the pseudo-label bias issue that pseudo-label biases mislead the weight derivation while maximizing the similarity between the weight gradient and gradient preference vectors. The proposed bias-aware gradient preference vectors help users derive loss weights automatically and save the time of choosing loss weights for model fine-tuning. The proposed ParetoSSL is generic as it can be employed as a fruit yield estimation component of crop management support systems, while at the same time being applied to counting frameworks of other objects. Xiaochun Mai, Meilu Zhu, Yixuan Yuan |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Progressive Distillation With Optimal Transport for Federated Incomplete Multi-Modal Learning of Brain Tumor SegmentationabstractMulti-modal Magnetic Resonance Imaging (MRI) provide sufficient complementary information for brain tumor segmentation, however, most current approaches rely on complete modalities and may collapse with incomplete modalities. Moreover, most existing endeavors focus on training with centralized databases, failing to make full use of distributed multi-silo datasets with rich patient data to learn a more robust brain tumor segmentation model. In this paper, considering the distributed training scenarios, we formulate Federated Incomplete Multi-modal Learning (FedIML) for brain tumor segmentation, and propose Progressive distiLlation with Optimal Transport (PLOT) framework to gradually train a modality robust segmentation model at each client and achieve compatible model aggregation at the server. Specifically, to remedy the issue of unstable local training caused by the random modality input, we present Modality Progressive Distillation (MPD), a multi-level knowledge distillation strategy guided by a modality routing mechanism. At each client, MPD provides a gradually learning course for a student model in an easy-to-hard manner to achieve a stable local training process. Moreover, to address the problem that the layer-wise knowledge from different models may contradict, at the server, we design Optimal Transport-guided Model Aggregation (OTMA) strategy, which yields a global alignment solution for model parameters via solving an optimal transport problem. OTMA can achieve a compatible parameter aggregation and boost the distributed training. Extensive experiments on the BraTS-2021 dataset demonstrate the effectiveness of the proposed framework over state-of-the-art methods. Qiushi Yang, Meilu Zhu, Yat Ming Peter Woo, Leanne Lai Chan, Yixuan Yuan |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | DEeR: Deviation Eliminating and Noise Regulating for Privacy-Preserving Federated Low-Rank AdaptationabstractIntegrating low-rank adaptation (LoRA) with federated learning (FL) has received widespread attention recently, aiming to adapt pretrained foundation models (FMs) to downstream medical tasks via privacy-preserving decentralized training. However, owing to the direct combination of LoRA and FL, current methods generally undergo two problems, i.e., aggregation deviation, and differential privacy (DP) noise amplification effect. To address these problems, we propose a novel privacy-preserving federated finetuning framework called Deviation Eliminating and Noise Regulating (DEeR). Specifically, we firstly theoretically prove that the necessary condition to eliminate aggregation deviation is guaranteeing the equivalence between LoRA parameters of clients. Based on the theoretical insight, a deviation eliminator is designed to utilize alternating minimization algorithm to iteratively optimize the zero-initialized and non-zero-initialized parameter matrices of LoRA, ensuring that aggregation deviation always be zeros during training. Furthermore, we also conduct an in-depth analysis of the noise amplification effect and find that this problem is mainly caused by the "linear relationship" between DP noise and LoRA parameters. To suppress the noise amplification effect, we propose a noise regulator that exploits two regulator factors to decouple relationship between DP and LoRA, thereby achieving robust privacy protection and excellent finetuning performance. Additionally, we perform comprehensive ablated experiments to verify the effectiveness of the deviation eliminator and noise regulator. DEeR shows better performance on public medical datasets in comparison with state-of-the-art approaches. The code is available at https://github.com/CUHK-AIM-Group/DEeR. Meilu Zhu, Axiu Mao, Jun Liu 0007, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2024 | Enhancing Clinical Information for Zero-Shot Medical Diagnosis by Prompting Large Language ModelabstractIn real clinical diagnosis workflow, unseen disease categories are commonly encountered, where most existing supervised deep learning methods are invalid to accurately recognize. Recent works utilizing large-scale image-report datasets to train vision-language models have witnessed impressive zero-shot capabilities, while they rely on high-quality diagnosis reports that are difficult to collect, especially on some rare diseases. In this work, we propose Bidirectional vision-language Clinical information Exploitation (BCE), a new paradigm towards superior generalized zero-shot learning for medical diagnosis by multi-modal information mining. To harvest sparse disease semantics in medical images, the Cross-modal Knowledge Interaction (CKI) is designed by matching the global textual information towards local visual representations, which encourages the model to capture dense correspondence from visual to textual information. Furthermore, instead of using category keywords as text prompts to yield fixed descriptions from large language models (LLM) in previous works, we propose a Modality-Guided model Tuning (MGT) to encourage the LLM to produce fine-grained clinical information conditioned on input visual information. MGT can efficiently update additional learnable parameters inserted into the LLM and dynamically adapt them to yield instance-aware clinical information. Finally, a Fine-grained text-image Alignment (FA) is present to provide reliable constraint for superior discrimination. Extensive experiments on various medical generalized zero-shot learning benchmarks demonstrate the superiority of the proposed framework. Qiushi Yang, Meilu Zhu, Yixuan Yuan |
BIBM | 2 |
| 2024 | Stealing Knowledge from Pre-trained Language Models for Federated Classifier Debiasing
Meilu Zhu, Qiushi Yang, Zhifan Gao, Jun Liu 0007, Yixuan Yuan |
MICCAI (10) | 1 |
| 2024 | Comprehensive learning and adaptive teaching: Distilling multi-modal knowledge for pathological glioma grading
Xiaohan Xing, Meilu Zhu, Zhen Chen 0013, Yixuan Yuan |
Medical Image Anal. | 2 |
| 2024 | CMCNet: Colorization-Aware Mix-Uncertainty-Adaptive Consistency Network for Semi-Supervised Fruit CountingabstractFruit counting is a fundamental and challenging task of automatic fruit yield estimation in the field of intelligent agriculture. In recent years, to relieve the burden of data annotation, semi-supervised counting methods have been studied. Though significant progress has been achieved, the state-of-the-art method estimates the uncertainty of binary segmentation to guide the consistency training of density maps, being prone to deficient uncertainty estimation. Moreover, the method treats pixels with different difficulty equally in each training iteration, being troubled by inflexible consistency training which results in high supervision loss at the beginning of training and even causes network collapse. To alleviate the above limitations, in this paper, we propose a novel semi-supervised counting method CMCNet for fruit counting. CMCNet designs image colorization as an auxiliary task to estimate the uncertainty for density map consistency. Note that this work is the first effort to utilize image colorization for uncertainty estimation in semi-supervised counting. To obtain accurate uncertainty estimation for density map consistency, CMCNet estimates density uncertainty on density maps to depict the difficulty of fruit pixels from the semantic perspective, while using image colorization for constructing colorization uncertainty to measure the difficulty of part of fruit pixels and background pixels from the visual perspective. Then we obtain a comprehensive uncertainty by mixing density uncertainty and colorization uncertainty. Further, we propose a mix-uncertainty-adaptive consistency (MUAC) module for consistency training of density maps. With mix-uncertainty, uncertainty distribution is estimated. By adaptively adjusting the uncertainty threshold, harder pixels will be selected first and easier ones will be added into consistency training gradually. To evaluate the effectiveness of CMCNet, extensive experiments are conducted on two fruit datasets. Experimental results show that our CMCNet is superior to state-of-the-art semi-supervised counting methods.Note to Practitioners—This work is motivated by the emerging need for semi-supervised counting methods in fruit yield estimation. The difficulty of training semi-supervised counting methods with unlabeled images is exacerbated by the noisy supervision issue that pseudo-labels of unlabeled images are noisy. The proposed colorization-aware uncertainty estimation strategy and mix-uncertainty-adaptive consistency approach help the fruit planter sufficiently utilize the information of a large amount of unlabeled data and save the annotation cost in fruit quantity estimation. The proposed method is generic as it can be employed as a fruit yield estimation component of crop management support systems, while at the same time being applied to counting frameworks of other objects. Xiaochun Mai, Meilu Zhu, Yixuan Yuan |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | MGIML: Cancer Grading With Incomplete Radiology-Pathology Data via Memory Learning and Gradient HomogenizationabstractTaking advantage of multi-modal radiology-pathology data with complementary clinical information for cancer grading is helpful for doctors to improve diagnosis efficiency and accuracy. However, radiology and pathology data have distinct acquisition difficulties and costs, which leads to incomplete-modality data being common in applications. In this work, we propose a Memory- and Gradient-guided Incomplete Modal-modal Learning (MGIML) framework for cancer grading with incomplete radiology-pathology data. Firstly, to remedy missing-modality information, we propose a Memory-driven Hetero-modality Complement (MH-Complete) scheme, which constructs modal-specific memory banks constrained by a coarse-grained memory boosting (CMB) loss to record generic radiology and pathology feature patterns, and develops a cross-modal memory reading strategy enhanced by a fine-grained memory consistency (FMC) loss to take missing-modality information from well-stored memories. Secondly, as gradient conflicts exist between missing-modality situations, we propose a Rotation-driven Gradient Homogenization (RG-Homogenize) scheme, which estimates instance-specific rotation matrices to smoothly change the feature-level gradient directions, and computes confidence-guided homogenization weights to dynamically balance gradient magnitudes. By simultaneously mitigating gradient direction and magnitude conflicts, this scheme well avoids the negative transfer and optimization imbalance problems. Extensive experiments on CPTAC-UCEC and CPTAC-PDA datasets show that the proposed MGIML framework performs favorably against state-of-the-art multi-modal methods on missing-modality situations. Pengyu Wang 0005, Huaqi Zhang, Meilu Zhu, Xi Jiang 0001, Harry Qin, Yixuan Yuan |
IEEE Trans. Medical Imaging | 3 |
| 2024 | FedOSS: Federated Open Set Recognition via Inter-Client Discrepancy and CollaborationabstractOpen set recognition (OSR) aims to accurately classify known diseases and recognize unseen diseases as the unknown class in medical scenarios. However, in existing OSR approaches, gathering data from distributed sites to construct large-scale centralized training datasets usually leads to high privacy and security risk, which could be alleviated elegantly via the popular cross-site training paradigm, federated learning (FL). To this end, we represent the first effort to formulate federated open set recognition (FedOSR), and meanwhile propose a novel Federated Open Set Synthesis (FedOSS) framework to address the core challenge of FedOSR: the unavailability of unknown samples for all anticipated clients during the training phase. The proposed FedOSS framework mainly leverages two modules, i.e., Discrete Unknown Sample Synthesis (DUSS) and Federated Open Space Sampling (FOSS), to generate virtual unknown samples for learning decision boundaries between known and unknown classes. Specifically, DUSS exploits inter-client knowledge inconsistency to recognize known samples near decision boundaries and then pushes them beyond decision boundaries to synthesize discrete virtual unknown samples. FOSS unites these generated unknown samples from different clients to estimate the class-conditional distributions of open data space near decision boundaries and further samples open data, thereby improving the diversity of virtual unknown samples. Additionally, we conduct comprehensive ablation experiments to verify the effectiveness of DUSS and FOSS. FedOSS shows superior performance on public medical datasets in comparison with state-of-the-art approaches. The source code is available at https://github.com/CityU-AIM-Group/FedOSS. Meilu Zhu, Jing Liao 0001, Jun Liu 0007, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2023 | FedPD: Federated Open Set Recognition with Parameter DisentanglementabstractExisting federated learning (FL) approaches are deployed under the unrealistic closed-set setting, with both training and testing classes belong to the same set, which makes the global model fail to identify the unseen classes as ‘unknown’. To this end, we aim to study a novel problem of federated open-set recognition (FedOSR), which learns an open-set recognition (OSR) model under federated paradigm such that it classifies seen classes while at the same time detects unknown classes. In this work, we propose a parameter disentanglement guided federated open-set recognition (FedPD) algorithm to address two core challenges of FedOSR: cross-client inter-set interference between learning closed-set and open-set knowledge and cross-client intra-set inconsistency by data heterogeneity. The proposed FedPD framework mainly leverages two modules, i.e., local parameter disentanglement (LPD) and global divide-and-conquer aggregation (GDCA), to first disentangle client OSR model into different subnetworks, then align the corresponding parts cross clients for matched model aggregation. Specifically, on the client side, LPD decouples an OSR model into a closed-set subnetwork and an open-set subnetwork by the task-related importance, thus preventing inter-set interference. On the server side, GDCA first partitions the two subnetworks into specific and shared parts, and subsequently aligns the corresponding parts through optimal transport to eliminate parameter misalignment. Extensive experiments on various datasets demonstrate the superior performance of our proposed method. Chen Yang 0026, Meilu Zhu, Yifan Liu 0010, Yixuan Yuan |
ICCV | 2 |
| 2023 | FedDM: Federated Weakly Supervised Segmentation via Annotation Calibration and Gradient De-ConflictingabstractWeakly supervised segmentation (WSS) aims to exploit weak forms of annotations to achieve the segmentation training, thereby reducing the burden on annotation. However, existing methods rely on large-scale centralized datasets, which are difficult to construct due to privacy concerns on medical data. Federated learning (FL) provides a cross-site training paradigm and shows great potential to address this problem. In this work, we represent the first effort to formulate federated weakly supervised segmentation (FedWSS) and propose a novel Federated Drift Mitigation (FedDM) framework to learn segmentation models across multiple sites without sharing their raw data. FedDM is devoted to solving two main challenges (i.e., local drift on client-side optimization and global drift on server-side aggregation) caused by weak supervision signals in FL setting via Collaborative Annotation Calibration (CAC) and Hierarchical Gradient De-conflicting (HGD). To mitigate the local drift, CAC customizes a distal peer and a proximal peer for each client via a Monte Carlo sampling strategy, and then employs inter-client knowledge agreement and disagreement to recognize clean labels and correct noisy labels, respectively. Moreover, in order to alleviate the global drift, HGD online builds a client hierarchy under the guidance of history gradient of the global model in each communication round. Through de-conflicting clients under the same parent nodes from bottom layers to top layers, HGD achieves robust gradient aggregation at the server side. Furthermore, we theoretically analyze FedDM and conduct extensive experiments on public datasets. The experimental results demonstrate the superior performance of our method compared with state-of-the-art approaches. The source code is available at https://github.com/CityU-AIM-Group/FedDM. Meilu Zhu, Zhen Chen 0013, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2022 | Discrepancy and Gradient-Guided Multi-modal Knowledge Distillation for Pathological Glioma Grading
Xiaohan Xing, Zhen Chen 0013, Meilu Zhu, Yuenan Hou, Zhifan Gao, Yixuan Yuan |
MICCAI (5) | 3 |
| 2022 | Instance importance-Aware graph convolutional network for 3D medical diagnosis
Zhen Chen 0013, Jie Liu 0044, Meilu Zhu, Yat Ming Peter Woo, Yixuan Yuan |
Medical Image Anal. | 3 |
| 2022 | Personalized Retrogress-Resilient Federated Learning Toward Imbalanced Medical DataabstractClinically oriented deep learning algorithms, combined with large-scale medical datasets, have significantly promoted computer-aided diagnosis. To address increasing ethical and privacy issues, Federated Learning (FL) adopts a distributed paradigm to collaboratively train models, rather than collecting samples from multiple institutions for centralized training. Despite intensive research on FL, two major challenges are still existing when applying FL in the real-world medical scenarios, including the performance degradation (i.e., retrogress) after each communication and the intractable class imbalance. Thus, in this paper, we propose a novel personalized FL framework to tackle these two problems. For the retrogress problem, we first devise a Progressive Fourier Aggregation (PFA) at the server side to gradually integrate parameters of client models in the frequency domain. Then, at the client side, we design a Deputy-Enhanced Transfer (DET) to smoothly transfer global knowledge to the personalized local model. For the class imbalance problem, we propose the Conjoint Prototype-Aligned (CPA) loss to facilitate the balanced optimization of the FL framework. Considering the inaccessibility of private local data to other participants in FL, the CPA loss calculates the global conjoint objective based on global imbalance, and then adjusts the client-side local training through the prototype-aligned refinement to eliminate the imbalance gap with such a balanced goal. Extensive experiments are performed on real-world dermoscopic and prostate MRI FL datasets. The experimental results demonstrate the advantages of our FL framework in real-world medical scenarios, by outperforming state-of-the-art FL methods with a large margin. The source code is available at https://github.com/CityU-AIM-Group/PRR-Imbalancehttps://github.com/CityU-AIM-Group/PRR-Imbalance. Zhen Chen 0013, Chen Yang 0026, Meilu Zhu, Zhe Peng, Yixuan Yuan |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Personalized Retrogress-Resilient Framework for Real-World Medical Federated Learning
Zhen Chen 0013, Meilu Zhu, Chen Yang 0026, Yixuan Yuan |
MICCAI (3) | 2 |
| 2021 | Mutual-Prototype Adaptation for Cross-Domain Polyp SegmentationabstractAccurate segmentation of the polyps from colonoscopy images provides useful information for the diagnosis and treatment of colorectal cancer. Despite deep learning methods advance automatic polyp segmentation, their performance often degrades when applied to new data acquired from different scanners or sequences (target domain). As manual annotation is tedious and labor-intensive for new target domain, leveraging knowledge learned from the labeled source domain to promote the performance in the unlabeled target domain is highly demanded. In this work, we propose a mutual-prototype adaptation network to eliminate domain shifts in multi-centers and multi-devices colonoscopy images. We first devise a mutual-prototype alignment (MPA) module with the prototype relation function to refine features through self-domain and cross-domain information in a coarse-to-fine process. Then two auxiliary modules: progressive self-training (PST) and disentangled reconstruction (DR) are proposed to improve the segmentation performance. The PST module selects reliable pseudo labels through a novel uncertainty guided self-training loss to obtain accurate prototypes in the target domain. The DR module reconstructs original images jointly utilizing prediction results and private prototypes to maintain semantic consistency and provide complement supervision information. We extensively evaluate the proposed model in polyp segmentation performance on three conventional colonoscopy datasets: CVC-DB, Kvasir-SEG, and ETIS-Larib. The comprehensive experimental results demonstrate that the proposed model outperforms state-of-the-art methods. Chen Yang 0026, Xiaoqing Guo, Meilu Zhu, Bulat Ibragimov, Yixuan Yuan |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | DSI-Net: Deep Synergistic Interaction Network for Joint Classification and Segmentation With Endoscope ImagesabstractAutomatic classification and segmentation of wireless capsule endoscope (WCE) images are two clinically significant and relevant tasks in a computer-aided diagnosis system for gastrointestinal diseases. Most of existing approaches, however, considered these two tasks individually and ignored their complementary information, leading to limited performance. To overcome this bottleneck, we propose a deep synergistic interaction network (DSI-Net) for joint classification and segmentation with WCE images, which mainly consists of the classification branch (C-Branch), the coarse segmentation (CS-Branch) and the fine segmentation branches (FS-Branch). In order to facilitate the classification task with the segmentation knowledge, a lesion location mining (LLM) module is devised in C-Branch to accurately highlight lesion regions through mining neglected lesion areas and erasing misclassified background areas. To assist the segmentation task with the classification prior, we propose a category-guided feature generation (CFG) module in FS-Branch to improve pixel representation by leveraging the category prototypes of C-Branch to obtain the category-aware features. In such way, these modules enable the deep synergistic interaction between these two tasks. In addition, we introduce a task interaction loss to enhance the mutual supervision between the classification and segmentation tasks and guarantee the consistency of their predictions. Relying on the proposed deep synergistic interaction mechanism, DSI-Net achieves superior classification and segmentation performance on public dataset in comparison with state-of-the-art methods. The source code is available at https://github.com/CityU-AIM-Group/DSI-Net. Meilu Zhu, Zhen Chen 0013, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2019 | Robust Facial Landmark Detection via Occlusion-Adaptive Deep NetworksabstractIn this paper, we present a simple and effective framework called Occlusion-adaptive Deep Networks (ODN) with the purpose of solving the occlusion problem for facial landmark detection. In this model, the occlusion probability of each position in high-level features are inferred by a distillation module that can be learnt automatically in the process of estimating the relationship between facial appearance and facial shape. The occlusion probability serves as the adaptive weight on high-level features to reduce the impact of occlusion and obtain clean feature representation. Nevertheless, the clean feature representation cannot represent the holistic face due to the missing semantic features. To obtain exhaustive and complete feature representation, it is vital that we leverage a low-rank learning module to recover lost features. Considering that facial geometric characteristics are conducive to the low-rank module to recover lost features, we propose a geometry-aware module to excavate geometric relationships between different facial components. Depending on the synergistic effect of three modules, the proposed network achieves better performance in comparison to state-of-the-art methods on challenging benchmark datasets. Meilu Zhu, Daming Shi 0001, Mingjie Zheng 0002, Muhammad Sadiq |
CVPR | 1 |
| 2019 | Deep Geometry Embedding Networks for Robust Facial Landmark DetectionabstractFacial landmark detection has witnessed substantial progress due to introducing convolutional neural networks. Nonetheless, current convolutional neural networks-based approaches ignore the useful geometric relationship between different facial locations. To address this issue, we propose a new module to model the facial geometric relationship. The module can be integrated into the convolutional neural networks architecture to obtain the geometric representation, whereafter we leverage bilinear pooling operation to embed it into high-level feature maps of original face image so as to produce the more discriminative face representation. Extensive evaluation experiments on multiple challenging benchmark datasets demonstrate that our captured geometric information is robust against occlusion and head pose variation and our proposed method outperforms state-of-the-art methods. Meilu Zhu, Daming Shi 0001 |
ICME | 1 |
| 2019 | Branched convolutional neural networks incorporated with Jacobian deep regression for facial landmark detection
Meilu Zhu, Daming Shi 0001, Junbin Gao |
Neural Networks | 1 |