EDBT 2026 Demo / reviewers in the wild / expert
Dongmei Fu
dblp:10/4939
· DBLP profile ↗
36ranked-venue papers
1as first author
23since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Multi-view hypergraph networks with aggregation and estrangement dependencies for multi-source spatiotemporal forecasting
Xiaomeng Wu, Dongmei Fu, Tao Yang 0010, Lizhen Shao |
Expert Syst. Appl. | 2 |
| 2026 | MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and ClassificationabstractNucleus detection and classification (NDC) in histopathology analysis is a fundamental task that underpins a wide range of high-level pathology applications. However, existing methods heavily rely on labor-intensive nucleus-level annotations and struggle to fully exploit large-scale unlabeled data for learning discriminative nucleus representations. In this work, we propose MUSE (MUlti-scale denSE self-distillation), a novel self-supervised learning method tailored for NDC. At its core is NuLo (Nucleus-based Local self-distillation), a coordinate-guided mechanism that enables flexible local self-distillation based on predicted nucleus positions. By removing the need for strict spatial alignment between augmented views, NuLo allows critical cross-scale alignment, thus unlocking the capacity of models for fine-grained nucleus-level representation. To support MUSE, we design a simple yet effective encoder-decoder architecture and a large field-of-view semi-supervised fine-tuning strategy that together maximize the value of unlabeled pathology images. Extensive experiments on three widely used benchmarks demonstrate that MUSE effectively addresses the core challenges of histopathological NDC. The resulting models not only surpass state-of-the-art supervised baselines but also outperform generic pathology foundation models. Zijiang Yang 0009, Hanqing Chao, Bokai Zhao, Yelin Yang, Yunshuo Zhang, Dongmei Fu, Junping Zhang, Le Lu 0001, Ke Yan 0006, Dakai Jin, Minfeng Xu, Yun Bian |
AAAI | 6 |
| 2026 | Shape-aware medical image segmentation via frequency domain partitioningabstractPrecise medical image segmentation is vital for computer-aided diagnosis, yet current methods struggle with subtle endoscopic areas where lesions and normal tissue appear similar. To address this, we propose a shape-aware partitioning model with a dual-branch architecture. Its high-frequency branch captures edges and fine details, while the low-frequency branch focuses on overall shape and color distribution. The proposed model integrates these features via a hybrid decoder and a chimeric wavelet block, facilitating continuous bilateral information interaction. We also introduce a dual-domain loss function to comprehensively evaluate model output against ground truth, especially when pixel value differences are small but frequency domain differences are significant. The proposed method markedly enhances the accuracy and efficiency of computer-aided diagnosis, particularly in polyp and skin lesion segmentation. By accurately capturing lesion shape and volume, it provides a robust tool crucial for disease grading and treatment planning. Moreover, it outperforms comparable hybrid architectures integrating convolutional neural networks and transformers on public endoscopic polyp segmentation benchmarks. Quantitatively, it achieves a 15.29% higher intersection over union than conventional hybrid networks with a 33.75 giga floating-point operations reduction. Furthermore, it shows a 5.93% improvement over hierarchical hybrid models with an 11.74 giga floating-point operations decrease. Jiayuan Huang, Dongmei Fu, Chuanjiang Qi |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Conditional disentangled information bottleneck for multi-view molecular graph representation learning
Jiaxin Dai, Dongmei Fu, Zhongwei Qiu, Lingwei Ma |
Expert Syst. Appl. | 2 |
| 2026 | Graph-enhanced & monotonic embeddings: A novel approach to tabular data representation
Xiaomeng Wu, Dongmei Fu |
Neurocomputing | 2 |
| 2026 | Taylor-Sensus Network: Embracing Noise to Enlighten Uncertainty for Scientific DataabstractUncertainty estimation is vital for machine learning with scientific data. Although existing methods effectively address inherent model uncertainty, they frequently overlook explicit modeling of complex noise in data devoid of temporal or spatial dependencies. This gap is especially challenging in structured scientific data, where such dependencies are commonly lacking. To address these challenges in scientific research, we propose the Taylor-Sensus Network (TSNet). TSNet innovatively uses a Taylor series expansion to model complex, heteroscedastic noise and proposes a deep Taylor block for aware noise distribution. TSNet includes a noise-aware contrastive learning module and a data density perception module for aleatoric and epistemic uncertainty. Additionally, an uncertainty combination operator is used to integrate these uncertainties, and the network is trained using a novel heteroscedastic mean square error loss. TSNet demonstrates superior performance over mainstream and state-of-the-art methods in experiments, highlighting its potential in scientific research and noise resistance. Guangxuan Song, Dongmei Fu, Zhongwei Qiu, Jintao Meng 0003 |
IEEE Trans. Big Data | 2 |
| 2025 | Knowledge graph information bottleneck enhanced molecular representation learning
Jiaxin Dai, Dongmei Fu, Zhongwei Qiu, Lingwei Ma |
Neural Networks | 2 |
| 2025 | VLAB: Enhancing Video Language Pretraining by Feature Adapting and BlendingabstractLarge-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited research on learning video-text representations for general video multimodal tasks based on these powerful features. Towards this goal, we propose a novel video-text pre-training method dubbed VLAB:VideoLanguage pre-training by featureAdapting andBlending, which transfers CLIP representations to video pre-training tasks and develops unified video multimodal models for a wide range of video-text tasks. Specifically, VLAB is founded on two key strategies: feature adapting and feature blending. In the former, we introduce a new video adapter module to address CLIP's deficiency in modeling temporal information and extend the model's capability to encompass both contrastive and generative tasks. In the latter, we propose an end-to-end training method that further enhances the model's performance by exploiting the complementarity of image and video features. We validate the effectiveness and versatility of VLAB through extensive experiments on highly competitive video multimodal tasks, including video text retrieval, video captioning, and video question answering. Remarkably, VLAB outperforms competing methods significantly and sets new records in video question answering on MSRVTT, MSVD, and TGIF datasets. It achieves an accuracy of 49.6, 60.9, and 79.0, respectively. Xingjian He, Fan Ma, Zhicheng Huang 0002, Xiaojie Jin 0004, Dongmei Fu, Yi Yang 0001, Jing Liu 0001, Jiashi Feng |
IEEE Trans. Multim. | 7 |
| 2025 | MM-NeRF: Multimodal-Guided 3D Multi-Style Transfer of Neural Radiance Fieldabstract3D style transfer aims to generate stylized views of 3D scenes with specified styles, which requires high-quality generating and keeping multi-view consistency. Existing methods still suffer the challenges of high-quality stylization with texture details and stylization with multimodal guidance. In this paper, we reveal that the common training method of stylization with NeRF, which generates stylized multi-view supervision by 2D style transfer models, causes the same object in supervision to show various states (color tone, details, etc.) in different views, leading NeRF to tend to smooth the texture details, further resulting in low-quality rendering for 3D multi-style transfer. To tackle these problems, we propose a novel Multimodal-guided 3D Multi-style transfer of NeRF, termed MM-NeRF. First, MM-NeRF projects multimodal guidance into a unified space to keep the multimodal styles consistency and extracts multimodal features to guide the 3D stylization. Second, a novel multi-head learning scheme is proposed to relieve the difficulty of learning multi-style transfer, and a multi-view style consistent loss is proposed to track the inconsistency of multi-view supervision data. Finally, a novel incremental learning mechanism is proposed to generalize MM-NeRF to any new style with small costs. Extensive experiments on several real-world datasets show that MM-NeRF achieves high-quality 3D multi-style stylization with multimodal guidance, and keeps multi-view consistency and style consistency between multimodal guidance. Zijiang Yang 0009, Zhongwei Qiu, Chang Xu 0002, Dongmei Fu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | PixelLM: Pixel Reasoning with Large Multimodal ModelabstractWhile large multimodal models (LMMs) have achieved remarkable progress, generating pixel-level masks for image reasoning tasks involving multiple open-world targets remains a challenge. To bridge this gap, we introduce PixelLM, an effective and efficient LMM for pixel-level reasoning and understanding. Central to PixelLM is a novel, lightweight pixel decoder and a comprehensive segmentation codebook. The decoder efficiently produces masks from the hidden embeddings of the codebook tokens, which encode detailed target-relevant information. With this design, PixelLM harmonizes with the structure of popular LMMs and avoids the need for additional costly segmentation models. Furthermore, we propose a target refinement loss to enhance the model's ability to differentiate between multiple targets, leading to substantially improved mask quality. To advance research in this area, we construct MUSE, a high-quality multi-target reasoning segmentation benchmark. PixelLM excels across various pixel-level image reasoning and understanding tasks, outperforming well-established methods in multiple benchmarks, including MUSE, single- and multi-referring segmentation. Comprehensive ablations confirm the efficacy of each proposed component. All code, models, and datasets will be publicly available. Zhongwei Ren, Zhicheng Huang 0002, Yunchao Wei, Yao Zhao 0001, Dongmei Fu, Jiashi Feng, Xiaojie Jin 0004 |
CVPR | 5 |
| 2024 | Multi-scale contrastive adaptor learning for segmenting anything in underperformed scenes
Zhongwei Qiu, Dongmei Fu |
Neurocomputing | 3 |
| 2024 | Contrastive Masked Autoencoders are Stronger Vision LearnersabstractMasked image modeling (MIM) has achieved promising results on various vision tasks. However, the limited discriminability of learned representation manifests there is still plenty to go for making a stronger vision learner. Towards this goal, we propose Contrastive Masked Autoencoders (CMAE), a new self-supervised pre-training method for learning more comprehensive and capable vision representations. By elaboratively unifying contrastive learning (CL) and masked image model (MIM) through novel designs, CMAE leverages their respective advantages and learns representations with both strong instance discriminability and local perceptibility. Specifically, CMAE consists of two branches where the online branch is an asymmetric encoder-decoder and the momentum branch is a momentum updated encoder. During training, the online encoder reconstructs original images from latent representations of masked images to learn holistic features. The momentum encoder, fed with the full images, enhances the feature discriminability via contrastive learning with its online counterpart. To make CL compatible with MIM, CMAE introduces two new components, i.e., pixel shifting for generating plausible positive views and feature decoder for complementing features of contrastive pairs. Thanks to these novel designs, CMAE effectively improves the representation quality and transfer performance over its MIM counterpart. CMAE achieves the state-of-the-art performance on highly competitive benchmarks of image classification, semantic segmentation and object detection. Notably, CMAE-Base achieves 85.3% top-1 accuracy on ImageNet and 52.5% mIoU on ADE20k, surpassing previous best results by 0.7% and 1.8% respectively. Zhicheng Huang 0002, Xiaojie Jin 0004, Chengze Lu, Qibin Hou, Ming-Ming Cheng, Dongmei Fu, Xiaohui Shen, Jiashi Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | DMIS: Dynamic Mesh-Based Importance Sampling for Training Physics-Informed Neural NetworksabstractModeling dynamics in the form of partial differential equations (PDEs) is an effectual way to understand real-world physics processes. For complex physics systems, analytical solutions are not available and numerical solutions are widely-used. However, traditional numerical algorithms are computationally expensive and challenging in handling multiphysics systems. Recently, using neural networks to solve PDEs has made significant progress, called physics-informed neural networks (PINNs). PINNs encode physical laws into neural networks and learn the continuous solutions of PDEs. For the training of PINNs, existing methods suffer from the problems of inefficiency and unstable convergence, since the PDE residuals require calculating automatic differentiation. In this paper, we propose Dynamic Mesh-based Importance Sampling (DMIS) to tackle these problems. DMIS is a novel sampling scheme based on importance sampling, which constructs a dynamic triangular mesh to estimate sample weights efficiently. DMIS has broad applicability and can be easily integrated into existing methods. The evaluation of DMIS on three widely-used benchmarks shows that DMIS improves the convergence speed and accuracy in the meantime. Especially in solving the highly nonlinear Schrödinger Equation, compared with state-of-the-art methods, DMIS shows up to 46% smaller root mean square error and five times faster convergence speed. Code is available at https://github.com/MatrixBrain/DMIS. Zijiang Yang 0009, Zhongwei Qiu, Dongmei Fu |
AAAI | 3 |
| 2023 | PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation with Progressive Video TransformersabstractExisting methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then performs single-person PSE with temporal model. However, the global spatio-temporal context among spatial instances can not be captured. In this paper, we propose a new end-to-end multi-person 3D Pose and Shape estimation framework with progressive Video Transformer, termed PSVT. In PSVT, a spatio-temporal encoder (STE) captures the global feature dependencies among spatial objects. Then, spatio-temporal pose decoder (STPD) and shape decoder (STSD) capture the global dependencies between pose queries and feature tokens, shape queries and feature tokens, respectively. To handle the variances of objects as time proceeds, a novel scheme of progressive decoding is used to update pose and shape queries at each frame. Besides, we propose a novel pose-guided attention (PGA) for shape decoder to better predict shape parameters. The two components strengthen the decoder of PSVT to improve performance. Extensive experiments on the four datasets show that PSVT achieves stage-of-the-art results. Zhongwei Qiu, Qiansheng Yang, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Chang Xu 0002, Dongmei Fu, Jingdong Wang 0001 |
CVPR | 8 |
| 2023 | Learning Degradation-Robust Spatiotemporal Frequency-Transformer for Video Super-ResolutionabstractVideo Super-Resolution (VSR) aims to restore high-resolution (HR) videos from low-resolution (LR) videos. Existing VSR techniques usually recover HR frames by extracting pertinent textures from nearby frames with known degradation processes. Despite significant progress, grand challenges remain to effectively extract and transmit high-quality textures from high-degraded low-quality sequences, such as blur, additive noises, and compression artifacts. This work proposes a novel degradation-robust Frequency-Transformer (FTVSR++) for handling low-quality videos that carry out self-attention in a combined space-time-frequency domain. First, video frames are split into patches and each patch is transformed into spectral maps in which each channel represents a frequency band. It permits a fine-grained self-attention on each frequency band so that real visual texture can be distinguished from artifacts. Second, a novel dual frequency attention (DFA) mechanism is proposed to capture the global and local frequency relations, which can handle different complicated degradation processes in real-world scenarios. Third, we explore different self-attention schemes for video processing in the frequency domain and discover that a "divided attention" which conducts joint space-frequency attention before applying temporal-frequency attention, leads to the best video enhancement quality. Extensive experiments on three widely-used VSR datasets show that FTVSR++ outperforms state-of-the-art methods on different low-quality videos with clear visual margins. Zhongwei Qiu, Huan Yang 0005, Jianlong Fu, Daochang Liu, Chang Xu 0002, Dongmei Fu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Weakly-supervised pre-training for 3D human pose estimation via perspective knowledgeabstractModern deep learning-based 3D pose estimation approaches require plenty of 3D pose annotations. However, existing 3D datasets lack diversity, which limits the performance of current methods and their generalization ability. Although existing methods utilize 2D pose annotations to help 3D pose estimation, they mainly focus on extracting 2D structural constraints from 2D poses, ignoring the 3D information hidden in the images. In this paper, we propose a novel method to extract weak 3D information directly from 2D images without 3D pose supervision. Firstly, we utilize 2D pose annotations and perspective prior knowledge to generate the relative depth of human joints . Then, we collect a 2D pose dataset (MCPC) and generate relative depth labels. Based on MCPC, we propose a weakly-supervised pre-training (WSP) strategy to distinguish the depth relationship between two points in an image. WSP enables the learning of the relative depth of two keypoints on lots of in-the-wild images, which is more capable of predicting depth and generalization ability for 3D human pose estimation. After fine-tuning the pose model on 3D pose datasets, WSP achieves state-of-the-art results on two widely-used benchmarks. Zhongwei Qiu, Kai Qiu 0001, Jianlong Fu, Dongmei Fu |
Pattern Recognit. | 4 |
| 2022 | Learning Spatiotemporal Frequency-Transformer for Compressed Video Super-Resolution
Zhongwei Qiu, Huan Yang 0005, Jianlong Fu, Dongmei Fu |
ECCV (18) | 4 |
| 2022 | Dynamic Graph Reasoning for Multi-person 3D Pose EstimationabstractMulti-person 3D pose estimation is a challenging task because of occlusion and depth ambiguity, especially in the cases of crowd scenes. To solve these problems, most existing methods explore modeling body context cues by enhancing feature representation with graph neural networks or adding structural constraints. However, these methods are not robust for their single-root formulation that decoding 3D poses from a root node with a pre-defined graph. In this paper, we propose GR-M3D, which models the Multi-person 3D pose estimation with dynamic Graph Reasoning. The decoding graph in GR-M3D is predicted instead of pre-defined. In particular, It firstly generates several data maps and enhances them with a scale and depth aware refinement module (SDAR). Then multiple root keypoints and dense decoding paths for each person are estimated from these data maps. Based on them, dynamic decoding graphs are built by assigning path weights to the decoding paths, while the path weights are inferred from those enhanced data maps. And this process is named dynamic graph reasoning (DGR). Finally, the 3D poses are decoded according to dynamic decoding graphs for each detected person. GR-M3D can adjust the structure of the decoding graph implicitly by adopting soft path weights according to input data, which makes the decoding graphs be adaptive to different input persons to the best extent and more capable of handling occlusion and depth ambiguity than previous methods. We empirically show that the proposed bottom-up approach even outperforms top-down methods and achieves state-of-the-art results on three 3D pose datasets. Zhongwei Qiu, Qiansheng Yang, Jian Wang 0066, Dongmei Fu |
ACM Multimedia | 4 |
| 2022 | IVT: An End-to-End Instance-guided Video Transformer for 3D Pose EstimationabstractVideo 3D human pose estimation aims to localize the 3D coordinates of human joints from videos. Recent transformer-based approaches focus on capturing the spatiotemporal information from sequential 2D poses, which cannot model the contextual depth feature effectively since the visual depth features are lost in the step of 2D pose estimation. In this paper, we simplify the paradigm into an end-to-end framework, Instance-guided Video Transformer (IVT), which enables learning spatiotemporal contextual depth information from visual features effectively and predicts 3D poses directly from video frames. In particular, we firstly formulate video frames as a series of instance-guided tokens and each token is in charge of predicting the 3D pose of a human instance. These tokens contain body structure information since they are extracted by the guidance of joint offsets from the human center to the corresponding body joints. Then, these tokens are sent into IVT for learning spatiotemporal contextual depth. In addition, we propose a cross-scale instance-guided attention mechanism to handle the variational scales among multiple persons. Finally, the 3D poses of each person are decoded from instance-guided tokens by coordinate regression. Experiments on three widely-used 3D pose estimation benchmarks show that the proposed IVT achieves state-of-the-art performances. Zhongwei Qiu, Qiansheng Yang, Jian Wang 0066, Dongmei Fu |
ACM Multimedia | 4 |
| 2021 | Seeing Out of the Box: End-to-End Pre-Training for Vision-Language Representation LearningabstractWe study joint learning of Convolutional Neural Network (CNN) and Transformer for vision-language pre-training (VLPT) which aims to learn cross-modal alignments from millions of image-text pairs. State-of-the-art approaches extract salient image regions and align regions with words step-by-step. As region-based visual features usually represent parts of an image, it is challenging for existing vision-language models to fully understand the semantics from paired natural languages. In this paper, we propose SOHO to "Seeing Out of tHe bOx" that takes a whole image as input, and learns vision-language representation in an end-to-end manner. SOHO does not require bounding box annotations which enables inference 10 times faster than region-based approaches. In particular, SOHO learns to extract comprehensive yet compact image features through a visual dictionary (VD) that facilitates cross-modal understanding. VD is designed to represent consistent visual abstractions of similar semantics. It is updated on-the-fly and utilized in our proposed pre-training task Masked Visual Modeling (MVM). We conduct experiments on four well-established vision-language tasks by following standard VLPT settings. In particular, SOHO achieves absolute gains of 2.0% R@1 score on MSCOCO text retrieval 5k test split, 1.5% accuracy on NLVR2test-P split, 6.7% accuracy on SNLI-VE test split, respectively. Zhicheng Huang 0002, Zhaoyang Zeng, Yupan Huang, Bei Liu 0001, Dongmei Fu, Jianlong Fu |
CVPR | 5 |
| 2021 | A Non-autoregressive Decoding Model Based on Joint Classification for 3D Human Pose Regression
Dongmei Fu, Tao Yang 0010 |
PRCV (2) | 2 |
| 2021 | Hard exudate segmentation in retinal image with attention mechanismabstractAbstract Diabetic retinopathy (DR) is the main reason that causes preventable blindness. Hard exudate is one of the earliest signs of diabetic retinopathy. Precise detection of hard exudate is helpful for the early diagnosis of diabetic retinopathy. Fully convolutional network (FCN) shows great performance on hard exudate segmentation task. However, there are limitations for fully convolutional network to build long‐range dependencies in different regions of the image. Convolution operator extract features in local area, segmentation results based on local features are likely to be wrong in some cases. Another channel attention method was proposed, and two different attention modules are used in the segmentation model. In this way, long‐range dependencies across different image regions are built efficiently in different stages of feature extraction. In addition, a new loss function is designed to deal with the data imbalance problem in hard exudate segmentation task. The proposed method was evaluated by two public datasets, and the comparative experiments show the effectiveness of the proposed method. Ze Si, Dongmei Fu, Zhicheng Huang 0002 |
IET Image Process. | 2 |
| 2021 | Non-fragile extended dissipative synchronization of Markov jump inertial neural networks: An event-triggered control strategy
Tian Fang, Shiyu Jiao, Dongmei Fu, Jing Wang 0071 |
Neurocomputing | 3 |
| 2020 | DGCN: Dynamic Graph Convolutional Network for Efficient Multi-Person Pose EstimationabstractMulti-person pose estimation aims to detect human keypoints from images with multiple persons. Bottom-up methods for multi-person pose estimation have attracted extensive attention, owing to the good balance between efficiency and accuracy. Recent bottom-up methods usually follow the principle of keypoints localization and grouping, where relations between keypoints are the keys to group keypoints. These relations spontaneously construct a graph of keypoints, where the edges represent the relations between two nodes (i.e., keypoints). Existing bottom-up methods mainly define relations by empirically picking out edges from this graph, while omitting edges that may contain useful semantic relations. In this paper, we propose a novel Dynamic Graph Convolutional Module (DGCM) to model rich relations in the keypoints graph. Specifically, we take into account all relations (all edges of the graph) and construct dynamic graphs to tolerate large variations of human pose. The DGCM is quite lightweight, which allows it to be stacked like a pyramid architecture and learn structural relations from multi-level features. Our network with single DGCM based on ResNet-50 achieves relative gains of 3.2% and 4.8% over state-of-the-art bottom-up methods on COCO keypoints and MPII dataset, respectively. Zhongwei Qiu, Kai Qiu 0001, Jianlong Fu, Dongmei Fu |
AAAI | 4 |
| 2020 | Belief-peaks clustering based on fuzzy label propagation
Jintao Meng 0003, Dongmei Fu |
Appl. Intell. | 2 |
| 2019 | U-Net with Attention Mechanism for Retinal Vessel Segmentation
Ze Si, Dongmei Fu |
ICIG (2) | 2 |
| 2019 | Learning Recurrent Structure-Guided Attention Network for Multi-person Pose EstimationabstractMulti-person pose estimation aims to localize tens of human joints (e.g., elbow, wrist, etc.) from multiple human bodies in an image. Existing approaches mainly adopt a two stage pipeline, which usually consists of a human detector (i.e., generating a bounding box for each person) and a single person pose estimator (i.e., generating human joints from each bounding box). However, these approaches neglect the challenges of large pose variations and heavy occlusions in each bounding box, which often results in imprecise human joint localization. In this paper, we propose a structure-guided attention network (SGAN) for multi-person pose estimation. Specifically, a structured pose representation is encoded by learning a joint confidence map and a joint association map, which can be further refined by a structure-guided attention network (SGAN) in a recurrent way. Note that SGAN enables a deep neural network to take initial pose estimation as references, and to discover multi-scale pose features as completion, and thus the learning of pose structures can be reinforced. Extensive experiments show the best single-model results against the state-of-the-art approaches, with a relative 3.5% mAP gain in the challenging COCO Keypoint dataset. Zhongwei Qiu, Kai Qiu 0001, Jianlong Fu, Dongmei Fu |
ICME | 4 |
| 2019 | Dempster-Shafer evidence theory-based multi-feature learning and fusion method for non-rigid 3D model retrievalabstractThis study introduces a novel multi‐feature‐based non‐rigid three‐dimensional (3D) model retrieval method. First, for each 3D model, compute the scale‐invariant heat kernel signature (SI‐HKS) descriptor and the wave kernel signature (WKS) descriptor of each vertex. Then, the normalised weighted bags of phrases feature is obtained and they are fed to the convolutional neural networks. The trust degree of each kind of descriptor is computed, and the total trust degree can be obtained. Finally, the fusion network is trained and the retrieval results can be obtained according to the ranking of the total trust degrees. For the training phase and the testing phase, the authors define different computation methods of the trust degrees and the total trust degrees. The Dempster–Shafer (DS) evidence‐based total trust degrees are used not only in the feature layer but also in the decision layer. The final decision results of the total trust degrees are used in the process of the network learning. So the proposed method can make full use of the complementary information of the SI‐HKS descriptor and the WKS descriptor. Extensive experiments have shown that the proposed multi‐feature fusion method has better performance than a single feature‐based method, and also outperforms other existing state‐of‐the‐art methods. Hui Zeng 0003, Xiuqing Wang, Dongmei Fu, Qingting Wei |
IET Comput. Vis. | 4 |
| 2019 | Optic disc segmentation in fundus images using adversarial trainingabstractGlaucoma is one of the leading causes of blindness in the world. Optic disc segmentation is an indispensable step for automatic detection of glaucoma with fundus images. In this study, the authors propose an automatic optic disc segmentation approach using adversarial training. The improved ‘U‐Net’ is used as the segmentation network to detect optic disc from fundus images, and then the authors add a ‘Patch‐level’ adversarial network to enhance higher‐order consistency between ground truth and the output from segmentation network, which further boosts the performance of segmentation network. In addition, a new loss function is designed to solve the problem of pixel‐level class imbalance in small target region extraction of medical images. All these improvements have effectively increased the segmentation accuracy on hard examples. Authors’ methods achieve Dice coefficient of 0.967 on Drishti‐GS dataset and 0.951 on RIM‐ONEv3 dataset, which outperform most of the existing methods. Dongmei Fu, Zhicheng Huang 0002, Hejun Tong |
IET Image Process. | 2 |
| 2018 | Semi-Supervised Soft Label Propagation Based on Mass Function for Community DetectionabstractWith the growing complexity of networks in practical applications, the accuracy and robustness of community detection approaches need to be improved. The semi-supervised label propagation (SLP) is known for its near-linear computation and making full use of the prior information. However, performance of existing SLP algorithms is seriously affected by outliers and the label information propagated in the process is categorical which leads to not completely accurate detection results. A semi-supervised soft label propagation based on mass function (SSLP) is proposed in this paper. We deal with the imprecise community assignment of each node by mass function which can be regarded as a soft label. And the mass assigned to the whole set of all communities measures the degree of outliers. Based on manifold assumption, the label of each node depends on its neighbors whose label information can be combined by Dempster's rule of combination. After convergence of SSLP, detection results will be given based on the mass matrix. Experiments on several real-world and artificial networks show that the proposed SSLP is competitive and even better. Jintao Meng 0003, Dongmei Fu, Tao Yang 0010 |
FUSION | 2 |
| 2018 | Classification of ADHD with bi-objective optimization
Lizhen Shao, Yadong Xu, Dongmei Fu |
J. Biomed. Informatics | 3 |
| 2017 | An Artery/Vein Classification Method Based on Color and Vascular Structure Information
Dongmei Fu, Haosen Ma |
ICIG (2) | 1 |
| 2017 | Adaptive Learning Compressive Tracking Based on Kalman Filter
Dongmei Fu, Yanan Shi, Chunhong Wu |
ICIG (3) | 2 |
| 2017 | Pedestrian tracking for infrared image sequence based on trajectory manifold of spatio-temporal slice
Tao Yang 0010, Dongmei Fu, Shu Pan |
Multim. Tools Appl. | 2 |
| 2016 | Laplacian Embedded Infinite Kernel Model for Semi-Supervised ClassificationabstractPromoted by its convexity and low time complexity, Laplacian embedded support vector regression (LapESVR) model based on manifold regularization (MR) has assumed an important role in semi-supervised classification. Conventionally, the LapESVR model is based on a single kernel function that is intrinsically capable of describing one feature mapping relation only. However, when the data to be processed is from a complex dataset where multiple features of the data are required to be treated, the classification performance using the LapESVR based on a single kernel substantially degrade, indicating that the classification requirement in this case is beyond the capability of the LapESVR. In addition, the processing data is often subject to the impact of abnormal data samples; therefore, in practice assigning a fixed value that is related to the average distance of the processing data as the parameter value of kernel function of the LapESVR is by no means optimal. To solve the problems as mentioned regarding the LapESVR, this paper proposes a Laplacian embedded infinite kernel regression (LapEIKR) model. The proposed model combines the multiple kernels linearly to improve its ability of characterization of the processing data, typical in semi-supervised classification of complex datasets, with multiple features. Further, the parameter setting of the multiple kernels of the LapEIKR model is turned into an optimization problem by formulating a corresponding minimum objective function and an iterative algorithm, and then the values of the settings are facilitated to be obtained by a formulated calculation, assuming the optimal values with respect to the designed objective function. Comparative experiments on the UCI datasets, benchmark datasets and Caltech256 datasets show that the proposed LapEIKR model is improving in terms of adaptivity and efficiency. Tao Yang 0010, Dongmei Fu, Chunhong Wu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2014 | Semi-supervised classification with Laplacian multiple kernel learning
Tao Yang 0010, Dongmei Fu |
Neurocomputing | 2 |