Nan Mu

dblp:137/4182 · DBLP profile ↗
← Back
48ranked-venue papers
15as first author
33since 2021 · last 2026
0000-0003-0476-7500ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 15 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 14 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 VFGS-Net: Frequency-Guided State-Space Learning for Topology-Preserving Retinal Vessel Segmentation
abstract
Accurate retinal vessel segmentation is a critical prerequisite for quantitative analysis of retinal images and computer-aided diagnosis of vascular diseases such as diabetic retinopathy. However, the elongated morphology, wide scale variation, and low contrast of retinal vessels pose significant challenges for existing methods, making it difficult to simultaneously preserve fine capillaries and maintain global topological continuity. To address these challenges, we propose the Vessel-aware Frequency-domain and Global Spatial modeling Network (VFGS-Net), an end-to-end segmentation framework that seamlessly integrates frequency-aware feature enhancement, dual-path convolutional representation learning, and bidirectional asymmetric spatial state-space modeling within a unified architecture. Specifically, VFGS-Net employs a dual-path feature convolution module to jointly capture fine-grained local textures and multi-scale contextual semantics. A novel vessel-aware frequency-domain channel attention mechanism is introduced to adaptively reweight spectral components, thereby enhancing vessel-relevant responses in high-level features. Furthermore, at the network bottleneck, we propose a bidirectional asymmetric Mamba2-based spatial modeling block to efficiently capture long-range spatial dependencies and strengthen the global continuity of vascular structures. Extensive experiments on four publicly available retinal vessel datasets demonstrate that VFGS-Net achieves competitive or superior performance compared to state-of-the-art methods. Notably, our model consistently improves segmentation accuracy for fine vessels, complex branching patterns, and low-contrast regions, highlighting its robustness and clinical potential.
Nan Mu
ICMR6
2025 Fusion meets Function: The Adaptive Selection-Generation Approach in Event Argument Extraction
abstract
Event Argument Extraction is a critical task of Event Extraction, focused on identifying event arguments within text. This paper presents a novel Fusion Selection-Generation-Based Approach, by combining the precision of selective methods with the semantic generation capability of generative methods to enhance argument extraction accuracy. This synergistic integration, achieved through fusion prompt, element-based extraction, and fusion learning, addresses the challenges of input, process, and output fusion, effectively blending the unique characteristics of both methods into a cohesive model. Comprehensive evaluations on the RAMS and WikiEvents demonstrate the model’s state-of-the-art performance and efficiency.
Guoxuan Ding, Tianshu Fu, Nan Mu, Daren Zha
COLING6
2025 A Diffusion Model over Directed Acyclic Graphs for Event Schema Generation
abstract
Event schema generation is crucial for understanding the structure and temporal relationships of complex events. In this paper, we introduce a novel Directed Acyclic Graph Diffusion Model (DAGDM) that integrates DAG characteristics within a diffusion framework to enhance the effectiveness of schema generation. Our method leverages DAG positional embeddings to capture the hierarchical structure of nodes within graphs, while employing a reachability-based attention to better extract structural relationships between events. To this end, we design a cross-generation strategy that separately generates event sequence and adjacency matrix. Experiments show that our model effectively captures long-range event sequences, significantly enhancing schema generation for complex events.1
Guoxuan Ding, Haotian Jin, Nan Mu, Daren Zha
ICASSP5
2025 PMA-Net: Parallel Mixed Attention Network for Predicting Intracranial Aneurysm Rupture Risk
abstract
Intracranial aneurysm (IA) is a life-threatening condition with high morbidity and mortality rates. Since preventive treatment of IA also carries inherent risks, accurate rupture risk prediction is crucial for optimizing clinical decisionmaking. However, current methods for IA rupture risk prediction remain limited in accuracy, generalization, and interpretability. To address the above issues, this study proposes PMA-Net, a novel framework based on a parallel mixed attention mechanism, to efficiently and accurately predict IA rupture risk from Computed Tomography Angiography (CTA) images. By jointly analyzing aneurysms and their surrounding regions, the model captures interdependent imaging features associated with rupture risk. A multi-branch attention module extracts both coarse-and fine-grained features, while clinical information is integrated to enhance prediction performance. In addition, interpretable visualization methods were employed to enhance the model interpretability. Experimental results show that the proposed method achieves superior accuracy, improving by 2.22 % and 1.08 % on internal and external test sets, reaching 92.22 % and 84.78 %, respectively.
Fuhao Zhang, Jingfeng Jiang, Nan Mu
ICTAI6
2025 Towards Robust Polyp Segmentation: Multi-Focus Attention Network with Fine-grained Polyp Cues
abstract
Colorectal cancer (CRC) is one of the prominent causes of cancer-related morbidity and mortality worldwide. More AI-assisted methods are conducted for early polyp detection and segmentation to improve the screening efficacy. However, previous solutions generally exhibit weak segmentation performance due to irregular structures of polyps, while the model robustness suffers from background noise of homogeneous neighbors. To this end, we propose a novel Multi-Focus Attention Network (MFANet) to encode multi-dimensional information (i.e., scale, contour, and shape) as fine-grained cues for polyp segmentation. Concretely, a Scale-Residual-Aware Attention (SRAA) is designed to apply the residual operation over each layer of the feature pyramid architecture, which could minimize the feature interference among different scales. To improve the model robustness, a Geometry-Structure-Aware Attention (GSAA) is formulated to integrate and refine multi-dimensional geometric features via a Channel-Wise Enhance Attention (CWEA), which condenses the spatial information and recalibrates the channel importance for adaptive feature recalibration. Experiments on six public datasets indicate the effectiveness of the proposed method. Notably, on the more challenging BKAI dataset, which is featured by tiny polyps with serious interference of homogeneous neighboring region, our MFANet can outperform the state-of-the-art (SOTA) methods. Additionally, it is experimentally verified that our approach consistently exhibits better segmentation performance with higher robustness against different attack strategies (i.e., FGSM, WaNet and PGD).
Nan Mu, Xianchao Zhang 0004, Yazhou Feng, Jingfeng Jiang
ICMR1
2025 EdgeSegDiff: Edge-Conditional Diffusion Model for Skin Lesion Segmentation
Tan Pan, Nan Mu
PRCV (14)3
2025 Joint Geometric Self-Attention and Boundary-Aware Search for High-Precision Intracranial Aneurysm Mesh Segmentation
Fuhao Zhang, Ling Wang 0005, Dapeng Chen, Jinshan Tang, Jingfeng Jiang, Nan Mu
SMC9
2025 Progressive Multi-Scale Vision Transformer for Hierarchical Myocardial Segmentation in Cardiac MRI
abstract
Myocardial infarction remains a global health challenge. Accurate myocardial segmentation in late gadolinium enhancement cardiac magnetic resonance imaging (LGE-CMRI) is critical for diagnosis and treatment planning. Although deep learning architectures have demonstrated excellent segmentation performance in conventional CMRI, their accuracy significantly declines in LGE-CMRI due to low tissue contrast and complex background interference. To address these challenges, we propose a Progressive Multi-scale Vision Transformer (PMVT) for myocardial segmentation in LGE-CMRI, which enhances spatial representation capabilities and improves adaptability in complex scenarios through multi-scale feature fusion and interaction. Specifically, PMVT includes a Multi-scale Progressive Attention Decoder (MPSD) for modeling both long and short-term dependencies, and a Multi-layer Hybrid Context Purification (MHCP) that combines different combinations of four prediction heads for prediction and loss calculation, effectively suppressing background interference. Experiments demonstrate that the proposed PMVT outperforms state-of-the-art (SOTA) models, achieving a Dice score of 89.61% for myocardial segmentation (a 1.56% improvement over the current SOTA). This result highlights its considerable potential for clinical applications in automated LGE-CMRI analysis.
Lei Pu, Yangjie Li, Yuanwei Xu, Jingfeng Jiang, Jinshan Tang, Nan Mu
SMC8
2025 A Transformer-Based Dual-Branch Mesh Convolutional Neural Network for Aortic Dissection Segmentation
abstract
Aortic dissection (AD) is a life-threatening condition caused by a tear in the aortic intima, allowing blood to enter the vessel wall and form a false lumen. Due to its high mortality rate, timely diagnosis and precise treatment are critical. Clinical diagnosis and treatment of AD rely heavily on accurate 3D vascular image segmentation. To address existing methods’ low segmentation accuracy and insufficient geometric detail preservation, this paper proposes a Transformer-based Dual-Branch Mesh Segmentation Network (TD-MSeg) for AD. This network employs a mesh-based self-attention mechanism to retain vascular geometric details while adopting a dual-branch decoder to effectively fuse features and model long-range dependencies. Specifically, TD-MSeg incorporates three key components: a Hierarchical Mesh Transformer (HMT) module that enhances feature modeling of critical anatomical structures (e.g., intimal tears), a dual-branch decoder that facilitates collaborative optimization of multi-scale local and global features, and a mesh label refinement module that uses a wide-path exploration algorithm to eliminate deformation artifacts and improve spatial label continuity. Moreover, experiments on two AD mesh segmentation datasets demonstrate that the proposed TD-MSeg achieves a 6% improvement in accuracy compared to traditional models and significantly enhances the recognition of complex vascular structures, thereby providing high-precision 3D reconstruction support for endovascular surgical planning.
Fuhao Zhang, Ling Wang 0005, Dapeng Chen, Jinshan Tang, Jingfeng Jiang, Nan Mu
SMC9
2025 DMSA-Net: Deformable Multi-Head Self-Attention Network for Unsupervised Medical Image Registration
abstract
Medical image registration is fundamental to clinical analysis, enabling structural comparison and the detection of pathological changes across subjects or time. Despite advances in Transformer-based models, accurate 3D registration remains challenging due to difficulties in capturing multi-scale deformations and multi-directional spatial correlations. Moreover, standard Transformer architectures are computationally expensive when applied to 3D volumes. To address these issues, we propose DMSA-Net, an unsupervised medical image registration network featuring a novel 3D Deformable Multi-Head Self-Attention mechanism. Specifically, we design a Shifted Window Deformable Attention (SWDA) module that dynamically adjusts key-value pairs within localized 3D windows, enhancing modeling of multi-scale deformations while ensuring computational efficiency. We also introduce a Multi-directional Convolution (MC) module that uses large convolution kernels applied in parallel along the height, width and depth to capture spatially diverse deformation features. Experiments on two public brain MRI datasets (e.g., LPBA and IXI) demonstrate the superiority of DMSA-Net, showing a 4.1% improvement in Dice similarity coefficient over VoxelMorph and a 2.1% gain over TransMorph on LPBA. These results confirm the effectiveness of our approach for medical image registration.
Linxiao Zheng, Nan Mu
SMC2
2025 Spatial Bi-Exploration for Robust Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) aims to segment camouflaged objects hidden within their environment. Existing COD models, aside from image features, mostly focus on a single coarse-grained spatial structure, such as depth information, texture information, or edge information. However, when faced with complex scenes where the target and background textures are similar and overlapping, or when subjected to noise interference, this design often leads to insufficient detection accuracy and robustness. To address these issues, we proposed a strategy for multiple spatial explorations and designedSpatial Bi-Exploration Network (SPNet). SPNet conducts a comprehensive analysis of complex camouflage scenarios by jointly exploring depth spatial, contour spatial, and image feature information, thereby enhancing detection performance and maintaining robustness. Unlike existing methods, SPNet leverages dual exploration of depth and contour spaces to mitigate the vulnerability of coarse structures to noise. Depth spatial information aids the model in recognizing the deep relationships between objects and the background, reducing the impact of noise on object boundaries, while contour spatial information improves edge detection accuracy. This dual approach significantly enhances robustness, especially in the face of adversarial attacks. Extensive experiments on benchmark datasets demonstrate that our model not only outperforms existing methods in detection performance but also exhibits superior robustness against adversarial attacks.
Xiao Wang 0029, Xin Yuan 0009, Nan Mu, Zheng Wang 0007
IEEE Signal Process. Lett.4
2025 A Wavelet-Guided Deep Unfolding Network for Single Image Reflection Removal
abstract
Removing unwanted reflections from images is a fundamental yet challenging problem in low-level computer vision. Recent deep learning-based Single Image Reflection Removal (SIRR) methods have made significant progress. However, separating reflections from transmission content remains difficult, particularly in complex scenes where the two exhibit high visual similarity. Upon careful analysis, we find that reflections predominantly reside in the high-frequency components of an image. These reflections tend to distort fine details in the high-frequency range, while the low-frequency information remains relatively less affected. This observation motivates us to explore a frequency-aware approach for SIRR by leveraging the Discrete Wavelet Transform (DWT). The wavelet decomposition enables us to distinguish and isolate reflective artifacts in the frequency domain while preserving the transmission information. Building on this insight, we propose a novel Wavelet-guided Deep Unfolding Network (WDUNet) that leverages the strengths of wavelet decomposition and deep unfolding techniques to improve interpretability and generalization in SIRR. Specifically, we formulate an optimization-based reflection removal model using DWT and convolutional dictionaries. The proposed model is optimized via a proximal gradient algorithm and then unfolded into a neural network architecture, where all parameters are learned end-to-end during training. By combining wavelet domain analysis with deep unfolding, WDUNet enhances both the interpretability and generalization of SIRR methods. Additionally, we design and integrate the Low-frequency Parameter Estimation Module (LPEM) and High-frequency Parameter Estimation Module (HPEM) modules into WDUNet, allowing the network to automatically learn and optimize the models' hyperparameters. Extensive experiments conducted on four benchmark datasets demonstrate that WDUNet consistently outperforms existing state-of-the-art methods in both objective evaluation metrics and subjective visual quality.
Qiufu Li, Xu Wu 0001, Nan Mu, LinLin Shen
IEEE Trans. Image Process.4
2025 Context-aware target texture perturbation attack for concealed object detection
Kui Jiang, Nan Mu
Vis. Comput.5
2024 A Meta-pattern-enhanced Generative Few-shot Attribute Extraction Framework for Open-world Sparse Corpora
abstract
Open-world attribute extraction is one of the most important tasks of information extraction aiming to mine all the valuable attributes of entities and their corresponding values from unstructured texts, usually in the form of (entity, attribute, value) triplets. However, existing methods have difficulty extracting attribute triplets from open-world sparse corpora where the attribute names are not previously given, especially in the few- shot scenario with only few manual annotations available. To solve the above problems, we propose a two-stage Meta-pattern-Enhanced Generative Few-shot Attribute Extraction (MEGFAE) framework which can be used to discover utmost valuable attribute triplets from open-world sparse corpora in a generative manner. For evaluation on open-world sparse corpora, we introduce a benchmark dataset called OSN-51511The dataset is available in https://github.com/sunshower-liu/OSN-515.. Experimental results verifies the effectiveness of our framework and inspires future explorations on the text mining on sparse corpora.
Xiyu Liu 0003, Xin Wang 0086, Zeyi Liu 0002, Nan Mu, Tianshu Fu, Ji Xiang
MSN5
2024 3D Deformable Convolution for Medical Image Registration
Nan Mu, Haoyang Xing
PRICAI (5)2
2024 HDF-SegNet: Dynamic Integration of Hierarchical Information for Cardiac Magnetic Resonance Image Segmentation
Zijuan Wang, Nan Mu
PRICAI (5)3
2024 Video salient object detection via self-attention-guided multilayer cross-stack fusion
Nan Mu, Jinjia Guo, Yiyue Hu, Rong Wang 0006
Multim. Tools Appl.2
2023 Exploring a Self-Attentive Multilayer Cross-Stacking Fusion Model for Video Salient Object Detection
abstract
As an effective measure to capture the object of interest in video sequence, video salient object detection (VSOD) requires the processing of information from spatial-motion modalities, although plenty of traditional VSOD models were dedicated to developing efficient spatial and motion features to obtain salient objects of global consistency, the highly redundant spatial information brought by consecutive identical objects will inevitably reduce the generalization ability of these VSOD model. Although exploring the integration of spatial and motion information can improve the inter-frame correlation of salient objects to some extent, previous models tend to focus only on simple spatio-temporal fusion, which can also lead to the generation of redundant information, resulting in poor detection performance. Therefore, it is necessary to focus on effectively fusing the feature information of different modalities to eliminate the effect of redundant information. In this research, we proposed a self-attentive multilayer cross-stacking fusion based VSOD model, which productively extracts the multimodal features for two-way information transfer, fully utilizes the spatial and temporal knowledge to complement each other, and refines the cross-stacking of the interacted information and spatial features for local and global saliency optimization. As a result, the redundant spatial information can be largely eliminated, reducing the misidentification of salient objects due to blurred backgrounds or moving objects, and adaptively activating more weights of the salient object to achieve globally consistent saliency. Comprehensive experiments on four publicly available VSOD datasets demonstrated that the model had superior performance compared to the latest multiple VSOD models.
Nan Mu, Jinjia Guo, Rong Wang 0006
SMC2
2023 A Multi-Distance Feature Dissimilarity-Guided Encoder-Decoder Network for Polyp Segmentation
abstract
Most colorectal cancers originate from adenomatous polyps, which start as single asymptomatic polyps and develop into malignant tumors. In clinical practice, colonoscopy is an extremely effective method for detecting polyps, and it provides important visual information for the accurate identification and removal of polyps. However, it is highly challenging to achieve accurate segmentation of various polyps due to the complex and variable size, shape, color, number, and growth background of polyps at different periods. To address these dilemmas, we propose a multi-distance feature dissimilarity-guided encoder-decoder network for automatic polyp segmentation, mainly consisting of the Multi-Distance Differential Module (MDDM) and the Hybrid Loss Module (HLM). The former mainly utilizes the multilayer feature subtraction (MLFS) operations to extract the difference information between short-distance adjacent layer features and short-distance and long-distance cross-layer features. Given this, the pyramid-inspired MDDM obtains discriminative features continuously between adjacent/cross layers, enhancing complementary features between different layers. The latter supervises the feature maps extracted at each network level to achieve finer predictions. Experiments on four challenge datasets confirm that the proposed model outperforms most state-of-the-art methods in six evaluation metrics while yielding reasonably accurate segmentation results.
Xianchao Zhang 0004, Jinjia Guo, Nan Mu, Jingfeng Jiang
SMC3
2023 Dynamic Scale-free Graph Embedding via Self-attention
abstract
Graph neural networks (GNNs) have recently become increasingly popular due to their ability to learn node representations in complex graphs. Existing graph representation learning methods mainly target static graphs in Euclidean space, whereas many graphs in practical applications are dynamic and evolve continuously over time. Recent work has demonstrated that real-world graphs exhibit hierarchical properties. Unfortunately, many methods typically do not account for these latent hierarchical structures. In this work, we propose a dynamic network in hyperbolic space via self-attention, referred to as DynHAT, which leverages both the hyperbolic geometry and attention mechanism to learn node representations. More specifically, DynHAT captures hierarchical information by mapping the structural graph onto hyperbolic space, and time-varying dynamic evolution by flexibly weighting historical representations. Through extensive experiments on three real-world datasets, we show the superiority of our model in embedding dynamic graphs in hyperbolic space and competing methods in a link prediction task. In addition, our results show that embedding dynamic graphs in hyperbolic space has competitive performance when necessitating low dimensions.
Dingyang Duan, Daren Zha, Jiahui Shen, Nan Mu
J. Web Eng.5
2023 An attention residual u-net with differential preprocessing and geometric postprocessing: Learning how to segment vasculature including intracranial aneurysms
abstract
OBJECTIVE: Intracranial aneurysms (IA) are lethal, with high morbidity and mortality rates. Reliable, rapid, and accurate segmentation of IAs and their adjacent vasculature from medical imaging data is important to improve the clinical management of patients with IAs. However, due to the blurred boundaries and complex structure of IAs and overlapping with brain tissue or other cerebral arteries, image segmentation of IAs remains challenging. This study aimed to develop an attention residual U-Net (ARU-Net) architecture with differential preprocessing and geometric postprocessing for automatic segmentation of IAs and their adjacent arteries in conjunction with 3D rotational angiography (3DRA) images. METHODS: The proposed ARU-Net followed the classic U-Net framework with the following key enhancements. First, we preprocessed the 3DRA images based on boundary enhancement to capture more contour information and enhance the presence of small vessels. Second, we introduced the long skip connections of the attention gate at each layer of the fully convolutional decoder-encoder structure to emphasize the field of view (FOV) for IAs. Third, residual-based short skip connections were also embedded in each layer to implement in-depth supervision to help the network converge. Fourth, we devised a multiscale supervision strategy for independent prediction at different levels of the decoding path, integrating multiscale semantic information to facilitate the segmentation of small vessels. Fifth, the 3D conditional random field (3DCRF) and 3D connected component optimization (3DCCO) were exploited as postprocessing to optimize the segmentation results. RESULTS: Comprehensive experimental assessments validated the effectiveness of our ARU-Net. The proposed ARU-Net model achieved comparable or superior performance to the state-of-the-art methods through quantitative and qualitative evaluations. Notably, we found that ARU-Net improved the identification of arteries connecting to an IA, including small arteries that were hard to recognize by other methods. Consequently, IA geometries segmented by the proposed ARU-Net model yielded superior performance during subsequent computational hemodynamic studies (also known as "patient-specific" computational fluid dynamics [CFD] simulations). Furthermore, in an ablation study, the five key enhancements mentioned above were confirmed. CONCLUSIONS: The proposed ARU-Net model can automatically segment the IAs in 3DRA images with relatively high accuracy and potentially has significant value for clinical computational hemodynamic analysis.
Nan Mu, Zonghan Lyu, Mostafa Rezaeitaleshmahalleh, Jinshan Tang, Jingfeng Jiang
Medical Image Anal.1
2022 GSDM: A Gated Semantic Discriminating Model for Knowledge Graph Completion
abstract
Knowledge representation learning is an automatic learning technique that can embeds a knowledge graph into a low-dimensional vector space. With use of this, knowledge becomes computable and various intelligent applications can be realized. Traditional semantic discriminating models suggest that the embeddings of entities should depend on the specific semantic environment. We find that the multiple latent information of relations has not been put to use by these models. In this paper, a gated semantic discriminating model (GSDM) is proposed to select useful latent information and neglect useless information according to the specific semantic environment for both entities and relations. Experiments show that GSDM achieves better performance than related state-of-the-art baselines on most indicators. The better trade-off between the discriminate parameter pressure and the model performance has proved the correctness and feasibility of semantic discriminating mechanism to some extent.
Neng Gao, Nan Mu, Yao Dong 0003, Lei Wang 0135, Yuanye He
CSCWD3
2022 Decision Tree Fusion and Improved Fundus Image Classification Algorithm
Yanhua Qiu, Jialing Wu, Qianying Zou, Nan Mu
GPC6
2022 Traffic Sign Image Segmentation Algorithm Based on Improved Spatio-Temporal Map Convolution
Qianying Zou, Nan Mu
GPC5
2022 BEFSR: A Multiple Attention-Based Model Considering Bidirectional Entity Information Flows and Few-Shot Relations
abstract
The traditional knowledge representation learning (KRL) models treat each triplet in a knowledge base independently, so they can not make full use of the neighborhood information across triplets. KRL models based on graph attention networks (GAT) can not only capture feature interactions across triplets, but also further distinguish the importance of neighbor entities. Recently, we find that there are two flaws in GAT-based KRL models: (1) Ignoring the bidirectionality of information flows leads to insufficient utilization of entity neighborhood information. When encapsulating the neighborhood information, only the forward information flows flowing into the target entity are considered, but the backward information flows flowing out are neglected. (2) The unified update process for all relations causes the useful information related the few-shot relations to be diluted. We propose a multiple attention-based model considering bidirectional entity information flows and few-shot relations (BEFSR). In our model, a GAT-based attention framework is used to integrate forward information flows and backward information flows of each triplet respectively to capture feature interactions across triplets, and an LSTM-based attention framework is adopted to gradually aggregate the entity-pair information for few-shot relations’ updating. In BEFSR, entities and relations can be updated more appropriately. Experiments demonstrate that BEFSR outperforms state-of-the-art KRL models in knowledge base completion task.
Neng Gao, Fali Wang, Nan Mu, Lei Wang 0135, Yao Dong 0003
ICPR4
2022 Dynamic Network Embedding in Hyperbolic Space via Self-attention
Dingyang Duan, Daren Zha, Nan Mu, Jiahui Shen
ICWE4
2022 MBNet: Detecting Salient Object in Low-Light Scenes
abstract
Benefiting from the powerful features created by deep learning techniques, salient object detection has recently made significant progress. Compared with the task of detecting salient object in well-light images, the detection of salient object in low-light scenes requires not only acquiring the spatial visual saliency of images under low light conditions, but also accurately identifying the multi-scale objects of interest. Mountain Basin Network (MBNet) is proposed for salient object detection to discriminate the pixel-level saliency of low-light images. To further refine the object localization and pixel classification performance, the proposed model integrates a high-low feature aggregation module (HLFA) to synergize the information from a high level branch (named Bal-Net) and a low level branch (named Mol-Net) to fuse the global and local context, and the hierarchical supervision modules (HSM) is embedded to assist in obtaining accurate salient objects, especially the small ones. Furthermore, multi-supervised integration strategy is leveraged to optimize the structure and boundaries of salient objects. Meanwhile, to facilitate further research and evaluation of the visual saliency models, we construct a new low-light dataset, which includes 13 categories with a total of 1000 low-light images. The experimental results show that the proposed model has state-of-the-art low-light saliency detection performance compared with seven existing methods.
Yiyue Hu, Nan Mu
SMC3
2021 First-order and High-order Information Fusion over Heterogeneous Information Network for Top-N Recommendation System
abstract
In recent years, more and more researchers pay attention to the recommendation system based on heterogeneous information network(HIN), because HIN is rich in various kinds of information, which can significantly improve the performance of the recommendation system. But the HIN based recommendation system faces the following problems: how to leverage high-order information to get semantic-level interaction characteristics between users and items; How to deeply fuse first-order and high-order information to enhance the representation ability of the system. To address these issues, we propose a novel model: First-order and High-order Information Fusion over Heterogeneous Information Network for Top-N Recommender System(FHRec). For first-order information, we use graph neural networks to generate the latent vectors of users and items. And for high-order information, we use a meta-path based semantic-level aggregation to get the interaction between users and items. Then we deeply integrate the first-order and high-order information and use neural collaborative filtering to improve the recommendation performance. Finally, we conduct comparative experiments of our model with other baseline algorithms on three real world datasets, and the experimental results prove the superiority of our model.
Nan Mu, Daren Zha
CSCWD1
2021 Gated Knowledge Graph Neural Networks for Top-N Recommendation System
abstract
In recent years, the knowledge graph based recommendation system is a research hotspot and scholars propose a propagation-based method, which combines graph neural networks with knowledge graph. But the previous work faces two problems: The first is that exist propagation methods generate neighours of target entity through random sampling strategy which will bring noise to the system. The second problem is that exist models directly aggregate the neighbors information of the target entity at each step, while ignoring the fact that the propagation of high-order information also needs to be selective and memorable. To solve these problems, we propose a novel model: Gated Knowledge Graph Neural Networks for Top-N Recommendation System(GKGNN). This model uses pretrain technique to generate neighbor set with high priority of the target entity in the graph. At the same time, this model introduces the gated mechanism into the propagation process and the valuable information is remembered and unimportant information is forgotten during the high-order information propagation. Finally, we conduct comparative experiments of our model with other baseline algorithms on three real world datasets, and the experimental results prove the superiority of our model.
Nan Mu, Daren Zha
CSCWD1
2021 Graph Attention Autoencoder for Collaborative Pair-wise Ranking
abstract
Recently, top-k recommendation system is getting more and more attention from researchers and unlike the rating prediction task, the purpose of top-k recommendation is to present the user with a list of items they are most interested in. Many rating prediction models do not produce good ranking results, so the top-k recommendation faces two challenges: the first is how to accurately obtain user and item latent vector from the user-item interaction rating matrix, and the second challenge is how to combine the ranking strategy with the recommendation model deeply. The booming deep learning technology, such as autoencoder and graph neural networks, brings us new solutions. So in this paper, we propose a novel model: Graph Attention Autoencoder for Collaborative Pair-wise Ranking. The autoencoder consists of graph attention encoder and collaborative neural decoder, which is used to generate user and item latent vector accurately. And then we use the pairwise ranking learning process to ensure the rating value accuracy and rating value pairwise ranking consistency. Finally, experimental results on three real-world datasets demonstrate the superiority of our model.
Nan Mu, Daren Zha, Yuanye He
CSCWD1
2021 Video Salient Object Detection Network with Bidirectional Memory and Spatiotemporal Constraints
abstract
Although deep learning-based video salient object detection networks have shown outstanding performance, there still exist problems such as the incompleteness of salient objects due to the lack of spatial saliency information, and the globally inconsistency of saliency results detected in long-term video sequences. To address these issues, we propose a novel video salient object detection network with bidirectional memory and spatiotemporal constraints. The global and local cues are first cascaded to extract spatial location information and internal structural details by residual skip connections. After obtaining reliable and fine-grained spatial saliency for a single frame, we design a DB-ConvLSTM module with bidirectional memory that retains the past and future information. A global temporal attention mechanism is added in the temporal dimension to correlate each frame with the whole video to obtain accurate salient objects of globally consistent. Extensive experiments have been conducted to demonstrate that the proposed model outperforms the other seven state-of-the-art models on four datasets, viz. DAVIS, FBMS, SegTrackV2, and UVSD.
Nan Mu
SMC2
2021 A Non-local Hierarchical Refinement Fully Convolutional Network for COVID-19 Infected Region Segmentation
abstract
To control the spread of COVID-19, accurate and efficient diagnosis of suspected cases is the crux of appropriate quarantine and prompt treatment. In addition to the diagnosis of COVID-19, it is vital to distinguish the severity of the confirmed cases, which is conducive to the selection and planning of treatment methods. At a critical juncture of the epidemic crisis, we are committed to developing a non-local hierarchical refinement fully convolutional network to assist experts in automatically segmenting the COVID-19-caused pneumonia infection regions. The architecture of the proposed model is a deep encoder-decoder framework. Specifically, we exploit a non-local perception module to capture more complementary coarse-structure information from different pyramid levels and a local refinement model to explicitly heighten the fine-detail information of each convolution layer. Moreover, the non-local and local features are aggregated through a boundary-aware multiple supervision strategy to produce the gratifying edge-preserving segmentation map. Comprehensive experiments demonstrate the effectiveness of the proposed model in boosting the learning ability of accurately identify the lung infected regions with clear contours. In particular, our model is superior remarkably to the state-of-the-art segmentation models both quantitatively and qualitatively on a real CT dataset of COVID-19.
Nan Mu
SMC2
2021 Progressive global perception and local polishing network for lung infection segmentation of COVID-19 CT images
Nan Mu, Jingfeng Jiang, Jinshan Tang
Pattern Recognit.1
2020 Collaborative Denoising Graph Attention Autoencoders for Social Recommendation
Nan Mu, Daren Zha
SEKE1
2019 Saliency Computational Model for Foggy Images by Fusing Frequency and Spatial Cues
abstract
A key challenge of saliency computation in foggy images is how to effectively detect salient objects which are less visible. The primary cause may lie in the fact that the light scattering through fog particles reduces image contrast. In this paper, we propose a frequency-spatial fusion saliency computational model based on discrete stationary wavelet transform (DSWT). The input image is firstly transformed into HSV color space, and the amplitude spectrum of each color channel is adjusted to generate the frequency domain saliency map. Then, the local-global superpixel contrast is measured to obtain the spatial domain saliency map. The DSWT is finally utilized to fuse the frequency-spatial cues. Experimental results indicate that the proposed model can efficiently reduce the influence of light scattering through fog particles, and can achieve the best performance in foggy images comparing to 16 state-of-the-art saliency models.
Xin Xu 0007, Nan Mu, Li Chen 0011, Jing Tian 0002
ICIP3
2019 Graph Attention Networks for Neural Social Recommendation
abstract
In recent years, social recommendation is a research hotspot because it contains social network information which can effectively solve the problem of data sparsity and cold start. But the social recommendation task faces two problems: one is that how to accurately learn user latent vector and item latent vector from user-item interaction graph and social graph, the other is that how to depict the intrinsic and complex interaction between users and items. With the development of graph neural networks, node embedding is becoming more and more accurate on the graph. Besides neural collaborative filtering explores the interaction of users and items deeply. So in this paper, we propose a novel model: graph attention networks for neural social recommendation (GAT-NSR). This model adopts multi-head attention mechanism for message passing on the two graphs, which get user & item latent vector from different perspectives. And also we design a neural collaborative recommendation module to capture the inherent characteristics of user-item interaction behavior for recommendation. Finally, detailed experimental results on two real-world datasets clearly prove the effectiveness of our proposed model.
Nan Mu, Daren Zha, Yuanye He, Zhihao Tang 0001
ICTAI1
2019 A Multi-granularity Neural Network for Answer Sentence Selection
abstract
In open-domain question answering system, the granularities of the answers vary with different types of questions. For example, for the questions asking about locations (Location type questions), their answers are usually short phrases. While for the questions asking about reasons (Description type questions), their answers are usually long clauses or sentences. This insight can be used to improve the performance of answer sentence selection, which is a crucial component of the open-domain QA system. In this paper, we propose a novel Multi-Granularity Neural Network (MGNN) model to better evaluate the semantic matching of questions and answers. First, MGNN has three classes of channels with each computing the similarity of question and answer pairs from one of the three granularity levels: clause level, phrase level and ngram level. Then, MGNN uses a parametrization weighting scheme which considers question types to combine these different granularity channels. We carry out experiments on a public available benchmark dataset for question answering. Empirical results show that our method outperforms state-of-the-art methods.
Chenggong Zhang, Weijuan Zhang, Daren Zha, Pengjie Ren, Nan Mu
IJCNN5
2019 Adaptive-Skip-TransE Model: Breaking Relation Ambiguities for Knowledge Graph Embedding
Shoukang Han, Lei Wang 0135, Zeyi Liu 0002, Nan Mu
KSEM (1)5
2019 A spatial-frequency-temporal domain based saliency model for low contrast video sequences
Nan Mu, Xin Xu 0007, Xiaolong Zhang 0002
J. Vis. Commun. Image Represent.1
2019 Finding autofocus region in low contrast surveillance images using CNN-based saliency algorithm
Nan Mu, Xin Xu 0007, Xiaolong Zhang 0002
Pattern Recognit. Lett.1
2018 Discrete stationary wavelet transform based saliency information fusion from frequency and spatial domain in low contrast images
Nan Mu, Xin Xu 0007, Xiaolong Zhang 0002, Xiaoli Lin
Pattern Recognit. Lett.1
2017 Particle Swarm Optimization Based Salient Object Detection for Low Contrast Images
Nan Mu, Xin Xu 0007, Xiaolong Zhang 0002, Li Chen 0011
ICONIP (3)1
2016 Covariance descriptor based convolution neural network for saliency computation in low contrast images
abstract
Saliency computational model with active environment perception can substantially facilitate a wide range of applications. Conventional saliency computational models primarily rely on hand-crafted low level image features, such as color or contrast. However, they may face great challenges in low lighting scenario, due to the lack of well-defined feature to represent saliency information in low contrast images. In this paper, we propose a novel deep neural network framework embedded with covariance descriptor for salient object detection in low contrast images. Several low-level features are extracted to compute their mutual covariance, which is then trained via a 7-layers convolutional neural network (CNN). The saliency map can be generated by estimating the final saliency score of each region via the pre-trained CNN model. Extensive experiments have been conducted on six challenging datasets to evaluate the performance of the proposed model against ten state-of-the-art models.
Xin Xu 0007, Nan Mu, Xiaolong Zhang 0002, Bo Li 0002
IJCNN2
2016 Hierarchical salient object detection model using contrast-based saliency and color spatial distribution
Xin Xu 0007, Nan Mu, Li Chen 0011, Xiaolong Zhang 0002
Multim. Tools Appl.2
2015 Hierarchical Features Fusion for Salient Object Detection in Low Contrast Images
Nan Mu, Xin Xu 0007
ICIC (2)1
2015 Salient object detection from distinctive features in low contrast images
abstract
Saliency computational model with active environment perception can be useful for many applications including image segmentation, image compression, image retrieval, and etc. Conventional saliency computational models rely on handcrafted low level features, such as color or contrast. These models face great difficulties in low lighting scenarios, due to the lack of well-defined feature to interpret saliency information in low contrast images. In this paper, a new approach is proposed to detect salient object from low contrast images. The proposed approach explores the most distinguishable salient information in low contrast images based on low level features. Extensive experiments have been conducted to evaluate the performance of the proposed method against the state-of-the-art saliency computational models.
Xin Xu 0007, Nan Mu, Hong Zhang 0022, Xiaowei Fu
ICIP2
2014 Block-Based Salient Region Detection Using a New Spatial-Spectral-Domain Contrast Measure
abstract
Visual saliency is an important cue in human visual system, it can identify salient region in image. Image contrast has been utilized as an effective feature to detect the salient region. The conventional contrast measures utilize both spectral and spatial properties of image in many salient region detection methods. However, they only consider the local characteristics of image region, consequently, the global characteristics are neglected. This paper presented a new contrast measure by exploiting both local and global characteristics of image regions. Furthermore, the proposed measure is utilized to perform salient region detection in image. Experiments are conducted on the MSRA test database to compare the performance of the proposed approach with the state-of-the-art salient region detection algorithms.
Nan Mu, Xin Xu 0007, Li Chen 0011, Jing Tian 0002
ISM1
2013 Theil-Equilibrium based Cooperation Mechanism for multi-services in ubiquitous stub enironments
Nan Mu, Lanlan Rui, Shao-Yong Guo 0001, Xuesong Qiu 0001
APNOMS1