Zhen Wang 0020

dblp:78/6727-20 · DBLP profile ↗
← Back
26ranked-venue papers
12as first author
25since 2021 · last 2027
0000-0002-5765-0827ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2027 GeoMoE: Geometry-Driven prompts with adaptive Mixture-of-Experts for remote sensing image semantic segmentation
Zhen Wang 0020, Nan Xu 0008, Zhu-Hong You
Expert Syst. Appl.1
2026 MSCANet: Multi-scale Cross-Attention Fusion Network for Multimodal Remote Sensing Image Semantic Segmentation
Zhen Wang 0020, Zhu-Hong You, Yu Li 0030
ICIC (8)3
2026 Cos-UMamba: Optimizing salient object detection with cosine scanning and bias-corrected feature fusion in optical remote sensing images
Zhen Wang 0020, Fu-Lin He, Nan Xu 0008, Zhu-Hong You
Expert Syst. Appl.1
2025 CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmark
abstract
While Large Language Models (LLMs) excel in question-answering (QA) tasks, their real reasoning abilities on multiple evidence retrieval and integration on Multi-hop QA tasks remain less explored. Firstly, LLMs sometimes generate answers that rely on internal memory rather than retrieving evidence and reasoning in the given context, which brings concerns about the evaluation quality of real reasoning abilities. Although previous counterfactual QA benchmarks can separate the internal memory of LLMs, they focus solely on final QA performance, which is insufficient for reporting LLMs' real reasoning abilities. Because LLMs are expected to engage in intricate reasoning processes that involve evidence retrieval and answering a series of sub-questions from given passages. Moreover, current factual Multi-hop QA (MHQA) benchmarks are annotated on open-source corpora such as Wikipedia, although useful for multi-step reasoning evaluation, they show limitations due to the potential data contamination in LLMs' pre-training stage. To address these issues, we introduce the Step-wise and Counterfactual benchmark (CofCA), a novel evaluation benchmark consisting of factual data and counterfactual data that reveals LLMs' real reasoning abilities on multi-step reasoning and reasoning chain evaluation. Our experimental results reveal a significant performance gap of several LLMs between Wikipedia-based factual data and counterfactual data, deeming data contamination issues in existing benchmarks. Moreover, we observe that LLMs usually bypass the correct reasoning chain, showing an inflated multi-step reasoning performance. We believe that our CofCA benchmark will enhance and facilitate the evaluations of trustworthy LLMs.
Jian Wu 0037, Linyi Yang, Zhen Wang 0020, Manabu Okumura, Yue Zhang 0004
ICLR3
2025 Fine-Tuning SAM for Forward-Looking Sonar With Collaborative Prompts and Embedding
abstract
The Segment Anything Model (SAM) represents a significant advancement in semantic segmentation, particularly for natural images, but encounters notable limitations when applied to forward-looking sonar (FLS) images. The primary challenges lie in the inherent boundary ambiguity of FLS images, which complicates the use of prompt strategies for accurate boundary delineation, and the lack of effective interaction between prompts and image features. In this letter, we introduce a collaborative prompting strategy to address these issues by generating dense prompt embeddings and sonar tokens that focus on contour and boundary features, thereby replacing the original dense prompt embedding and IoU token. To further enhance segmentation, we employ embedding compensation techniques based on Mamba and KAN, which increase boundary information to image embedings and improve the fusion of prompts within image embeddings. We conducted comprehensive experiments, including comparative analyses and ablation studies, to validate the superiority of our proposed approach. Results show that our method significantly improves segmentation performance for FLS images, effectively addressing boundary ambiguity and optimizing prompt utilization. The source code and dataset will be available on https://github.com/darkseid-arch/FLSSAM.
Zhen Wang 0020, Nan Xu 0008, Zhu-Hong You
IEEE Geosci. Remote. Sens. Lett.2
2025 Vision Foundation Model-Driven Multiscale Expert Tuning for Multimodal Remote Sensing Semantic Segmentation
abstract
Multimodal remote sensing semantic segmentation based on Optical and Digital Surface Model (Opt-DSM) data is pivotal for comprehensive scene interpretation. However, prevailing methodologies often lack a unified vision foundation model and encounter significant challenges in bridging modality gaps and achieving effective feature fusion. Conventional models, such as the Segment Anything Model (SAM), exhibit inherent limitations when addressing the unique complexities of multimodal remote sensing, particularly in managing cross-modal discrepancies and intricate surface structures. In this study, we present VF-MET (Vision Foundation Model-Driven Multi-Scale Expert Tuning), an innovative framework meticulously tailored for Opt-DSM semantic segmentation tasks. VF-MET incorporates an adaptive Multi-Scale Expert Tuning (AMET) strategy, which substantially enhances the feature extraction capabilities of vision foundation models. This enables the robust capture of cross-scale and morphologically irregular objects, while simultaneously preserving superior generalization ability. To further address the segmentation of densely distributed and weakly correlated regions, we propose a collaborative Box-Point Prompt Mechanism (CBPM), which significantly improves spatial localization and contextual discrimination. Moreover, we introduce a Two-Stage Mask Decoder (TSMD) that facilitates efficient multimodal feature fusion and augments contextual understanding. Extensive experiments conducted on public Opt-DSM benchmark datasets unequivocally demonstrate that VF-MET achieves state-of-the-art performance. Comprehensive ablation studies further substantiate the indispensable contributions of each constituent module within the proposed architecture. The source code and datasets are publicly accessible at https://github.com/NWPUFranklee/VF-MET.git.
Zhen Wang 0020, Nan Xu 0008, Zhu-Hong You, De-Shuang Huang
IEEE Trans. Geosci. Remote. Sens.2
2025 Combining Airborne LiDAR Data and Optical Imagery for Improved National-Scale Beach Topography Estimation: A Case Study in New Zealand
abstract
Accurate beach topography mapping is crucial for understanding coastal dynamics and mitigating climate change impacts. However, traditional methods such as airborne LiDAR have limitations, leading to substantial gaps in national-scale elevation data. This study presents an innovative framework to reconstruct missing elevation data along New Zealand’s coastline by integrating airborne LiDAR, Sentinel-2 optical imagery, and geometric features (distance) using machine learning methods. Our results show that Artificial Neural Network (ANN) emerged as the best model (test set: R²=0.79, RMSE=0.91 m; validation set: 0.79, RMSE=0.93 m), outperforming other models in accuracy. The produced 10-m DEM for national-scale sandy beaches expands area coverage by 286.6% (114.15 km²), filling gaps in 1249 beaches, including remote areas such as Stewart Island. This novel framework offers a scalable solution for improving the comprehensiveness and accuracy of beach topography. It provides essential support for inundation prediction, habitat management, and the development of climate adaptation strategies, thereby facilitating more informed decision-making in coastal zone management and climate change mitigation efforts.
Conghong Huang, Yue Ma 0002, Xin Ma 0007, Yifu Ou, Chunpeng Chen, Shaoguang Zhou, Dongzhen Jia, Zhen Wang 0020, Qingquan Li 0001, Nan Xu 0008
IEEE Trans. Geosci. Remote. Sens.10
2025 FDMamba: Frequency-Driven Dual-Branch Mamba Network for Road Extraction From Remote Sensing Images
abstract
Road extraction from remote sensing imagery is crucial for a variety of applications, including transportation monitoring, disaster response, and urban planning. However, existing methods often fail to accurately delineate sparse, curvilinear, and boundary-blurred road structures in high-resolution images, leading to incomplete detail preservation and inadequate contextual understanding. To address these challenges, we propose a novel Frequency-Driven Dual-Branch Mamba Network (FDMamba) for precise road extraction from remote sensing imagery. The proposed FDMamba integrates frequency-aware modeling with a dual-branch architecture, enabling collaborative learning of fine-grained edge details and global spatial dependencies. Specifically, FDMamba comprises three key modules: a Fourier Reconstruction Attention Mechanism (FRAM) to enhance high-frequency boundary information and low-frequency structural representation; a Rotation-aware Mamba Module (RAMamba) that leverages multi-path state space modeling for robust directional perception of road structures; and a Phase-guided Feature Fusion Module (PFFM) for effective cross-scale alignment and fusion of high- and low-frequency features. Furthermore, to mitigate the issue of blurred or ambiguous boundaries, we introduce a hybrid loss function that combines binary cross-entropy, focal loss, and frequency-aware loss, explicitly guiding the model to focus on edge structure and multi-frequency complementary information. Extensive experiments on three benchmark datasets, CHN6-CUG, DeepGlobe, and Massachusetts, demonstrate that FDMamba consistently outperforms state-of-the-art methods in terms of F1-score and IoU, achieving superior boundary clarity and structural continuity while preserving overall geometric integrity. The code is available at https://github.com/darkseid-arch/RE-FDMamba.
Zhen Wang 0020, Shen-Ao Yuan, Nan Xu 0008, Zhu-Hong You, De-Shuang Huang
IEEE Trans. Geosci. Remote. Sens.1
2025 UAVSeg: Dual-Encoder Cross-Scale Attention Network for UAV Images' Semantic Segmentation
abstract
Benefiting from the powerful feature extraction and feature correlation modeling capabilities of convolutional neural networks (CNNs) and Transformer models, these techniques have been widely used in unmanned aerial vehicle (UAV) aerial image semantic segmentation tasks. However, the ground objects in aerial images contain feature information with different scales, and existing methods directly cascade low-level visual features and high-level semantic features without processing, resulting in low semantic segmentation precision. To address these challenges, we propose a dual-encoder cross-scale attention network, which efficiently extracts local and global context information from aerial images and performs fine-grained fusion of multiscale features to improve semantic segmentation performance. First, we introduce the dual-CNN-Transformer encoder, which embeds the scan-focus window Transformer (SFWT) into CNNs as an auxiliary encoder to supplement the local feature information lost in the global context information extraction process. Second, the cross-scale lightweight integration (CSLI) module is designed, which uses a light dot-product attention mechanism (DPAM) to fusion multiscale features and reduce model calculation parameters. Finally, the linear multilayer perceptron (LMLP) is used to restore the feature map resolution while expanding the deconvolution receptive field. To validate the effectiveness of the proposed method, we conducted extensive experiments on real aerial scene datasets, including UAVid, Urban Drone, and AeroScapes. The experimental results show that our method achieves state-of-the-art performance while maintaining superior real-time efficiency. Implementation codes will be available athttps://github.com/darkseid-arch/UAVSeg.
Zhen Wang 0020, Zhu-Hong You, Nan Xu 0008, Chuanlei Zhang, De-Shuang Huang
IEEE Trans. Geosci. Remote. Sens.1
2025 Local-Global Structure-Aware Geometric Equivariant Graph Representation Learning for Predicting Protein-Ligand Binding Affinity
abstract
Predicting protein-ligand binding affinities is a critical problem in drug discovery and design. A majority of existing methods fail to accurately characterize and exploit the geometrically invariant structures of protein-ligand complexes for predicting binding affinities. In this study, we propose Geo-protein-ligand binding affinity (PLA), a geometric equivariant graph representation learning framework with local-global structure awareness, to predict binding affinity by capturing the geometric information of protein-ligand complexes. Specifically, the local structural information of 3-D protein-ligand complexes is extracted by using an equivariant graph neural network (EGNN), which iteratively updates node representations while preserving the equivariance of coordinate transformations. Meanwhile, a graph transformer is utilized to capture long-range interactions among atoms, offering a global view that adaptively focuses on complex regions with a significant impact on binding affinities. Furthermore, the multiscale information from the two channels is integrated to enhance the predictive capability of the model. Extensive experimental studies on two benchmark datasets confirm the superior performance of Geo-PLA. Moreover, the visual interpretation of the learned protein-ligand complexes further indicates that our model offers valuable biological insights for virtual screening and drug repositioning.
Zhu-Hong You, Xuequn Shang 0001, Lei Wang 0121, Zhen Wang 0020
IEEE Trans. Neural Networks Learn. Syst.7
2024 ControversialQA: Exploring Controversy in Question Answering
abstract
Controversy is widespread online. Previous studies mainly define controversy based on vague assumptions of its relation to sentiment such as hate speech and offensive words. This paper introduces the first question-answering dataset that defines content controversy by user perception, i.e., votes from plenty of users. It contains nearly 10K questions, and each question has a best answer and a most controversial answer. Experimental results reveal that controversy detection in question answering is essential and challenging, and there is no strong correlation between controversy and sentiment tasks. We also show that controversial answers and most acceptable answers cannot be distinguished by retrieval-based QA models, which may cause controversy issues. With these insights, we believe ControversialQA can inspire future research on controversy in QA systems.
Zhen Wang 0020, Peide Zhu, Jie Yang 0028
LREC/COLING1
2024 MRHF: Multi-stage Retrieval and Hierarchical Fusion for Textbook Question Answering
Peide Zhu, Zhen Wang 0020, Manabu Okumura, Jie Yang 0028
MMM (2)2
2023 SEBGLMA: Semantic Embedded Bipartite Graph Network for Predicting lncRNA-miRNA Associations
abstract
Identifying the association between long noncoding RNA (lncRNA) and micro‐RNA (miRNA) is of great significance for the treatment of diseases by interfering with the combination of miRNA and messenger RNA (mRNA). Although many efforts and resources have been invested to identify lncRNA‐miRNA associations (LMAs), clinical trials are still expensive and laborious. Nevertheless, the experiments also need to consult a large number of side effects. Therefore, novel computer‐aided models are urgently needed to predict LMAs. This paper proposed a semantic embedded bipartite graph network for predicting lncRNA‐miRNA associations (SEBGLMA), which provided a novel feature extraction method by integrating K‐mer segmentation, word2vec, Gaussian interaction profile (GIP), and graph convolution network (GCN). Concretely, the attribute characteristics of RNA sequences are extracted by K‐mer segmentation and word2vec modules. Afterward, the adjacent matrix is completed by GIP self‐similarity. Then, the attribute characteristics and adjacent matrix are fed into GCN for embedding behavior features. Finally, the features are sent into the rotation forest (RoF) for detecting potential LMAs. The average accuracy, precision, sensitivity, specificity, Matthews correlation coefficient, and F1‐Score are 87.09%, 87.66%, 87.03%, 87.84%, 74.18%, and 86.99% on the benchmark data set. For fairly validating the performance of our model, we conducted various comparisons with different classifiers. Furthermore, the case studies of hsa‐miR‐497‐5P and NONHSAT022145.2 are also established. The results of comparisons and case studies further illustrated that our method is anticipated to become a robust and reliable tool for the identification of LMAs.
Zhengyang Zhao 0002, Jie Lin 0005, Zhen Wang 0020, Jianxin Guo, Xinke Zhan, Chuan Shi 0005, Wenzhun Huang
Int. J. Intell. Syst.3
2023 Enhanced Graph Neural Network with Multi-Task Learning and Data Augmentation for Semi-Supervised Node Classification
abstract
Graph neural networks (GNNs) have achieved impressive success in various applications. However, training dedicated GNNs for small-scale graphs still faces many problems such as over-fitting and deficiencies in performance improvements. Traditional methods such as data augmentation are commonly used in computer vision (CV) but are barely applied to graph structure data to solve these problems. In this paper, we propose a training framework named MTDA (Multi-Task learning with Data Augmentation)-GNN, which combines data augmentation and multi-task learning to improve the node classification performance of GNN on small-scale graph data. First, we use Graph Auto-Encoders (GAE) as a link predictor, modifying the original graphs’ topological structure by promoting intra-class edges and demoting interclass edges, in this way to denoise the original graph and realize data augmentation. Then the modified graph is used as the input of the node classification model. Besides defining the node pair classification as an auxiliary task, we introduce multi-task learning during the training process, forcing the predicted labels to conform to the observed pairwise relationships and improving the model’s classification ability. In addition, we conduct an adaptive dynamic weighting strategy to distribute the weight of different tasks automatically. Experiments on benchmark data sets demonstrate that the proposed MTDA-GNN outperforms traditional GNNs in graph-based semi-supervised node classification.
Cheng Fan 0005, Buhong Wang, Zhen Wang 0020
Int. J. Pattern Recognit. Artif. Intell.3
2023 Hidden Feature-Guided Semantic Segmentation Network for Remote Sensing Images
abstract
For semantic segmentation of remote sensing images, convolutional neural networks (CNNs) have proven to be powerful tools. However, the existing CNN-based methods have the problems of feature information loss, serious interference by clutter information, and ignoring the correlation between different scale features. To solve these problems, this article proposes a novel hidden feature-guided semantic segmentation network (HFGNet) for remote sensing images, which achieves accurate semantic segmentation by hierarchically extracting and fusing valuable feature information. Specifically, the hidden feature extraction module (HFE-M) is introduced to suppress the salient feature representation to mine more valuable hidden features. Meanwhile, the multifeature interactive fusion module (MIF-M) establishes the correlation between different features to achieve hierarchical feature fusion. The multiscale feature calibration module (MSFC) is constructed to enhance the diversity and refinement representation of hierarchical fusion features. Besides, the local-channel attention mechanism (LCA-M) is designed to improve the feature perception capability of the object region and suppress background information interference. We conducted extensive experiments on the widely used ISPRS 2-D Semantic Labeling dataset and the 15-Class Gaofen Image dataset. Experimental results demonstrate that the proposed HFGNet has advantages over several state-of-the-art methods. The source code and models are available athttps://github.com/darkseid-arch/RS-HFGNet.
Zhen Wang 0020, Shanwen Zhang, Chuanlei Zhang, Buhong Wang
IEEE Trans. Geosci. Remote. Sens.1
2022 N24News: A New Dataset for Multimodal News Classification
abstract
Current news datasets merely focus on text features on the news and rarely leverage the feature of images, excluding numerous essential features for news classification. In this paper, we propose a new dataset, N24News, which is generated from New York Times with 24 categories and contains both text and image information in each news. We use a multitask multimodal method and the experimental results show multimodal news classification performs better than text-only news classification. Depending on the length of the text, the classification accuracy can be increased by up to 8.11%. Our research reveals the relationship between the performance of a multimodal classifier and its sub-classifiers, and also the possible improvements when applying multimodal in news classification. N24News is shown to have great potential to prompt the multimodal news studies.
Zhen Wang 0020, Xu Shan, Xiangxie Zhang, Jie Yang 0028
LREC1
2022 Lightweight Convolution Neural Network Based on Multi-Scale Parallel Fusion for Weed Identification
abstract
Accurate identification of weed species is the premise for controlling weeds in field. But it is a challenging task due to the complexity and high-dimensional nonlinearity of the weed images in natural field. Convolutional neural networks (CNNs) model has been widely applied to image identification, but most of the CNNs models have the problems of large parameters, low identification accuracy, and single feature scale. This paper presents a novel deep neural network structure, named as MPF-Net for weed species identification. In MPF-Net, firstly, the weed images is sent into two different scales of depthwise separable convolution layers; secondly, the parallel output feature information is cross-fused, and uses the residual learning structure to increase the network model depth and feature extraction ability; finally the lightweight model PL-Model and the scale reduction module SR-Model are stacked together to construct the lightweight network. We have performed extensive experiments on real weed datasets, and compared the proposed MPF-Net against several variations of lightweight networks. The experimental results on the weed image dataset show that the proposed method is effective and feasible for weed species identification.
Zhen Wang 0020, Jianxin Guo, Shanwen Zhang
Int. J. Pattern Recognit. Artif. Intell.1
2022 Adversarial Attacks and Defenses for Deep-Learning-Based Unmanned Aerial Vehicles
abstract
The introduction of deep learning (DL) technology can improve the performance of cyber–physical systems (CPSs) in many ways. However, this also brings new security issues. To tackle these challenges, this article explores the vulnerabilities of DL-based unmanned aerial vehicles (UAVs), which are typical CPSs. Although many research works have been reported previously on adversarial attacks of DL models, only few of them are concerned about safety-critical CPSs, especially regression models in such systems. In this article, we analyze the problem of adversarial attacks against DL-based UAVs and propose two adversarial attack methods against regression models in UAVs. The experiments demonstrate that the proposed nontargeted and targeted attack methods both can craft imperceptible adversarial images and pose a considerable threat to the navigation and control of UAVs. To address this problem, adversarial training and defensive distillation methods are further investigated and evaluated, increasing the robustness of DL models in UAVs. To our knowledge, this is the first study on adversarial attacks and defenses against DL-based UAVs, which calls for more attention to the security and safety of such safety-critical applications.
Jiwei Tian, Buhong Wang, Rongxiao Guo, Zhen Wang 0020, Kunrui Cao
IEEE Internet Things J.4
2022 Exploring Targeted and Stealthy False Data Injection Attacks via Adversarial Machine Learning
abstract
State estimation methods used in cyber–physical systems (CPSs), such as smart grid, are vulnerable to false data injection attacks (FDIAs). Although substantial deep learning methods have been proposed to detect such attacks, deep neural networks (DNNs) are highly susceptible to adversarial attacks, which modify input of DNNs with unnoticeable but malicious perturbations. This article proposes a method to explore targeted and stealthy FDIAs via adversarial machine learning. We pose FDIAs as sparse optimization problems to achieve initial attack objectives and remain stealthy during attacks. We propose a parallel optimization algorithm to efficiently solve the problems and explore additional sparse-state attacks. The experimental results show that for IEEE 14-bus and 118-bus systems, the success rate of two-state sparse attacks with small-scale targets is as high as 80%. In addition, the attack success rate can continue to increase as the number of attack states increases. The proposed attacks demonstrate that attackers can implement attacks that can bypass both bad data detectors and neural network detectors while keeping the initial attack objectives unchanged, which is a critical and urgent security threat in CPS.
Jiwei Tian, Buhong Wang, Zhen Wang 0020, Mete Ozay
IEEE Internet Things J.4
2022 Joint Adversarial Example and False Data Injection Attacks for State Estimation in Power Systems
abstract
Although state estimation using a bad data detector (BDD) is a key procedure employed in power systems, the detector is vulnerable to false data injection attacks (FDIAs). Substantial deep learning methods have been proposed to detect such attacks. However, deep neural networks are susceptible to adversarial attacks or adversarial examples, where slight changes in inputs may lead to sharp changes in the corresponding outputs in even well-trained networks. This article introduces the joint adversarial example and FDIAs (AFDIAs) to explore various attack scenarios for state estimation in power systems. Considering that perturbations added directly to measurements are likely to be detected by BDDs, our proposed method of adding perturbations to state variables can guarantee that the attack is stealthy to BDDs. Then, malicious data that are stealthy to both BDDs and deep learning-based detectors can be generated. Theoretical and experimental results show that our proposed state-perturbation-based AFDIA method (S-AFDIA) can carry out attacks stealthy to both conventional BDDs and deep learning-based detectors, while our proposed measurement-perturbation-based adversarial FDIA method (M-AFDIA) succeeds if only deep learning-based detectors are used. The comparative experiments show that our proposed methods provide better performance than state-of-the-art methods. Besides, the ultimate effect of attacks can also be optimized using the proposed joint attack methods.
Jiwei Tian, Buhong Wang, Zhen Wang 0020, Kunrui Cao, Mete Ozay
IEEE Trans. Cybern.3
2022 Global Perception Network for Salient Object Detection in Remote Sensing Images
abstract
Despite recent works that have achieved remarkable progress on salient object detection for natural scene images, to detect various types and scales of objects, complex backgrounds in remote sensing images are still challenging. In this study, a novel global perception network (GPNet) is constructed for the salient object detection of remote sensing images. The proposed GPNet includes a global perception module (GPM), an axial attention block (AAB), and a feature distillation structure (FDS). The GPM is used to preserve the relationships of the entire dataset, the AAB is designed to capture the dependencies between the space and channel, the FDS is introduced to enable the helpful multilevel information flow into deep layers to enhance feature generation, and the global and the local attention information are mutually fused to enhance the network mode. Extensive experiments on three public datasets demonstrate that the proposed method outperforms other compared state-of-the-art methods both qualitatively and quantitatively (https://github.com/liuyu1002/GPnet).
Yu Liu 0107, Shanwen Zhang, Zhen Wang 0020, Baoping Zhao, Lincheng Zou
IEEE Trans. Geosci. Remote. Sens.3
2022 Multiscale Feature Enhancement Network for Salient Object Detection in Optical Remote Sensing Images
abstract
Aircraft detection in synthetic aperture radar (SAR) images plays an essential role in satellite observation and military decisions. Due to discrete scattering properties, speckle noise interference, and various aircraft types, many existing methods struggle to achieve the desired detection performance. In this article, we propose an innovative semantic condition constraint guided feature aware network (SCFNet) for detecting different aircraft categories in SAR images. First, considering the discrete scattering properties of aircraft, we design a local-global feature aware module (LGA-M) and morphological-semantic feature aware module (MSF-M), which can effectively extract the fine-grained feature information contained in SAR images. Second, to effectively fuse different feature information, we construct a feature fusion pyramid (FFP), which uses different branches and paths to reasonably merge multiple feature information types and suppresses background information interference. Third, according to the structure characteristics of aircraft, the global coordinate attention mechanism (G-CAT) is presented to highlight foreground target features and suppress speckle noise interference. Finally, we construct semantic condition constraints, including constraint condition setting, semantic information calculation, and template matching, to improve aircraft localization and recognition accuracy. Extensive experiments demonstrate that the proposed SCFNet can obtain state-of-the-art performance on the SAR aircraft detection dataset, which achieves AP and F1 Score of 94.83% and 95.58%, respectively. The related implementation codes will be made publicly available at https://github.com/darkseid-arch/AirDetection.
Zhen Wang 0020, Jianxin Guo, Chuanlei Zhang, Buhong Wang
IEEE Trans. Geosci. Remote. Sens.1
2022 MLFFNet: Multilevel Feature Fusion Network for Object Detection in Sonar Images
abstract
Sonar image object detection is essential in underwater rescue and resource exploration. Although many convolution neural network (CNN)-based object detection algorithms have achieved great success in natural images. However, for underwater sonar images, problems, such as seabed reverberation noise interference, low proportion of foreground object region pixels, and poor imaging resolution, present considerable challenges to achieving accurate underwater object detection. To address these problems, we propose a novel sonar image object detector called the multilevel feature fusion network (MLFFNet). The detector consists of multiscale convolution module (MS-Conv), multilevel feature extraction module (ML-FEM), multilevel feature fusion module (ML-FFM), neighborhood channel attention mechanism (N-CAM), multiscale feature pyramid module (MS-FPN), and feature association module (FA). First, we use the MS-Conv to extract different scale feature information in the object region. Second, the ML-FEM and ML-FFM are used to obtain the local detail and global context features. Third, the N-CAM and MS-FPN are used to obtain the foreground objects’ semantic feature and position feature, and suppress the background region noise interference. Finally, we use the FA module to enhance the category and feature correlation of different objects. Extensive experiments are conducted on the real scene sonar image dataset. The experimental results demonstrate that MLFFNet performs better than other state-of-the-art object detection methods. Code and dataset are publicly athttps://github.com/darkseid-arch/SonarMLFFNet.
Zhen Wang 0020, Jianxin Guo, Leya Zeng, Chuanlei Zhang, Buhong Wang
IEEE Trans. Geosci. Remote. Sens.1
2022 SCFNet: Semantic Condition Constraint Guided Feature Aware Network for Aircraft Detection in SAR Images
abstract
Aircraft detection in synthetic aperture radar (SAR) images plays an essential role in satellite observation and military decisions. Due to discrete scattering properties, speckle noise interference, and various aircraft types, many existing methods struggle to achieve the desired detection performance. In this article, we propose an innovative semantic condition constraint guided feature aware network (SCFNet) for detecting different aircraft categories in SAR images. First, considering the discrete scattering properties of aircraft, we design a local-global feature aware module (LGA-M) and morphological-semantic feature aware module (MSF-M), which can effectively extract the fine-grained feature information contained in SAR images. Second, to effectively fuse different feature information, we construct a feature fusion pyramid (FFP), which uses different branches and paths to reasonably merge multiple feature information types and suppresses background information interference. Third, according to the structure characteristics of aircraft, the global coordinate attention mechanism (G-CAT) is presented to highlight foreground target features and suppress speckle noise interference. Finally, we construct semantic condition constraints, including constraint condition setting, semantic information calculation, and template matching, to improve aircraft localization and recognition accuracy. Extensive experiments demonstrate that the proposed SCFNet can obtain state-of-the-art performance on the SAR aircraft detection dataset, which achieves AP and F1 Score of 94.83% and 95.58%, respectively. The related implementation codes will be made publicly available at https://github.com/darkseid-arch/AirDetection.
Zhen Wang 0020, Nan Xu 0008, Jianxin Guo, Chuanlei Zhang, Buhong Wang
IEEE Trans. Geosci. Remote. Sens.1
2022 Fused Adaptive Receptive Field Mechanism and Dynamic Multiscale Dilated Convolution for Side-Scan Sonar Image Segmentation
abstract
Side-scan sonar (SSS) is a vital sensor for marine survey, which is widely used in military and civilian fields. The accurate segmentation of SSS images is critical in sonar image intelligent interpretation. Existing SSS image segmentation methods have several limitations, such as insufficient feature extraction, relatively worse segmentation results for tiny target categories, and serious interference by seabed reverberation noise and bright shadow region. To overcome these issues, we propose a novel encoder-decoder architecture SSS image segmentation method based on convolution neural network (CNN). First, we extract the multi-scale feature information contained in target region using the dynamic multi-scale dilated convolution (DMDC_Conv). Second, to further obtain the global and detail feature information, we construct the adaptive receptive field mechanism block (ARFM_Block). Third, we design a feature fusion attention mechanism block (FFAM_Block) to fuse high-level and low-level feature information with different scales and suppress background information interference. Final, we construct a tree structure optimization module (TSOM) to solve the problem of pixel misclassification and obtain refine SSS image segmentation results. Extensive experiments are carried out on the constructed real scene SSS image dataset. The experimental results show that the proposed method achieves 93.24% and 90.82% of MPA and MIoU, respectively, which outperforms other state-of-the-art methods and has a substantial advantage in inference speed and calculation parameters.
Zhen Wang 0020, Shanwen Zhang, Lutz Gross, Chuanlei Zhang, Buhong Wang
IEEE Trans. Geosci. Remote. Sens.1
2013 SFAPS: An R package for structure/function analysis of protein sequences based on informational spectrum method
abstract
The R package SFAPS has been developed for structure/function analysis of protein sequences based on information spectrum method. The informational spectrum method employs the electron-Ion interaction potential parameter as the numerical representation for the protein sequence, and obtains the characteristic frequency of a particular protein interaction after computing the Discrete Fourier Transform (DFT) for protein sequences. The informational spectrum method is often used to analyze protein sequences, so we developed this software, which is implemented as an add-on package to the freely available and widely used statistical language R. Our package is distributed as open source code for Linux, Unix and Microsoft Windows. It is released under the GNU General Public License.
Suping Deng, Jing-Hua Yuan, De-Shuang Huang, Zhen Wang 0020
BIBM4